JobMatch
← Back to jobs

Site Reliability Engineer

Hamilton Barnes 🌳

RemoteUnited KingdomGBP 120kseniorFull-time
Posted
today
Source
LinkedIn (remote, Europe)
Field
Engineering

Skills

KubernetesPythonJiraAI

Description

Job Title: Site Reliability Engineer (GPU Infrastructure) Location: Remote (UK Wide) Salary: Up to £120,000 depending on experience A high growth GPU cloud provider is hiring a Site Reliability Engineer to keep large scale production GPU infrastructure fast, stable and observable. This isn't a sandbox or legacy estate, you'd be working directly on live infrastructure supporting demanding AI workloads, sitting where reliability, networking and platform engineering meet. The role is fully remote from anywhere in the UK, within a small and technically strong engineering led team. What's in it for you? - Real production scope from day one on live GPU infrastructure, not a ticket queue or sandbox environment - Rare breadth across reliability, high performance networking and platform engineering, with room to grow into whichever area interests you most - Fully remote working from anywhere in the UK - A team focused on building durable systems and fixing root causes through automation, rather than repeating the same manual fixes every week Responsibilities - Own the reliability and performance of production GPU infrastructure, including the high performance networking underneath it - Design, build and maintain observability, monitoring and dashboards, improving signal quality through correlation, enrichment, suppression and deduplication - Build Python based automation for incident triage, runbook execution and routine operational tasks - Integrate observability, ITSM and infrastructure APIs to enrich alerts and automate operational workflows - Deliver internal tools and self service capabilities, including CLI utilities, ChatOps integrations and dashboards - Turn post incident learnings into better tooling, automation and operational standards Skills Required - Proven track record as an SRE or Production/Infrastructure Engineer with hands on GPU infrastructure in a live production environment - Strong Python for automation and tooling - Deep observability and dashboarding experience (Prometheus, Grafana or similar), with alert design and incident management - UK based and able to work fully remote Strong Differentiators - Low latency networking or InfiniBand experience - Platform engineering background (Kubernetes, infrastructure as code, internal developer platforms) - HPC infrastructure exposure, such as Slurm managed clusters or research computing environments - ChatOps and ITSM integration experience (Slack/Teams bots, ServiceNow or Jira Service Management APIs) Apply now if you want to work on large scare GPU infrastructure with real breadth of scope!

JobMatch aggregates public listings. Always apply through the original posting.