Senior Site Reliability Engineer

Atarusgroup — United Kingdom · Posted ~2 hours ago

Senior Onsite

Skills

AWS Kubernetes Site reliability engineering Cloud infrastructure Edge infrastructure Software engineering Automation Observability Monitoring CI/CD Incident response

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior SRE opportunity at a fast-growing technology organization developing cloud-native systems at significant scale. You will own reliability across AWS and edge infrastructure, build software and automation, manage Kubernetes environments, improve observability and CI/CD, strengthen incident response, and eliminate reliability bottlenecks.

Highlights

Senior onsite role with compensation up to £130,000 plus equity, ownership of reliability at significant scale, and substantial technical scope across cloud, edge infrastructure, automation, observability, and developer tooling.

Description

🚀 Senior Site Reliability Engineer (SRE) 📍 London (Onsite) 💰 Up to £130K + Equity We’re partnered with one of the UK’s fastest-growing AI companies, building cloud-native video intelligence products deployed across a rapidly growing fleet of connected devices. They’re looking for a Senior SRE to own reliability across both their AWS platform and edge infrastructure — combining software engineering, automation and production infrastructure at serious scale. What You’ll Be Doing 👇 Build software and tooling to improve the reliability and scalability of production systemsOwn and evolve large-scale AWS & Kubernetes environmentsBuild observability, monitoring and telemetry across cloud and edge infrastructureImprove deployment, CI/CD and incident response processesAutomate provisioning and lifecycle management across connected devicesIdentify reliability bottlenecks and engineer them out of the platformImprove developer experience through internal tooling and automation Tech ⚙️ AWS • Kubernetes / EKS • Terraform / Pulumi • Python / Go • Grafana • Prometheus • IoT / Edge What They’re Looking For 🧠 Strong SRE or software engineering backgroundProduction coding experience with Python or GoDeep experience operating systems on AWS & KubernetesStrong Infrastructure-as-Code and automation experienceSolid understanding of observability, incident management and production reliabilityComfortable debugging complex distributed systemsEdge / IoT experience is a bonus, not essential Why It’s Interesting 💡 SRE role where you'll actually write software, not just manage infrastructureReliability challenges spanning cloud + thousands of physical devicesReal-world AI systems with meaningful scale and availability requirementsSmall engineering team with significant ownership and architectural influenceStrong commercial traction, significant funding and meaningful equity If you’re an SRE who enjoys building systems rather than babysitting them, this is worth a conversation.