Site Reliability Engineer

Atarusgroup — United Kingdom · Posted ~14 hours ago

Senior Onsite

Skills

AWS Kubernetes Observability CI/CD Automation Incident Management

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

We need a senior SRE to architect and operate scalable cloud and edge platforms, leveraging AWS and Kubernetes to ensure reliability of AI‑driven video intelligence services.

Highlights

Join a fast‑growing AI company to design and maintain highly reliable cloud and edge infrastructure for large‑scale video intelligence services.

Description

🚀 Senior Site Reliability Engineer (SRE) 📍 London (Onsite) 💰 Up to £130K + Equity We’re partnered with one of the UK’s fastest-growing AI companies, building cloud-native video intelligence products deployed across a rapidly growing fleet of connected devices. They’re looking for a Senior SRE to own reliability across both their AWS platform and edge infrastructure — combining software engineering, automation and production infrastructure at serious scale. What You’ll Be Doing 👇 Build software and tooling to improve the reliability and scalability of production systemsOwn and evolve large-scale AWS & Kubernetes environmentsBuild observability, monitoring and telemetry across cloud and edge infrastructureImprove deployment, CI/CD and incident response processesAutomate provisioning and lifecycle management across connected devicesIdentify reliability bottlenecks and engineer them out of the platformImprove developer experience through internal tooling and automation Tech ⚙️ AWS • Kubernetes / EKS • Terraform / Pulumi • Python / Go • Grafana • Prometheus • IoT / Edge What They’re Looking For 🧠 Strong SRE or software engineering backgroundProduction coding experience with Python or GoDeep experience operating systems on AWS & KubernetesStrong Infrastructure-as-Code and automation experienceSolid understanding of observability, incident management and production reliabilityComfortable debugging complex distributed systemsEdge / IoT experience is a bonus, not essential Why It’s Interesting 💡 SRE role where you'll actually write software, not just manage infrastructureReliability challenges spanning cloud + thousands of physical devicesReal-world AI systems with meaningful scale and availability requirementsSmall engineering team with significant ownership and architectural influenceStrong commercial traction, significant funding and meaningful equity If you’re an SRE who enjoys building systems rather than babysitting them, this is worth a conversation.