Site Reliability Engineer - Kubernetes

Client Server — United Kingdom · Posted ~4 hours ago

Senior Full-time Hybrid Visa History ✓

Skills

Site reliability engineering Kubernetes Reliability engineering System observability Automation Tooling Incident investigation System resilience Performance monitoring Operational efficiency Observability Monitoring SRE tooling Infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An advanced cybersecurity organization is seeking a Site Reliability Engineer specializing in Kubernetes. You will embed reliability practices across teams, improve observability and resilience, build automation and operational tooling, monitor system performance, and investigate incidents and systemic issues. The role combines infrastructure engineering with opportunities to improve developer experience and operational efficiency.

Highlights

An SRE opportunity combining Kubernetes, reliability engineering, observability, automation, and complex incident investigation. The role provides scope to improve system resilience, reduce downtime, develop engineering tooling and best practices, and enhance developer experience within a collaborative technical environment.

Description

Site Reliability Engineer (SRE Kubernetes) Cambridge / WFH to £70k Opportunity to progress your career at the world's most advanced cybersecurity technology business that uses AI technology to protect clients across the globe from advanced cyber threats, working alongside a team of friendly and supportive people and enjoying a host of perks and benefits. As a Site Reliability Engineer / SRE you will work across teams to embed reliability best practices, improve system resilience and solve complex operational challenges. You will take a proactive approach to identifying risks, improving system observability and enhancing the developer experience through automation and tooling. Key responsibilities will include developing tooling, frameworks and best practices to improve reliability and operational efficiency, monitoring system performance and implementing improvements to reduce incidents and downtime and conducting deep-dive investigations into incidents and systemic issues, driving long-term fixes. Location / WFH: You'll join a highly talented, diverse team in the Cambridge office twice a week where you can enjoy a great team atmosphere with free lunches and problem solving sessions. About you: You have experience in Site Reliability Engineering, DevOps or infrastructure engineering You have a strong understanding of distributed systems, reliability principles and system design You can script / code with Go, Python or similarYou have experience with cloud platforms: AWS, Azure or GCPYou have experience with containerisation and orchestration including KubernetesYou're familiar with observability practices; monitoring, logging, tracingYou have strong problem solving and critical thinking skillsYou're collaborative with great communication skillsYou are degree educated, having achieved a 2.1 or above in a STEM discipline from a top tier university (e.g. Russel Group, Oxbridge) What's in it for you: As a Site Reliability Engineer / SRE you will earn a competitive package: Salary to £70kPensionPrivate Medical InsuranceLife AssuranceEnhanced parental leaveEmployee Assistance Program23 days holiday plus an additional one for your birthdayCharity giving schemesPersonal training and development budgets Apply now to find out more about this Site Reliability Engineer (SRE Kubernetes) opportunity.