Site Reliability Engineer

Client Server — United Kingdom · Posted ~1 day ago

Mid Hybrid

Skills

Site Reliability Engineering Observability Automation Monitoring Incident investigation System resilience

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary

An advanced cybersecurity organization is seeking a Site Reliability Engineer to embed reliability practices across teams and improve the resilience of critical systems. You will develop automation and tooling, enhance observability, monitor performance, investigate incidents and systemic issues, and drive improvements that reduce downtime.

Highlights

Work on complex reliability challenges in an advanced cybersecurity environment, with opportunities to improve resilience, observability, automation, developer experience, and long-term operational efficiency while benefiting from a supportive team culture.

Description

Site Reliability Engineer SRE Cambridge / WFH to £70k Opportunity to progress your career at the world's most advanced cybersecurity technology business that uses AI technology to protect clients across the globe from advanced cyber threats, working alongside a team of friendly and supportive people and enjoying a host of perks and benefits. As a Site Reliability Engineer / SRE you will work across teams to embed reliability best practices, improve system resilience and solve complex operational challenges. You will take a proactive approach to identifying risks, improving system observability and enhancing the developer experience through automation and tooling. Key responsibilities will include developing tooling, frameworks and best practices to improve reliability and operational efficiency, monitoring system performance and implementing improvements to reduce incidents and downtime and conducting deep-dive investigations into incidents and systemic issues, driving long-term fixes. Location / WFH: You'll join a highly talented, diverse team in the Cambridge office twice a week where you can enjoy a great team atmosphere with free lunches and problem solving sessions. About You You have experience in Site Reliability Engineering, DevOps or infrastructure engineering You have a strong understanding of distributed systems, reliability principles and system design You can script / code with Go, Python or similarYou have experience with cloud platforms: AWS, Azure or GCPYou have experience with containerisation and orchestration e.g. KubernetesYou're familiar with observability practices; monitoring, logging, tracingYou have strong problem solving and critical thinking skillsYou're collaborative with great communication skills What's In It For You As a Site Reliability Engineer / SRE you will earn a competitive package: Salary to £70kPensionPrivate Medical InsuranceLife AssuranceEnhanced parental leaveEmployee Assistance Program23 days holiday plus an additional one for your birthdayCharity giving schemesPersonal training and development budgets Apply now to find out more about this SRE Lead (Site Reliability Engineer) opportunity.