Summary
✨ AI‑Generated
Join an inclusive reliability engineering team responsible for improving the operational characteristics of critical products and services. You will drive availability, performance, monitoring, security, incident response, capacity planning, and safe change management while partnering closely with feature teams and stakeholders. The role is predominantly remote with regular office collaboration.
Highlights
Mostly remote senior SRE role with only one office day per week, substantial stakeholder interaction, collaborative engineering practices, and strong emphasis on innovation and professional development.
Description
Senior Site Reliability Engineer (London) - We’re working in collaboration to source a Senior Site Reliability Engineer for a large UK client.
The role is mostly working remotely, with only 1 day per week being required to work in the London office.
In this key role, you’ll improve, drive, and embed non-functional and operational characteristics such as availability, performance, efficiency, change management, monitoring, security, incident response, and capacity planning of our products and servicesYou’ll enjoy significant stakeholder interaction, working in collaboration with engineers to ensure a principled approach to deliver change in a safe and secure wayThis is a chance to join an inclusive team with a collaborative ethos and a commitment to innovation and professional developmentYou'll work from home some of the time, but you'll also spend a significant amount of time working from an office or hub
What you'll do
Work closely with our feature team and other colleagues to meet defined service level objectives and continually improve systems and environments.Define error budgets that support finding the right balance between risk and reliability.Provide structure and help to our release process, suggesting and making improvements where possible.Help scale systems sustainably through mechanisms like automation, evolving them by pushing for changes that improve reliability and velocity.Coach and provide guidance to colleagues and the wider team, leading where required.
In addition to this, you’ll:
Proactively contribute new ideas and innovations to meet short-term and longer-term goalsContinually balance and manage any potential risksBe accountable for the day-to-day health of both production and non-production environments and respond to any incidents as requiredProvide technical expertise and input to establish the risk tolerance of products and servicesCommunicate incident status updates clearly and frequently to other teams, customers and stakeholders
The skills you'll need
At least 10 years of hands-on experience, including as a Senior SRE with a proactive approach to spotting problems, areas for improvement and performance bottlenecks.
Experience working with cloud-native microservices, including containerisation, management of Kubernetes workloads and API management.Hands-on experience with Azure, Infrastructure as Code (IaC), and technologies such as PowerShell, JSON, Azure Bicep, ARM and Azure DevOps.The client is moving to Terraform, which is essential, moving from Bicep (desirable).Experience with full-stack observability using tools such as Grafana Stack, Log Analytics, AppInsightsExcellent knowledge of DevOps processes and principlesKnowledge of IT Service Management and automation of IT fulfilment processes through Orchestration and ServiceNowStrong communication skills with the ability to proactively engage with a wide range of stakeholders
SALARY INCLUDES 10% Benefits-As-Cash