Summary
✨ AI‑Generated
An experienced Senior Site Reliability Engineer is sought for a hybrid role focused on production cloud platforms. You will use strong Azure engineering skills, automation, observability and operational practices to improve the reliability, security and scalability of business-critical systems.
Highlights
A highly hands-on senior SRE role focused on keeping business-critical cloud platforms observable, secure, scalable, measurable, and reliable. The position combines Azure expertise with automation and operational leadership in a technically important environment.
Description
Senior Site Reliability Engineer - Azure Cloud Platform
Southampton - Hybrid
We are working alongside a global provider of advanced SaaS solutions to organisations operating within the public safety and justice sectors.
Its technology supports multimedia evidence management and emergency contact centre operations for customers around the world.
Due to continued growth they are expanding its Cloud Platform Engineering function and are looking for an experienced Senior Site Reliability Engineer.
This is a highly hands-on position focused on ensuring that business-critical cloud platforms remain observable, measurable, secure, scalable and reliable.
The successful candidate will combine strong Azure engineering expertise with platform automation, observability and operational leadership.
Role Responsibilities
Work as part of the Site Reliability Engineering team responsible for protecting and improving production environments.Manage and prioritise a technical backlog of reliability, scalability and operational improvements.Lead investigations into service outages, performance degradation, platform reliability and cloud expenditure.Conduct detailed root-cause analysis and ensure corrective actions are implemented.Identify repetitive operational activities and replace them with sustainable automation.Provide technical leadership and guidance to Cloud Operations, Support, DevOps and Engineering teams.Establish and maintain service level objectives, service level agreements, service level indicators and error budgets.Design and implement monitoring, alerting and dashboarding across cloud platforms and microservices.Deploy and configure observability technologies including Grafana, Prometheus, Azure Monitor and OpenTelemetry.Develop custom application and platform metrics to improve operational visibility.Create advanced queries, dashboards and alerts for distributed microservices.Develop reusable Bicep or Terraform modules for monitoring and cloud infrastructure.Support and improve production Kubernetes environments, particularly Azure Kubernetes Service.Contribute to cloud architecture, technical scoping and the implementation of scalable platform solutions.Support continuous improvement across deployment, provisioning and operational processes.Help ensure platforms and working practices meet relevant security, governance and compliance requirements.Explore opportunities to use AI-assisted tools to improve automation, troubleshooting and engineering productivity.
Essential Skills and Experience
Demonstrable experience supporting business-critical cloud platforms and live production services.Strong hands-on knowledge of Microsoft Azure.Production experience with Kubernetes and containerised workloads, ideally using AKS.Extensive experience in platform engineering, cloud provisioning and observability.Strong monitoring, alerting and dashboarding experience using technologies such as: Azure Monitor, Grafana, Prometheus, OpenTelemetry, Elasticsearch
Experience creating custom metrics, queries, dashboards and alerts for microservices.Advanced scripting or software development skills using PowerShell, Python, C# or a comparable language.Strong Infrastructure as Code experience using Bicep, ARM or Terraform.Experience using Git or another version-control platform.Good knowledge of Microsoft SQL Server, Elasticsearch and structured data formats including YAML, JSON and XML.Strong understanding of microservices architecture, cloud platforms and containerisation.Experience defining or working with SLOs, SLAs, SLIs and error budgets.Excellent troubleshooting and root-cause analysis skills.Experience designing scalable, secure and maintainable cloud solutions.Strong understanding of cybersecurity principles, governance and compliance.Security requirement: Applicants must have lived in the UK continuously for the past five years and be eligible to obtain NPPV3 and UK Security Clearance.
Spectrum IT Recruitment (South) Limited is acting as an Employment Agency in relation to this vacancy.