Head of Site Reliability Engineering

Mark Williams Recruitment โ€” United Arab Emirates ยท Posted ~3 days ago

Head Full-time

Skills

Site Reliability Engineering Cloud platforms Kubernetes Containers CI/CD Infrastructure automation Observability Python Go Bash GCP Docker Prometheus Grafana Datadog

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

A senior technology leadership position responsible for creating an SRE function, defining reliability standards, leading automation initiatives, and improving operational excellence across enterprise systems.

Highlights

Leadership role focused on building engineering capabilities, improving reliability, scaling cloud platforms, and mentoring technical teams.

Description

We are partnering with an organisation that is looking to appoint a Head of Site Reliability Engineering (SRE) to establish and lead its SRE function. This is an excellent opportunity for a highly technical leader who enjoys building teams, defining best practices, and driving reliability, scalability, and operational excellence across modern cloud platforms. We are looking for someone who can remain hands on while developing a high performing SRE capability. Key Responsibilities - Build and lead the Site Reliability Engineering function, defining the vision, operating model, and best practices - Develop SRE principles, standards, and processes to improve platform reliability, availability, and performance - Lead the adoption of cloud native technologies while supporting both modern cloud applications and legacy workloads - Partner with engineering, infrastructure, and product teams to improve system resilience, scalability, and operational efficiency - Drive automation across monitoring, deployment, incident management, and platform operations - Define and measure SLIs, SLOs, and error budgets to improve service reliability - Lead incident management, root cause analysis, and continuous improvement initiatives - Mentor and develop engineers, creating a strong engineering culture focused on reliability and operational excellence Requirements - 10+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Engineering - Experience building or scaling an SRE function within a large enterprise or technology organisation - Strong hands-on experience with Google Cloud Platform (GCP) and cloud native technologies - Proven experience supporting both cloud native applications and legacy enterprise workloads - Strong understanding of Kubernetes, containers, infrastructure automation, CI/CD, observability, and platform engineering - Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, or similar - Strong scripting or programming experience using languages such as Python, Go, or Bash