Description
Senior Site Reliability Engineer (GCP)
Manchester, Leeds or Edinburgh |Hybrid | Full-time
Contract : Inside IR35
ROLE PURPOSE Improve the reliability, scalability, security and operational excellence of cloud-hosted services on Google Cloud Platform.
This is a hands-on engineering role spanning SRE practices, production Kubernetes, infrastructure automation, observability, CI/CD, incident response and continuous service improvement.
About the role
The Senior Site Reliability Engineer will work with Cloud Platform, Software Engineering, Product, Security and Service teams to design and operate resilient services.
The role will use software engineering and automation to reduce operational toil, improve service health and turn incident learning into measurable reliability improvements.
Key responsibilities
Lead reliability, availability, scalability and performance improvements across GCP-hosted applications, platforms and shared services.Define and operate service level indicators, service level objectives and error-budget practices that connect technical health to customer impact.Design actionable observability using Dynatrace, including instrumentation, dashboards, distributed tracing, service health views and SLO-based alerting.Build modular, reusable and maintainable Terraform code for secure cloud infrastructure, platform services and environment provisioning.Administer production Kubernetes environments, covering cluster lifecycle, workload deployment, capacity, upgrades, networking and platform troubleshooting.Automate operational activities and repetitive support work using Python, Groovy, Bash or PowerShell to reduce toil and improve consistency.Build and enhance CI/CD pipelines using Jenkins, Azure DevOps, GitHub Actions or equivalent tooling, with automated quality and deployment controls.Lead incident response, complex troubleshooting, root-cause analysis and post-incident reviews; ensure corrective actions are tracked to completion.Embed security, resilience, monitoring and supportability into platform and application designs from the outset.Coach engineers, contribute reusable standards and patterns, and promote SRE and operational excellence across engineering communities.Core outcomes
Reliable and measurable production services with clear ownership, SLOs and actionable alerts.Less manual effort through automation, self-service and repeatable engineering patterns.Faster restoration, stronger incident learning and sustained reduction in repeat failures.
Essential skills and experience
Senior Site Reliability Engineer (GCP)
Capability
Required experience
Google Cloud Platform
Strong hands-on GCP engineering and operations experience across production services.
Relevant GCP certification is preferred.
Site Reliability Engineering
Demonstrable SRE experience improving availability, resilience, scalability, supportability and operational performance.
Dynatrace & observability
Instrumentation, dashboards, metrics, logs, traces, service health modelling and SLO-based alerting using Dynatrace.
Terraform
Strong Infrastructure as Code capability, including modular design, reusable components, state management, code review and maintainability.
Kubernetes & containers
Production Kubernetes administration and Docker knowledge, including deployments, upgrades, networking, capacity and troubleshooting.
CI/CD engineering
Hands-on Jenkins, Azure DevOps, GitHub Actions or comparable tooling for build, test, infrastructure and deployment automation.
Scripting & coding
Practical automation using Python, Groovy, Bash or PowerShell, with sound software engineering and source-control practices.
Incident management
Major incident response, structured troubleshooting, root-cause analysis, problem management and post-incident improvement.
Cloud security
Working knowledge of IAM, least privilege, secrets, secure configuration, policy controls and operational risk management.
Networking & APIs
Strong understanding of Linux/Unix, DNS, TCP/IP, routing, load balancing, firewalls, VPC connectivity, HTTP and API troubleshooting.
DevOps & toil reduction
Automation-first mindset and experience replacing repetitive operational activity with reliable engineering solutions.
Stakeholder collaboration
Ability to explain service risk and technical options clearly and work effectively across engineering, product, security and operations teams.
Technical environment
Cloud: Google Cloud Platform, cloud-native services and managed platform capabilities.Observability: Dynatrace, telemetry, dashboards, SLI/SLO measures, logs, metrics and distributed traces.Platform: Kubernetes, Docker, Linux/Unix, cloud networking and APIs.Automation: Terraform, Python, Groovy, Bash and PowerShell.Delivery: Jenkins, Azure DevOps, GitHub Actions, Git and automated release controls.
Operational and engineering expectations
Senior Site Reliability Engineer (GCP)
Reliability and observability
Create service health models that combine application, infrastructure and customer-impact signals.Use SLOs and error budgets to guide prioritisation between reliability improvement and feature delivery.Tune alerts to be actionable, reduce noise and provide clear runbooks and escalation paths.Perform production-readiness reviews covering capacity, dependencies, failure modes, recovery and monitoring.Platform automation and delivery
Create reusable Terraform modules and automated environment patterns that are secure, reviewable and repeatable.Integrate infrastructure and application changes into CI/CD with validation, approvals, auditability and rollback options.Automate cluster operations, routine maintenance, evidence collection and common remediation activities.Use code reviews, testing and version control to maintain engineering quality across infrastructure and operational tooling.Incident and problem management
Coordinate technical diagnosis during major incidents and communicate impact, findings and recovery actions clearly.Complete evidence-based root-cause analysis and convert learning into engineering backlog improvements.Identify recurring failure patterns and remove underlying causes rather than relying on repeated manual intervention.Improve runbooks, diagnostics and recovery automation based on operational experience.Desirable experience
Professional GCP certification or equivalent evidence of advanced Google Cloud capability.OpenTelemetry, Google Cloud Monitoring, service mesh or distributed tracing experience.GitOps, policy as code, secrets management and automated compliance controls.Hybrid or multi-cloud environments and large-scale enterprise cloud networking.Financial services or another highly regulated, security-conscious environment.On-call operations, resilience testing, disaster recovery and capacity engineering.Candidate profile
Hands-on engineer who can move confidently between code, cloud infrastructure, Kubernetes and production diagnostics.Uses evidence and service data to prioritise improvements and make reliability trade-offs transparent.Communicates calmly during incidents and builds constructive relationships across technical and non-technical teams.Promotes learning, automation and continuous improvement while maintaining strong security and control standards.Job title matches:
Site Reliability Engineer (SRE) | Cloud SRE | Google Cloud SRE | DevOps Engineer with Observability | Cloud Platform Engineer | Observability Engineer
Success measures
Improved availability, service performance and operational resilience.Reduced alert noise, incident recurrence and manual operational toil.Consistent, secure infrastructure delivery through reusable automation.Clear service ownership, measurable SLOs and effective operational readiness.