Senior Site Reliability Engineer

Pracyva Ltd — United Kingdom · Posted ~4 hours ago

Senior Contract Hybrid

Skills

Google Cloud Platform Kubernetes infrastructure automation observability CI/CD incident response SRE cloud engineering Terraform

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior Site Reliability Engineer is sought for a hands-on role focused on improving the reliability, scalability, security, and operational excellence of cloud-hosted services. You’ll work across Kubernetes, infrastructure automation, observability, CI/CD, incident response, and service improvement in a collaborative engineering environment.

Highlights

Hands-on work across cloud reliability, scalability, security, automation, observability, and continuous service improvement in a cross-functional engineering environment.

Description

Senior Site Reliability Engineer (GCP) Manchester, Leeds or Edinburgh |Hybrid | Full-time Contract : Inside IR35 ROLE PURPOSE Improve the reliability, scalability, security and operational excellence of cloud-hosted services on Google Cloud Platform. This is a hands-on engineering role spanning SRE practices, production Kubernetes, infrastructure automation, observability, CI/CD, incident response and continuous service improvement. About the role The Senior Site Reliability Engineer will work with Cloud Platform, Software Engineering, Product, Security and Service teams to design and operate resilient services. The role will use software engineering and automation to reduce operational toil, improve service health and turn incident learning into measurable reliability improvements. Key responsibilities Lead reliability, availability, scalability and performance improvements across GCP-hosted applications, platforms and shared services.Define and operate service level indicators, service level objectives and error-budget practices that connect technical health to customer impact.Design actionable observability using Dynatrace, including instrumentation, dashboards, distributed tracing, service health views and SLO-based alerting.Build modular, reusable and maintainable Terraform code for secure cloud infrastructure, platform services and environment provisioning.Administer production Kubernetes environments, covering cluster lifecycle, workload deployment, capacity, upgrades, networking and platform troubleshooting.Automate operational activities and repetitive support work using Python, Groovy, Bash or PowerShell to reduce toil and improve consistency.Build and enhance CI/CD pipelines using Jenkins, Azure DevOps, GitHub Actions or equivalent tooling, with automated quality and deployment controls.Lead incident response, complex troubleshooting, root-cause analysis and post-incident reviews; ensure corrective actions are tracked to completion.Embed security, resilience, monitoring and supportability into platform and application designs from the outset.Coach engineers, contribute reusable standards and patterns, and promote SRE and operational excellence across engineering communities.Core outcomes Reliable and measurable production services with clear ownership, SLOs and actionable alerts.Less manual effort through automation, self-service and repeatable engineering patterns.Faster restoration, stronger incident learning and sustained reduction in repeat failures. Essential skills and experience Senior Site Reliability Engineer (GCP) Capability Required experience Google Cloud Platform Strong hands-on GCP engineering and operations experience across production services. Relevant GCP certification is preferred. Site Reliability Engineering Demonstrable SRE experience improving availability, resilience, scalability, supportability and operational performance. Dynatrace & observability Instrumentation, dashboards, metrics, logs, traces, service health modelling and SLO-based alerting using Dynatrace. Terraform Strong Infrastructure as Code capability, including modular design, reusable components, state management, code review and maintainability. Kubernetes & containers Production Kubernetes administration and Docker knowledge, including deployments, upgrades, networking, capacity and troubleshooting. CI/CD engineering Hands-on Jenkins, Azure DevOps, GitHub Actions or comparable tooling for build, test, infrastructure and deployment automation. Scripting & coding Practical automation using Python, Groovy, Bash or PowerShell, with sound software engineering and source-control practices. Incident management Major incident response, structured troubleshooting, root-cause analysis, problem management and post-incident improvement. Cloud security Working knowledge of IAM, least privilege, secrets, secure configuration, policy controls and operational risk management. Networking & APIs Strong understanding of Linux/Unix, DNS, TCP/IP, routing, load balancing, firewalls, VPC connectivity, HTTP and API troubleshooting. DevOps & toil reduction Automation-first mindset and experience replacing repetitive operational activity with reliable engineering solutions. Stakeholder collaboration Ability to explain service risk and technical options clearly and work effectively across engineering, product, security and operations teams. Technical environment Cloud: Google Cloud Platform, cloud-native services and managed platform capabilities.Observability: Dynatrace, telemetry, dashboards, SLI/SLO measures, logs, metrics and distributed traces.Platform: Kubernetes, Docker, Linux/Unix, cloud networking and APIs.Automation: Terraform, Python, Groovy, Bash and PowerShell.Delivery: Jenkins, Azure DevOps, GitHub Actions, Git and automated release controls. Operational and engineering expectations Senior Site Reliability Engineer (GCP) Reliability and observability Create service health models that combine application, infrastructure and customer-impact signals.Use SLOs and error budgets to guide prioritisation between reliability improvement and feature delivery.Tune alerts to be actionable, reduce noise and provide clear runbooks and escalation paths.Perform production-readiness reviews covering capacity, dependencies, failure modes, recovery and monitoring.Platform automation and delivery Create reusable Terraform modules and automated environment patterns that are secure, reviewable and repeatable.Integrate infrastructure and application changes into CI/CD with validation, approvals, auditability and rollback options.Automate cluster operations, routine maintenance, evidence collection and common remediation activities.Use code reviews, testing and version control to maintain engineering quality across infrastructure and operational tooling.Incident and problem management Coordinate technical diagnosis during major incidents and communicate impact, findings and recovery actions clearly.Complete evidence-based root-cause analysis and convert learning into engineering backlog improvements.Identify recurring failure patterns and remove underlying causes rather than relying on repeated manual intervention.Improve runbooks, diagnostics and recovery automation based on operational experience.Desirable experience Professional GCP certification or equivalent evidence of advanced Google Cloud capability.OpenTelemetry, Google Cloud Monitoring, service mesh or distributed tracing experience.GitOps, policy as code, secrets management and automated compliance controls.Hybrid or multi-cloud environments and large-scale enterprise cloud networking.Financial services or another highly regulated, security-conscious environment.On-call operations, resilience testing, disaster recovery and capacity engineering.Candidate profile Hands-on engineer who can move confidently between code, cloud infrastructure, Kubernetes and production diagnostics.Uses evidence and service data to prioritise improvements and make reliability trade-offs transparent.Communicates calmly during incidents and builds constructive relationships across technical and non-technical teams.Promotes learning, automation and continuous improvement while maintaining strong security and control standards.Job title matches: Site Reliability Engineer (SRE) | Cloud SRE | Google Cloud SRE | DevOps Engineer with Observability | Cloud Platform Engineer | Observability Engineer Success measures Improved availability, service performance and operational resilience.Reduced alert noise, incident recurrence and manual operational toil.Consistent, secure infrastructure delivery through reusable automation.Clear service ownership, measurable SLOs and effective operational readiness.