Summary
✨ AI‑Generated
Take ownership of the reliability and operational excellence of cloud-hosted services in a hands-on senior engineering role. You will improve availability, scalability, performance, and security while working with Kubernetes, infrastructure automation, observability, CI/CD, and incident response. The role combines software engineering with practical SRE methods to reduce operational toil and turn production learnings into measurable improvements.
Highlights
Hands-on senior engineering role focused on reliability, scalability, security, and operational excellence. The position offers cross-functional collaboration and significant influence over resilient cloud services and continuous service improvement.
Description
ROLE PURPOSE Improve the reliability, scalability, security and operational excellence of cloud-hosted services on Google Cloud Platform.
This is a hands-on engineering role spanning SRE practices, production Kubernetes, infrastructure automation, observability, CI/CD, incident response and continuous service improvement.
About the role
The Senior Site Reliability Engineer will work with Cloud Platform, Software Engineering, Product, Security and Service teams to design and operate resilient services.
The role will use software engineering and automation to reduce operational toil, improve service health and turn incident learning into measurable reliability improvements.
Key responsibilities
Lead reliability, availability, scalability and performance improvements across GCP-hosted applications, platforms and shared services.Define and operate service level indicators, service level objectives and error-budget practices that connect technical health to customer impact.Design actionable observability using Dynatrace, including instrumentation, dashboards, distributed tracing, service health views and SLO-based alerting.Build modular, reusable and maintainable Terraform code for secure cloud infrastructure, platform services and environment provisioning.Administer production Kubernetes environments, covering cluster lifecycle, workload deployment, capacity, upgrades, networking and platform troubleshooting.Automate operational activities and repetitive support work using Python, Groovy, Bash or PowerShell to reduce toil and improve consistency.Build and enhance CI/CD pipelines using Jenkins, Azure DevOps, GitHub Actions or equivalent tooling, with automated quality and deployment controls.Lead incident response, complex troubleshooting, root-cause analysis and post-incident reviews; ensure corrective actions are tracked to completion.Embed security, resilience, monitoring and supportability into platform and application designs from the outset.Coach engineers, contribute reusable standards and patterns, and promote SRE and operational excellence across engineering communities.