Summary
✨ AI‑Generated
A Site Reliability Engineer is sought to keep large-scale production systems reliable and resilient. The role focuses on Kubernetes infrastructure across major cloud platforms, modern observability, infrastructure automation, CI/CD, incident response, and continuous reliability improvements. You will partner with engineering and platform teams to reduce operational toil and design systems that recover gracefully from failures.
Highlights
Owns large-scale reliability and observability, with strong opportunities to automate operations, improve resilience, lead incident response, and collaborate closely with engineering and platform teams.
Description
About The Role
The role keeps production systems running at scale — owning the reliability, observability, and infrastructure that powers services handling millions of requests per day.
You will work closely with software engineers and platform teams to eliminate toil, automate operations, and design systems that fail gracefully instead of catastrophically.
Key Responsibilities
Build and maintain Kubernetes-based infrastructure across multiple cloud environments (AWS or GCP), including clusters, ingress, autoscaling policies, and workload orchestrationDesign and implement observability stacks using Prometheus, Grafana, and OpenTelemetry — dashboards, SLOs, and alerting that catch real incidents without generating noiseLead incident response for production outages: triage, mitigation, root cause analysis, and blameless postmortems with concrete follow-up actionsAutomate infrastructure provisioning and configuration with Terraform, Ansible, and CI/CD pipelines (GitHub Actions or GitLab CI) to make deployments boring and repeatableReduce toil through automation — writing tooling in Python or Go that eliminates manual operational workPartner with development teams on capacity planning, performance tuning, and reliability reviews for new services before they hit productionContribute to chaos engineering practices, failure testing, and disaster recovery runbooks for critical systems
What We Are Looking For
3–7 years of experience in SRE, DevOps, or production infrastructure engineeringDeep hands-on experience with Kubernetes in production — deployment, debugging, RBAC, networking, and scalingStrong proficiency in at least one scripting or systems language: Python, Go, or BashProduction experience with infrastructure-as-code (Terraform) and modern CI/CD pipelinesSolid grasp of distributed systems fundamentals: load balancing, caching, queuing, replication, and failure modesExperience running on-call and leading incident response for customer-facing servicesBachelor's degree in Computer Science, Engineering, or equivalent practical experienceBonus: Experience with service meshes (Istio, Linkerd), chaos engineering tools (Gremlin, Litmus), cloud security posture tooling, or contributing to open-source infrastructure projects