Site Reliability Engineer

Evlo Ai — United States · Posted ~3 hours ago

Senior Full-time

Skills

Kubernetes Cloud infrastructure SRE Infrastructure as Code Monitoring SLO/SLI Terraform Helm Prometheus Grafana OpenTelemetry

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A hands-on SRE role responsible for operating scalable cloud platforms, improving reliability, automating infrastructure, and driving operational best practices.

Highlights

Advanced reliability engineering role working on large-scale distributed systems, automation, and cloud-native infrastructure.

Description

About The Role The role owns the reliability of large-scale, distributed production systems — Kubernetes-based infrastructure, service meshes, and data pipelines serving millions of requests per hour. When systems fail, this role is on point for detection, mitigation, and root cause elimination. You will work closely with product engineering teams to build paved-road tooling, automate away toil, and drive SLO-based operations across the platform. This is a hands-on engineering role, not a ticket-queue role: expect to write code, design systems, and influence architecture. Key Responsibilities Build and maintain production Kubernetes clusters (EKS/GKE) at scale, including cluster upgrades, autoscaling policies, and multi-region failoverDefine, instrument, and drive SLOs/SLIs with error budgets for critical services using Prometheus, Grafana, and OpenTelemetryDesign and implement Infrastructure as Code with Terraform and Helm, enforcing review workflows and drift detection across environmentsLead incident response for high-severity production events — triage, mitigation, blameless postmortems, and durable follow-up actionsAutomate away toil: build self-service tooling and CI/CD pipelines (GitHub Actions, ArgoCD) that reduce manual operational burden measurably each quarterHarden reliability fundamentals: capacity planning, chaos engineering exercises, disaster recovery drills, and load testing for peak traffic eventsPartner with development teams on system design reviews, latency budgets, dependency management, and graceful degradation strategies What We Are Looking For 3–6 years of experience in SRE, DevOps, or production systems engineering, with proven on-call ownership of customer-facing servicesDeep hands-on expertise with Kubernetes: operating clusters in production, debugging workloads, writing operators or controllers in Go or PythonStrong proficiency with observability tooling — Prometheus/Grafana, distributed tracing (Jaeger/Tempo), and building actionable alerting (not noise)Production experience with Terraform and at least one major cloud (AWS, GCP, or Azure), including networking (VPCs, load balancers, DNS, CDN)Solid scripting and automation skills in Python, Go, or Bash, with comfort reading and modifying application codeTrack record of running incident response and writing postmortems that lead to systemic fixes, not blameBonus: experience with service mesh (Istio/Linkerd), chaos tooling (Litmus, Gremlin), multi-cloud or on-prem hybrid environments, and formal SRE program leadership (error budget policy, capacity reviews)