Site Reliability Engineer

Evlo Ai — United States · Posted ~2 hours ago

Senior Full-time

Skills

Kubernetes AWS or GCP Prometheus Grafana OpenTelemetry Terraform Ansible CI/CD incident response root cause analysis infrastructure automation AWS GCP GitHub Actions GitLab

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A Site Reliability Engineer is sought to keep large-scale production systems reliable and resilient. The role focuses on Kubernetes infrastructure across major cloud platforms, modern observability, infrastructure automation, CI/CD, incident response, and continuous reliability improvements. You will partner with engineering and platform teams to reduce operational toil and design systems that recover gracefully from failures.

Highlights

Owns large-scale reliability and observability, with strong opportunities to automate operations, improve resilience, lead incident response, and collaborate closely with engineering and platform teams.

Description

About The Role The role keeps production systems running at scale — owning the reliability, observability, and infrastructure that powers services handling millions of requests per day. You will work closely with software engineers and platform teams to eliminate toil, automate operations, and design systems that fail gracefully instead of catastrophically. Key Responsibilities Build and maintain Kubernetes-based infrastructure across multiple cloud environments (AWS or GCP), including clusters, ingress, autoscaling policies, and workload orchestrationDesign and implement observability stacks using Prometheus, Grafana, and OpenTelemetry — dashboards, SLOs, and alerting that catch real incidents without generating noiseLead incident response for production outages: triage, mitigation, root cause analysis, and blameless postmortems with concrete follow-up actionsAutomate infrastructure provisioning and configuration with Terraform, Ansible, and CI/CD pipelines (GitHub Actions or GitLab CI) to make deployments boring and repeatableReduce toil through automation — writing tooling in Python or Go that eliminates manual operational workPartner with development teams on capacity planning, performance tuning, and reliability reviews for new services before they hit productionContribute to chaos engineering practices, failure testing, and disaster recovery runbooks for critical systems What We Are Looking For 3–7 years of experience in SRE, DevOps, or production infrastructure engineeringDeep hands-on experience with Kubernetes in production — deployment, debugging, RBAC, networking, and scalingStrong proficiency in at least one scripting or systems language: Python, Go, or BashProduction experience with infrastructure-as-code (Terraform) and modern CI/CD pipelinesSolid grasp of distributed systems fundamentals: load balancing, caching, queuing, replication, and failure modesExperience running on-call and leading incident response for customer-facing servicesBachelor's degree in Computer Science, Engineering, or equivalent practical experienceBonus: Experience with service meshes (Istio, Linkerd), chaos engineering tools (Gremlin, Litmus), cloud security posture tooling, or contributing to open-source infrastructure projects