Site Reliability Engineer

Evlo Ai — United States · Posted ~12 hours ago

Senior Full-time

Skills

Terraform Kubernetes cloud infrastructure Prometheus Grafana observability distributed systems incident response SLOs SLIs cloud platforms

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology-focused organization is seeking a reliability engineer to build and operate scalable cloud infrastructure. The role focuses on automation, infrastructure as code, monitoring, incident management, and improving the resilience of critical distributed systems.

Highlights

Opportunity to own large-scale cloud reliability initiatives, automate infrastructure operations, improve observability, and work on high-availability production systems.

Description

About The Role The role owns the reliability, scalability, and performance of distributed cloud infrastructure supporting core production workloads at scale. The team works closely with platform and software engineering squads to automate infrastructure provisioning, eliminate toil, and maintain high availability standards. Key Responsibilities Architect, deploy, and manage production infrastructure using Terraform and Kubernetes across multi-region cloud environmentsDesign and implement comprehensive observability pipelines using Prometheus, Grafana, and distributed tracing tools to monitor system health and latencyDefine and maintain service level objectives (SLOs), service level indicators (SLIs), and error budgets for all critical production servicesAutomate incident response, root cause analysis, and remediation workflows to minimize downtime and prevent recurring outagesConduct thorough post-mortems and implement systematic preventive measures across the infrastructure stackParticipate in an on-call rotation to handle production incidents and ensure rapid recovery from system anomalies What We Are Looking For 3–6 years of experience in site reliability engineering, DevOps, or systems engineering in large-scale production environmentsDeep expertise in Kubernetes administration, containerization, and cloud-native networking principlesStrong proficiency in infrastructure-as-code tooling, specifically Terraform and AnsibleHands-on experience with modern observability stacks (Prometheus, Grafana, Datadog) and log aggregation platforms (ELK, Loki)Proficiency in at least one scripting or programming language such as Python, Go, or Bash for automation tasksBonus: Experience with service mesh technologies like Istio, chaos engineering practices, or holding AWS/GCP professional certifications