Site Reliability Engineer

Evlo Ai — United States · Posted ~22 hours ago

Senior

Skills

Site reliability engineering Kubernetes Terraform Cloud infrastructure SLOs SLIs Observability Incident response CI/CD Python Go Distributed systems Cloud SLO SLI

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A Site Reliability Engineer is sought to own the reliability, scalability, and observability of production infrastructure serving millions of users. The role focuses on Kubernetes and cloud infrastructure, infrastructure as code, SLO/SLI monitoring, automated incident response, CI/CD, and improving distributed system resilience.

Highlights

Own reliability and scalability for infrastructure serving millions of users, work with modern distributed systems, automate operations, and directly influence system availability, performance, and resilience.

Description

About The Role The role owns the reliability, scalability, and observability of core production infrastructure supporting millions of users daily. The team works alongside backend and platform engineers to build robust distributed systems where latency, high availability, and automated recovery are critical. Key Responsibilities Design, build, and maintain production Kubernetes clusters and underlying cloud infrastructure using TerraformDefine and track critical service level objectives (SLOs), service level indicators (SLIs), and comprehensive alerting pipelinesAutomate incident response, failover mechanisms, and routine operational tasks using Python or GoConduct root cause analysis for production incidents and implement preventative architectural improvementsManage continuous integration and continuous deployment (CI/CD) pipelines to ensure safe, zero-downtime releasesCollaborate with engineering teams to review system architectures for scalability, security, and performance bottlenecks What We Are Looking For 3–6 years of experience in Site Reliability Engineering, DevOps, or systems engineering in cloud-native environmentsStrong hands-on experience with Kubernetes, Docker, and container orchestration at production scaleDeep proficiency with Infrastructure as Code tooling, specifically Terraform and AnsibleSolid understanding of networking concepts, TCP/IP, DNS, load balancing, and secure cloud architecturesExperience with observability platforms such as Prometheus, Grafana, Datadog, or OpenTelemetryBonus: Contributions to open-source infrastructure projects or certifications such as CKA or AWS Certified DevOps Engineer