Site Reliability Engineer

Evlo Ai — United States · Posted ~1 hour ago

Mid Full-time

Skills

Site Reliability Engineering AWS Kubernetes Linux Terraform Helm CI/CD Infrastructure automation Observability Incident response Service-level objectives Error budgets Operational readiness Distributed systems GitHub Actions

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Own the reliability of distributed production services as a Site Reliability Engineer. You will operate cloud-native infrastructure, automate provisioning and deployments, establish measurable reliability standards, and improve observability and incident response as systems scale.

Highlights

High-impact reliability role focused on highly available production systems, automation, observability, incident response, and scalable cloud infrastructure. The position works closely with platform, application, and security teams to improve reliability and delivery velocity.

Description

About The Role The Site Reliability Engineer owns the systems and practices that keep production services available, performant, and recoverable at scale. The role spans Kubernetes-based workloads, cloud infrastructure, deployment automation, observability, and incident response across a distributed production environment. This engineer will work with platform, application, and security teams to reduce operational risk through automation and sound architecture. The work directly affects uptime, deployment velocity, customer experience, and the team's ability to operate reliable services as traffic and system complexity grow. Key Responsibilities Design, operate, and improve highly available services running on AWS, Kubernetes, and LinuxAutomate infrastructure provisioning and configuration management with Terraform, Helm, and GitHub Actions or equivalent CI/CD toolingDefine and maintain service-level objectives, error budgets, and operational readiness standards for production systemsBuild observability using Prometheus, Grafana, OpenTelemetry, and centralized logging to identify latency, capacity, and reliability issuesLead incident response, including mitigation, stakeholder communication, root-cause analysis, and durable follow-up actionsEngineer safe deployment and recovery workflows, including progressive delivery, automated rollback, backup validation, and disaster recovery testingPartner with software engineers to eliminate recurring toil, improve system performance, and strengthen production design reviews What We Are Looking For 3–8 years of experience in site reliability engineering, DevOps, platform engineering, or a closely related infrastructure roleHands-on experience operating production workloads on AWS or another major cloud provider, including networking, IAM, compute, storage, and managed databasesStrong Kubernetes and Linux administration skills, with practical experience debugging containers, networking, resource constraints, and distributed systemsProficiency with infrastructure as code and automation using Terraform plus Python, Go, Bash, or a comparable programming languageExperience building production observability with metrics, logs, traces, dashboards, and actionable alerting; familiarity with SLO and error-budget practicesDemonstrated experience participating in an on-call rotation and managing high-severity incidents in a blameless, process-oriented environmentBachelor’s degree in computer science, engineering, or a related technical field, or equivalent professional experience; Bonus: experience with service mesh technologies, Argo CD, Kafka, PostgreSQL, compliance-driven environments, or multi-region disaster recovery