Site Reliability Engineer

Evlo Ai — United States · Posted ~6 hours ago

Senior Full-time Remote

Skills

AWS Kubernetes Terraform Infrastructure as Code CI/CD Observability Incident response Automation Prometheus Grafana Datadog

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Own the reliability, scalability, and operational readiness of production systems running on cloud infrastructure. You will design highly available services, manage Kubernetes and infrastructure-as-code environments, build robust observability, lead incident response, and automate operational workflows while partnering closely with engineering teams.

Highlights

Fully remote role focused on reliability and scalability, with substantial ownership of cloud infrastructure, automation, observability, incident response, and deployment safety. The position offers close collaboration with software and platform engineering teams and meaningful technical impact.

Description

About The Role The Site Reliability Engineer owns the reliability, scalability, and operational readiness of production systems running across cloud infrastructure. The role focuses on Kubernetes, infrastructure as code, observability, incident response, and automation that keeps customer-facing services available and responsive at scale. You will partner with software engineers and platform teams to eliminate recurring operational work, improve deployment safety, and build systems that are easier to operate. This role is remote, with preference for candidates based in or near Atlanta, GA, and requires participation in an on-call rotation. Key Responsibilities Design and operate highly available services and infrastructure on AWS, using Kubernetes, Terraform, and automated deployment pipelinesBuild and maintain observability systems with Prometheus, Grafana, Datadog, or equivalent tools, including service-level indicators, dashboards, and actionable alertsLead incident response for production outages, coordinate mitigation, document root-cause analyses, and drive corrective actions through completionAutomate provisioning, scaling, release processes, and routine operational tasks with Python, Go, Bash, or similar languagesDefine and improve service-level objectives, error budgets, capacity plans, and disaster recovery procedures for critical systemsPartner with development teams to improve application performance, deployment safety, resilience testing, and production readinessReview infrastructure and architecture changes for security, reliability, cost efficiency, and operational risk What We Are Looking For 3–8 years of experience in site reliability engineering, DevOps, platform engineering, or production infrastructure rolesHands-on experience operating Kubernetes and containerized workloads in AWS, GCP, or Azure production environmentsStrong proficiency with infrastructure as code, preferably Terraform, and experience managing CI/CD systems such as GitHub Actions, GitLab CI, or Argo CDPractical expertise in observability across metrics, logs, and distributed traces using tools such as Prometheus, Grafana, Datadog, OpenTelemetry, or ELKProficiency in Python, Go, Bash, or another programming language used to build reliable automation and operational toolingBachelor’s degree in computer science, engineering, or a related technical field, or equivalent professional experienceBonus: Experience with service mesh technologies, Kafka or other distributed systems, chaos engineering, PostgreSQL reliability, and zero-downtime migrations