Site Reliability Engineer

Evlo Ai — United States · Posted ~9 hours ago

Mid Remote

Skills

AWS Kubernetes Terraform Infrastructure as Code Cloud networking Observability Incident response Distributed systems CI/CD Service-level objectives GitHub Actions Argo CD Jenkins

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A Site Reliability Engineer is sought to design, automate, and operate highly available production infrastructure at scale. The role focuses on Kubernetes, AWS, Terraform, cloud networking, observability, incident response, distributed systems, and automated delivery workflows, with ownership of service-level objectives and continuous reliability improvements. The position is remote.

Highlights

Remote SRE role with direct ownership of reliability, production readiness, availability, and performance. Provides hands-on work with AWS, Kubernetes, cloud networking, observability, infrastructure as code, and modern deployment automation.

Description

About The Role The Site Reliability Engineer will design, automate, and operate the infrastructure supporting production services at scale. The role focuses on Kubernetes-based platforms, cloud networking, observability, incident response, and the reliability of distributed systems running across AWS environments. Working with application engineers, security, and platform teams, the engineer will turn operational requirements into resilient systems and repeatable delivery workflows. The role is remote from New York, NY, with direct ownership of service-level objectives, production readiness, and continuous improvements to availability and performance. Key Responsibilities Design and operate highly available production infrastructure across AWS using Kubernetes, Terraform, and infrastructure-as-code best practicesBuild and maintain CI/CD pipelines with tools such as GitHub Actions, Argo CD, or Jenkins to automate testing, deployments, rollbacks, and environment provisioningDefine and enforce service-level objectives, error budgets, and production readiness standards for critical servicesImplement observability using Prometheus, Grafana, OpenTelemetry, and centralized logging platforms to reduce detection and resolution timeLead incident response, coordinate technical remediation, and produce blameless postmortems with measurable follow-up actionsImprove platform resilience through capacity planning, performance testing, disaster recovery exercises, and automated failure remediationPartner with development teams to improve application architecture, containerization, security controls, and operational runbooks What We Are Looking For 3–8 years of experience in site reliability engineering, DevOps, platform engineering, or a related production infrastructure roleStrong hands-on experience with AWS services such as EC2, EKS, IAM, VPC, S3, RDS, and CloudWatchProficiency with Kubernetes, Docker, Terraform, and Linux systems administration in production environmentsExperience building and operating CI/CD pipelines and managing deployment strategies such as canary releases, blue-green deployments, and automated rollbacksWorking knowledge of observability, distributed systems, networking, and incident management practices, including SLOs, SLIs, and error budgetsProficiency in Python, Go, or Bash for automation, tooling, and operational workflows; bachelor’s degree in computer science, engineering, or equivalent practical experienceBonus: Experience with Argo CD, Helm, Prometheus, Grafana, OpenTelemetry, service meshes, multi-region architectures, or chaos engineering