Site Reliability Engineer

Evlo Ai — United States · Posted ~1 hour ago

Skills

AWS Kubernetes Terraform Infrastructure as Code SRE SLOs SLIs Monitoring Incident response Root cause analysis CI/CD Infrastructure security Prometheus Grafana Datadog GitHub Actions ArgoCD

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization operating high-traffic global systems is seeking an SRE to own the reliability, scalability, and security of distributed production infrastructure. You will build AWS and Kubernetes infrastructure with Terraform, establish observability and service objectives, lead incident response and root-cause analysis, optimize resources, and scale CI/CD pipelines. The role offers significant ownership of resilient cloud systems.

Highlights

Own reliability, scalability, and security for distributed production infrastructure serving millions of users. The role combines cloud architecture, automation, observability, incident leadership, performance optimization, and CI/CD engineering.

Description

About The Role The role owns the reliability, scalability, and security of distributed infrastructure powering high-traffic production systems serving millions of users globally. You will work alongside platform architects and software engineering teams to design resilient cloud architectures, automate deployment pipelines, and minimize system downtime. Key Responsibilities Design, build, and maintain cloud infrastructure on AWS and Kubernetes using Terraform and Infrastructure as Code best practicesDefine and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and comprehensive monitoring metrics using Prometheus, Grafana, and DatadogLead incident response and root cause analysis (RCA) for production outages, implementing automated remediation to prevent recurrenceOptimize cloud resource utilization, compute performance, and infrastructure security posture across all environmentsBuild and scale CI/CD deployment pipelines using GitHub Actions or ArgoCD to support rapid, safe software releases What We Are Looking For 3–6 years of experience in Site Reliability Engineering, DevOps, or systems engineering in cloud-native environmentsStrong hands-on experience with Kubernetes, Docker, and container orchestration at scaleProficiency in infrastructure automation and configuration management using Terraform and AnsibleDeep understanding of Linux systems internals, networking fundamentals (TCP/IP, DNS, TLS), and load balancingBonus: Experience with service meshes like Istio, chaos engineering practices, and software development proficiency in Go or Python