Site Reliability Engineer

Evlo Ai — United States · Posted ~8 hours ago

Senior Full-time

Skills

SRE AWS GCP Terraform Pulumi Prometheus Grafana OpenTelemetry Datadog incident response Kubernetes GitHub Actions ArgoCD cloud cost optimization database performance

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Own the reliability, scalability, and performance of production distributed systems operating at massive traffic scale across multi-region cloud environments. You will build infrastructure with Infrastructure as Code, establish comprehensive observability, lead incident response and post-mortems, automate releases, and optimize infrastructure and database performance while partnering closely with software engineering teams.

Highlights

High-scale reliability role focused on multi-region cloud systems, strong observability, automation, incident learning, cost optimization, and close collaboration with software engineering teams.

Description

About The Role The role owns the reliability, scalability, and performance of production distributed systems handling massive traffic scale across multi-region cloud environments. The team works closely with software engineering squads to embed resilience into architecture, automate operational toil, and maintain strict SLAs. Key Responsibilities Design, build, and maintain production infrastructure on AWS or GCP using Infrastructure as Code tools such as Terraform and PulumiImplement comprehensive observability stacks using Prometheus, Grafana, OpenTelemetry, and Datadog for real-time monitoring and alertingDrive incident response and conduct thorough blameless post-mortems to continuously improve system resilience and prevent recurrenceAutomate deployment pipelines and release engineering processes using GitHub Actions, ArgoCD, and KubernetesOptimize cloud infrastructure costs, resource utilization, and database performance without compromising system reliabilityEstablish and enforce security compliance, IAM policies, and disaster recovery runbooks across all production environments What We Are Looking For 3–7 years of experience in Site Reliability Engineering, DevOps, or systems engineering roles within high-growth tech environmentsDeep expertise in Kubernetes administration, containerization (Docker), and service mesh architecturesStrong proficiency in scripting languages such as Python or Go for automation and tool developmentHands-on experience with modern CI/CD pipelines, GitOps workflows, and infrastructure as code frameworksSolid understanding of networking fundamentals (TCP/IP, DNS, TLS, load balancing) and distributed systems failure modesBonus: Experience with Chaos Engineering practices, eBPF for observability, or holding professional cloud certifications (AWS/GCP)