Site Reliability Engineer

Evlo Ai โ€” United States ยท Posted ~1 day ago

Mid

Skills

Site Reliability Engineering DevOps Linux administration Networking Kernel tuning Docker Kubernetes Terraform Infrastructure as Code Prometheus Grafana Datadog Incident response Root cause analysis CI/CD Python, Go, or Bash Linux Ansible GitOps Python Go Bash Istio

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

Help operate and evolve mission-critical infrastructure serving a large-scale user base. You will build Kubernetes platforms, implement observability and reliability standards, automate infrastructure with Terraform and configuration tools, lead incident reviews, and improve cloud efficiency and deployment safety. Candidates should have 3โ€“7 years of relevant experience and strong Linux, networking, container, infrastructure-as-code, and automation skills.

Highlights

Opportunity to own reliability and performance for high-scale infrastructure, collaborate closely with engineering teams, and improve resilient multi-cloud systems through automation and operational excellence.

Description

About The Role The role owns the reliability, scalability, and performance of mission-critical production infrastructure supporting millions of daily active users. The team works closely with software engineering squads to design resilient systems, automate operational toil, and maintain high availability standards across distributed cloud environments. Key Responsibilities Design, build, and maintain production Kubernetes clusters and container orchestration pipelines across multi-cloud environmentsImplement comprehensive monitoring, alerting, and observability frameworks using Prometheus, Grafana, and Datadog to track SLOs and SLAsAutomate infrastructure provisioning and configuration management using Terraform, Ansible, and GitOps workflowsLead incident response, root cause analysis, and post-mortem reviews to continuously improve system resilience and fault toleranceOptimize cloud resource utilization, security postures, and CI/CD deployment pipelines to accelerate developer velocity safely What We Are Looking For 3โ€“7 years of experience in site reliability engineering, DevOps, or systems engineering in high-scale production environmentsStrong proficiency in Linux systems administration, networking fundamentals, and kernel tuningHands-on expertise with containerization (Docker, Kubernetes) and infrastructure-as-code (Terraform)Proficiency in at least one modern programming language for automation and tooling, such as Python, Go, or BashBachelor's degree in Computer Science, Engineering, or equivalent practical experienceBonus: Experience with service mesh technologies like Istio, chaos engineering practices, and multi-region cloud disaster recovery