Skills
Site Reliability Engineering
DevOps
Linux administration
Networking
Kernel tuning
Docker
Kubernetes
Terraform
Infrastructure as Code
Prometheus
Grafana
Datadog
Incident response
Root cause analysis
CI/CD
Python, Go, or Bash
Linux
Ansible
GitOps
Python
Go
Bash
Istio
Summary
Help operate and evolve mission-critical infrastructure serving a large-scale user base. You will build Kubernetes platforms, implement observability and reliability standards, automate infrastructure with Terraform and configuration tools, lead incident reviews, and improve cloud efficiency and deployment safety. Candidates should have 3โ7 years of relevant experience and strong Linux, networking, container, infrastructure-as-code, and automation skills.
Highlights
Opportunity to own reliability and performance for high-scale infrastructure, collaborate closely with engineering teams, and improve resilient multi-cloud systems through automation and operational excellence.
Description
About The Role
The role owns the reliability, scalability, and performance of mission-critical production infrastructure supporting millions of daily active users.
The team works closely with software engineering squads to design resilient systems, automate operational toil, and maintain high availability standards across distributed cloud environments.
Key Responsibilities
Design, build, and maintain production Kubernetes clusters and container orchestration pipelines across multi-cloud environmentsImplement comprehensive monitoring, alerting, and observability frameworks using Prometheus, Grafana, and Datadog to track SLOs and SLAsAutomate infrastructure provisioning and configuration management using Terraform, Ansible, and GitOps workflowsLead incident response, root cause analysis, and post-mortem reviews to continuously improve system resilience and fault toleranceOptimize cloud resource utilization, security postures, and CI/CD deployment pipelines to accelerate developer velocity safely
What We Are Looking For
3โ7 years of experience in site reliability engineering, DevOps, or systems engineering in high-scale production environmentsStrong proficiency in Linux systems administration, networking fundamentals, and kernel tuningHands-on expertise with containerization (Docker, Kubernetes) and infrastructure-as-code (Terraform)Proficiency in at least one modern programming language for automation and tooling, such as Python, Go, or BashBachelor's degree in Computer Science, Engineering, or equivalent practical experienceBonus: Experience with service mesh technologies like Istio, chaos engineering practices, and multi-region cloud disaster recovery