Summary
✨ AI‑Generated
Own the reliability, scalability, and performance of production distributed systems operating at massive traffic scale across multi-region cloud environments. You will build infrastructure with Infrastructure as Code, establish comprehensive observability, lead incident response and post-mortems, automate releases, and optimize infrastructure and database performance while partnering closely with software engineering teams.
Highlights
High-scale reliability role focused on multi-region cloud systems, strong observability, automation, incident learning, cost optimization, and close collaboration with software engineering teams.
Description
About The Role
The role owns the reliability, scalability, and performance of production distributed systems handling massive traffic scale across multi-region cloud environments.
The team works closely with software engineering squads to embed resilience into architecture, automate operational toil, and maintain strict SLAs.
Key Responsibilities
Design, build, and maintain production infrastructure on AWS or GCP using Infrastructure as Code tools such as Terraform and PulumiImplement comprehensive observability stacks using Prometheus, Grafana, OpenTelemetry, and Datadog for real-time monitoring and alertingDrive incident response and conduct thorough blameless post-mortems to continuously improve system resilience and prevent recurrenceAutomate deployment pipelines and release engineering processes using GitHub Actions, ArgoCD, and KubernetesOptimize cloud infrastructure costs, resource utilization, and database performance without compromising system reliabilityEstablish and enforce security compliance, IAM policies, and disaster recovery runbooks across all production environments
What We Are Looking For
3–7 years of experience in Site Reliability Engineering, DevOps, or systems engineering roles within high-growth tech environmentsDeep expertise in Kubernetes administration, containerization (Docker), and service mesh architecturesStrong proficiency in scripting languages such as Python or Go for automation and tool developmentHands-on experience with modern CI/CD pipelines, GitOps workflows, and infrastructure as code frameworksSolid understanding of networking fundamentals (TCP/IP, DNS, TLS, load balancing) and distributed systems failure modesBonus: Experience with Chaos Engineering practices, eBPF for observability, or holding professional cloud certifications (AWS/GCP)