Summary
✨ AI‑Generated
A Site Reliability Engineer is sought to own the reliability, scalability, and observability of production infrastructure serving millions of users. The role focuses on Kubernetes and cloud infrastructure, infrastructure as code, SLO/SLI monitoring, automated incident response, CI/CD, and improving distributed system resilience.
Highlights
Own reliability and scalability for infrastructure serving millions of users, work with modern distributed systems, automate operations, and directly influence system availability, performance, and resilience.
Description
About The Role
The role owns the reliability, scalability, and observability of core production infrastructure supporting millions of users daily.
The team works alongside backend and platform engineers to build robust distributed systems where latency, high availability, and automated recovery are critical.
Key Responsibilities
Design, build, and maintain production Kubernetes clusters and underlying cloud infrastructure using TerraformDefine and track critical service level objectives (SLOs), service level indicators (SLIs), and comprehensive alerting pipelinesAutomate incident response, failover mechanisms, and routine operational tasks using Python or GoConduct root cause analysis for production incidents and implement preventative architectural improvementsManage continuous integration and continuous deployment (CI/CD) pipelines to ensure safe, zero-downtime releasesCollaborate with engineering teams to review system architectures for scalability, security, and performance bottlenecks
What We Are Looking For
3–6 years of experience in Site Reliability Engineering, DevOps, or systems engineering in cloud-native environmentsStrong hands-on experience with Kubernetes, Docker, and container orchestration at production scaleDeep proficiency with Infrastructure as Code tooling, specifically Terraform and AnsibleSolid understanding of networking concepts, TCP/IP, DNS, load balancing, and secure cloud architecturesExperience with observability platforms such as Prometheus, Grafana, Datadog, or OpenTelemetryBonus: Contributions to open-source infrastructure projects or certifications such as CKA or AWS Certified DevOps Engineer