Summary
✨ AI‑Generated
A Site Reliability Engineer will design and operate infrastructure for highly available production services across multiple environments. The role covers cloud and container platforms, infrastructure as code, observability, capacity planning, incident response, and reliability standards, with a strong focus on sustainable automation.
Highlights
Reliability-focused engineering role with ownership of highly available cloud infrastructure, incident response, observability, capacity planning, and safe delivery. Strong emphasis on turning recurring operational issues into durable automation and measurable reliability improvements.
Description
About The Role
The Site Reliability Engineer will design, operate, and improve the infrastructure behind highly available production services running across AWS, Kubernetes, and Linux-based systems.
The role focuses on reliability engineering, incident response, observability, capacity planning, and safe delivery at scale.
This role will partner with application engineers and platform teams to reduce operational risk and improve developer velocity.
The team needs an engineer who can turn recurring production issues into durable automation, clear service-level objectives, and measurable improvements in availability and performance.
Key Responsibilities
Build and operate production infrastructure using AWS, Kubernetes, Terraform, and Helm across multiple environmentsDefine and maintain service-level objectives, error budgets, runbooks, and escalation procedures for critical servicesImprove observability through Prometheus, Grafana, OpenTelemetry, and centralized logging platforms such as Datadog or ElasticsearchLead incident response, coordinate technical investigations, and deliver blameless postmortems with clearly tracked remediation itemsAutomate provisioning, deployment, and operational workflows using Python, Go, or Bash and integrate them into CI/CD pipelinesAnalyze system performance, capacity, and failure patterns to eliminate bottlenecks and improve availability, latency, and recovery timePartner with security and engineering teams to implement least-privilege access, secrets management, patching, and resilient disaster recovery practices
What We Are Looking For
3–8 years of experience in site reliability engineering, DevOps, platform engineering, or production infrastructure operationsStrong hands-on experience with AWS services, Linux systems, Kubernetes, Docker, Terraform, and infrastructure-as-code practicesProficiency in Python, Go, or a similar programming language for automation, tooling, and systems integrationExperience operating CI/CD systems such as GitHub Actions, GitLab CI, Jenkins, or Argo CD in production environmentsSolid understanding of distributed systems, networking, databases, observability, incident management, and reliability metricsBachelor’s degree in computer science, engineering, or a related technical field, or equivalent professional experienceBonus: Experience with service meshes, multi-region architectures, Kafka, PostgreSQL, chaos engineering, or formal SRE practices