Summary
✨ AI‑Generated
A Site Reliability Engineer role responsible for building highly reliable, scalable, and performant production systems. You will automate infrastructure, reduce operational toil, manage Kubernetes environments, operate GitOps workflows, define reliability objectives, and lead incident troubleshooting and root-cause analysis.
Highlights
Focus on automation, observability, scalability, and continuous improvement while working closely with development teams. The role provides hands-on ownership of production reliability, incident response, infrastructure optimization, and modern GitOps practices.
Description
Role Overview
As a Site Reliability Engineer, you will be responsible for ensuring the reliability, scalability, and performance of production systems.
You will apply engineering principles to operations, focusing on automation, observability, and continuous improvement to reduce manual effort and enhance system resilience.
You will work closely with development teams to build and operate highly reliable services, embedding reliability into the full software lifecycle.
Main Responsibilities
Monitor system reliability and define, implement, and track SLIs, SLOs, and error budgets for production servicesLead incident response activities, including troubleshooting, root cause analysis, and resolutionAutomate infrastructure processes and reduce operational toil through scripting and tooling enhancementsManage and optimize Kubernetes clusters, including resource configurations, manifests, and kustomize overlaysMaintain and monitor GitOps workflows using ArgoCD to ensure consistent and reliable application deploymentsDesign, build, and maintain CI/CD pipelines using Jenkins Job DSL (Groovy)Collaborate with development teams to improve system reliability, scalability, and deployment readinessMaintain and upgrade platform tools such as Jenkins, ArgoCD, and container registriesDevelop and maintain runbooks, postmortems, and internal technical documentation
Requirements
3+ years of experience in Site Reliability Engineering (SRE), DevOps, or infrastructure engineeringSolid understanding of SRE principles, including SLIs, SLOs, and error budget frameworksExperience operating and managing systems in cloud-native environments (e.g., AWS, GCP, Azure)Strong proficiency in containerization and orchestration technologies, particularly Docker and KubernetesExperience with CI/CD pipelines, automation, and deployment strategies (e.g., Jenkins, GitLab CI, GitHub Actions)Familiarity with GitOps practices and tools such as ArgoCD or FluxCDHands-on experience with observability and monitoring tools (e.g., Prometheus, Grafana, logging and alerting systems)Proficiency in scripting and automation (e.g., Bash, Python, Groovy) and Infrastructure as Code (e.g., Terraform, Ansible)Strong foundation in Linux/Unix systems and Git-based version control workflows