Site Reliability Engineer

Allofresh — Indonesia · Posted ~2 hours ago

Mid Full-time

Skills

SRE Kubernetes Incident response Automation Monitoring GitOps ArgoCD SLI/SLO

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology team is hiring an SRE to improve system reliability, automate operational processes, manage cloud-native infrastructure, and support production services.

Highlights

Improve reliability of production systems, automate operations, and work with modern infrastructure technologies.

Description

Role Overview As a Site Reliability Engineer, you will be responsible for ensuring the reliability, scalability, and performance of production systems. You will apply engineering principles to operations, focusing on automation, observability, and continuous improvement to reduce manual effort and enhance system resilience. You will work closely with development teams to build and operate highly reliable services, embedding reliability into the full software lifecycle. Main Responsibilities Monitor system reliability and define, implement, and track SLIs, SLOs, and error budgets for production servicesLead incident response activities, including troubleshooting, root cause analysis, and resolutionAutomate infrastructure processes and reduce operational toil through scripting and tooling enhancementsManage and optimize Kubernetes clusters, including resource configurations, manifests, and kustomize overlaysMaintain and monitor GitOps workflows using ArgoCD to ensure consistent and reliable application deploymentsDesign, build, and maintain CI/CD pipelines using Jenkins Job DSL (Groovy)Collaborate with development teams to improve system reliability, scalability, and deployment readinessMaintain and upgrade platform tools such as Jenkins, ArgoCD, and container registriesDevelop and maintain runbooks, postmortems, and internal technical documentation Requirements 3+ years of experience in Site Reliability Engineering (SRE), DevOps, or infrastructure engineeringSolid understanding of SRE principles, including SLIs, SLOs, and error budget frameworksExperience operating and managing systems in cloud-native environments (e.g., AWS, GCP, Azure)Strong proficiency in containerization and orchestration technologies, particularly Docker and KubernetesExperience with CI/CD pipelines, automation, and deployment strategies (e.g., Jenkins, GitLab CI, GitHub Actions)Familiarity with GitOps practices and tools such as ArgoCD or FluxCDHands-on experience with observability and monitoring tools (e.g., Prometheus, Grafana, logging and alerting systems)Proficiency in scripting and automation (e.g., Bash, Python, Groovy) and Infrastructure as Code (e.g., Terraform, Ansible)Strong foundation in Linux/Unix systems and Git-based version control workflows