Lead Site Reliability Engineer

Avrioctech — United Arab Emirates · Posted ~21 hours ago

Lead Onsite

Skills

Cloud infrastructure architecture Site Reliability Engineering CI/CD Jenkins Argo CD SLOs and SLIs Observability Elastic Stack Prometheus Grafana Dynatrace New Relic Kubernetes Helm Auto-healing systems Chaos engineering Chaos Mesh Litmus AWS FIS Incident management Infrastructure documentation AWS

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A lead-level reliability engineering role focused on architecting highly available cloud infrastructure and embedding operational excellence into software delivery. You will optimize CI/CD, define reliability objectives, build comprehensive observability, manage container orchestration platforms, automate recovery, and lead resilience and chaos testing. The role involves close collaboration with engineering and product teams and strong ownership of infrastructure and incident practices.

Highlights

Leadership role focused on designing scalable cloud infrastructure, improving reliability and deployment automation, building advanced observability, and driving resilience engineering. Offers broad technical ownership and collaboration across engineering and product teams.

Description

HIRING: Site Reliability Engineer - Lead | Abu Dhabi, UAE We’re looking for a Site Reliability Engineering (SRE) Lead to design, scale, and elevate our cloud infrastructure and observability ecosystem. Key Responsibilities: • Architect and deploy scalable, highly available cloud infrastructure • Lead SRE best practices to ensure reliability, performance, and scalability • Optimize CI/CD pipelines (Jenkins, Argo CD or similar) for seamless deployments • Define and track SLOs & SLIs to maintain uptime and service health • Build robust observability frameworks (Elastic Stack, Prometheus, Grafana, Dynatrace, New Relic) • Manage Kubernetes clusters and Helm charts for efficient orchestration • Implement auto-healing systems and proactive monitoring • Drive chaos engineering and resilience testing (Chaos Mesh, Litmus, AWS FIS) • Collaborate with engineering and product teams to embed reliability into development • Maintain clear infrastructure and incident documentation What We’re Looking For: • 8+ years of experience in DevOps/SRE, including leadership in enterprise environments • Hands-on experience with AWS, GCP, or Azure • Strong expertise in Infrastructure as Code (Terraform, CloudFormation, Ansible) • Proven experience in CI/CD, monitoring, and incident response • Deep knowledge of observability tools and practices • Strong Kubernetes and Helm experience at scale • Experience with databases like MySQL, Cassandra, etc. • Proficiency in Python, Bash, or Go • Experience in BCP/DR planning and capacity management • Strong communication, troubleshooting, and documentation skills