Site Reliability Engineer

Superbot Bot — United Arab Emirates · Posted ~4 hours ago

Senior Full-time

Skills

AWS Terraform Kubernetes Infrastructure as Code Prometheus Grafana SLI/SLO/SLA Incident Management

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A reliability-focused engineering role responsible for building resilient cloud infrastructure, improving system observability, managing container platforms, and driving operational excellence through automation and best practices.

Highlights

Opportunity to influence large-scale infrastructure reliability, automation practices, observability, and cloud operations while working on scalable systems.

Description

About The Role You’ll sit at the intersection of software engineering and infrastructure, partnering with product and platform teams to keep our systems observable, scalable, and resilient. From shaping our on-call culture to driving infrastructure-as-code adoption, you’ll have real influence over how BOT builds and operates software at scale — for 500+ people and growing. Key Responsibilities Own reliability across critical services — define and track SLIs/SLOs/SLAs, lead blameless post-mortems, and drive down MTTR through systematic incident management.Design, build, and maintain cloud infrastructure on AWS (primary) using Terraform, ensuring environments are reproducible, version-controlled, and auditable.Scale and optimize our Kubernetes-based container platform — capacity planning, resource tuning, autoscaling, and cluster lifecycle management.Strengthen observability end-to-end: instrument services with Prometheus and Grafana, build actionable alerting, and reduce alert noise through continuous tuning.Accelerate CI/CD pipelines (GitHub Actions / GitLab CI) to support frequent, safe deployments — shift reliability left by embedding checks into the delivery workflow.Partner with engineering teams as an internal reliability advisor — run game days, advocate for SRE best practices, and help developers build services that operate well from the start. Job Requirements 4+ years in an SRE, DevOps, or platform engineering role with production ownership at scale.Hands-on Kubernetes experience — deployment, scaling, networking, and troubleshooting in a production environment.Infrastructure-as-code fluency with Terraform (or Pulumi) across a major cloud provider (AWS preferred).Observability stack experience — Prometheus, Grafana, and/or equivalent tools with a track record of building meaningful dashboards and alerts.Proficiency in at least one scripting or programming language (Python, Go, or Bash) for automation and tooling.AWS certification (Solutions Architect, DevOps Engineer, or SysOps Administrator). What We Offer Competitive Compensation: Enjoy a salary package tailored to your skills and experience, along with performance-based bonuses.Comprehensive Benefits: We support your well-being with accommodation, meal allowances, and assistance with work visa processing.Work-Life Balance: Unwind with generous holiday and New Year bonuses.Top-Tier Equipment: Stay productive with the latest tools, including a MacBook and iPhone.Thriving Culture: Immerse yourself in a dynamic, inclusive work environment that fosters growth.