Summary
A site reliability engineering role focused on maintaining scalable production environments, improving observability, automating operations, and supporting cloud-native platforms.
Highlights
Opportunity to improve reliability of high-availability digital platforms through automation, cloud operations, and modern monitoring practices.
Description
Location: Abu Dhabi, UAE
Experience: 5+ Years
Role Overview
We are seeking Site Reliability Engineer (SRE) to ensure the reliability, availability, and scalability of enterprise digital banking platforms.
The role focuses on observability, automation, incident management, cloud-native operations, and continuous improvement of high-availability production environments.
Key Responsibilities
Reliability Engineering
Define and implement SLIs, SLOs, and error budgets.Build observability solutions using Dynatrace, Prometheus, Grafana, and ELK.Leverage AIOps for proactive monitoring and anomaly detection.Lead incident response, root cause analysis, and post-incident reviews.
Platform Engineering
Support Kubernetes-based microservices and cloud-native platforms.Automate operational workflows, remediation, and deployment processes.Improve deployment strategies using canary and blue-green releases.Integrate reliability controls into CI/CD pipelines.
Cloud Operations
Optimize AWS-based production environments.Support Kafka, Redis, RabbitMQ, Aurora, and RDS services.Collaborate with DevOps, Cloud, Product, and Engineering teams.Promote operational excellence and continuous reliability improvements.
Required Skills
Site Reliability EngineeringKubernetesTerraformAWSDynatracePrometheus & GrafanaELK StackPython / Bash / Go
Preferred Skills
AIOps (Davis AI)Kafka & RabbitMQAurora & RDSCI/CDBanking or FinTech Experience
๐ฉ Apply Now: careers@d4insight.com