AI Systems Resiliency Engineers
Damacdigital — United Arab Emirates · Posted ~4 weeks ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
DAMAC AI runs GPU cluster infrastructure at the edge of what's commercially deployed — B300 and GB300 Blackwell-based DGX SuperPOD and HGX systems.
We need the engineer who keeps it running when things break, and makes sure it doesn't break the same way twice.
The Role
You will own reliability and resilience for mission-critical AI compute infrastructure — SLA management, disaster recovery, incident response, and chaos engineering.
This is SRE work with deep GPU infrastructure specificity, not generic cloud ops.
What You Will Do
Own SLA and uptime commitments for DGX SuperPOD and HGX GPU clustersDesign and run disaster recovery and business continuity plansLead incident management and root cause analysis for GPU infrastructure failuresDesign and execute chaos engineering and resilience testingOwn capacity planning and performance managementBuild observability and automation for proactive failure detection
What You Bring
Strong SRE background — GPU or HPC infrastructure exposure idealHands-on with GPU cluster failure modes — NVIDIA DGX/HGX or equivalentProven SLA and uptime ownership at scaleDirect disaster recovery and chaos engineering execution experienceObservability tooling and automation skills — Prometheus, Grafana, Python, or equivalentComfortable in a 24/7 mission-critical environment
Based in Dubai or Noida.
Immediate requirement.
We have 85,785 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 85,785 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume