DevOps Engineer

Quantum Talent Group — United Arab Emirates · Posted ~22 hours ago

Contract Onsite

Skills

DevOps Terraform Ansible CloudFormation cloud infrastructure on-premises infrastructure Prometheus Grafana ELK Stack real-time data streaming queuing systems

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A 12-month onsite DevOps position focused on designing and maintaining scalable infrastructure for AI workloads. You will automate provisioning, optimize cloud and on-premises environments, manage real-time data systems, and implement robust monitoring, logging, and alerting for highly available production platforms.

Highlights

A 12-month onsite opportunity working on scalable infrastructure for AI model deployment, with broad exposure to cloud, on-premises systems, automation, monitoring, streaming, and production reliability.

Description

A data analytics company based in Abu Dhabi is looking for a DevOps Engineer to join them on a 12-month contract. The role requires onsite work in Abu Dhabi, and preference will be given to candidates who are currently based in the UAE and available locally. Key Responsibilities: Design, build, and maintain scalable infrastructure to support AI model deployment and end-to-end lifecycle management.Automate infrastructure provisioning, configuration, and management using tools such as Terraform, Ansible, and CloudFormation.Optimize cloud and on-premises infrastructure to improve scalability, performance, and cost efficiency.Manage and optimize queuing systems and real-time data streaming architectures for reliable performance.Monitor, maintain, and troubleshoot production environments to ensure high availability and optimal system performance.Implement robust logging, monitoring, and alerting solutions using tools such as Prometheus, Grafana, and the ELK Stack.Establish comprehensive monitoring for infrastructure health, system metrics, and machine learning model performance.Perform root cause analysis and conduct post-incident reviews to identify issues and continuously improve system reliability. Qualifications: At least 5-8 years of related experience Proficiency in at least one scripting language: Python, Bash, or Go.Hands-on experience with cloud platforms like AWS, Azure, or Google Cloud.Skilled in containerization and orchestration with Docker and Kubernetes.Experience using CI/CD tools such as Azure DevOps, Jenkins, GitLab CI/CD, or CircleCI.Knowledge of monitoring and observability tools like Prometheus, Datadog, New Relic, Grafana, or PagerDuty.Familiarity with real-time streaming architectures for AI and data applications