Summary
✨ AI‑Generated
A DevOps/MLOps engineering role dedicated to infrastructure and deployment pipelines for AI, machine learning, LLM, and microservices workloads. You will manage GPU-enabled compute, Linux environments, cloud and on-premises infrastructure, containerized applications, and automated deployment workflows while improving reliability, scalability, and performance. The position is offered on a renewable one-year contract.
Highlights
A renewable one-year contract focused on modern AI and machine learning infrastructure. The role offers hands-on ownership of GPU-enabled environments, cloud and on-premises systems, containerized deployments, automation, scalability, and reliability for AI workloads.
Description
Job Code: 6315
Job Title: DevOps / MLOps Engineer (AI & LLM Platforms)
Location: Abu Dhabi
Contract: 1 year and renewable
Experience: 7+ years
Role Purpose
The DevOps / MLOps Engineer is responsible for the setup, automation, and maintenance of infrastructure and deployment pipelines for AI/ML and microservices-based applications.
This role focuses on enabling efficient development, testing, and deployment of AI solutions, including LLM workloads, while ensuring system reliability, scalability, and performance.
Key Responsibilities
Infrastructure Support & Environment ManagementSet up and maintain compute infrastructure, including GPU-enabled environments.Configure and manage Linux-based systems for development and production environments.Provisioning and configuration of cloud and on-prem infrastructure.Monitor system resources and assist in performance tuning and optimization.
Containerization & DeploymentBuild and manage containerized applications using Docker.Deploy and manage applications on Kubernetes clusters under guidance from senior engineers.Creating deployment configurations, Helm charts, and environment setups.Support scaling and orchestration of microservices and AI workloads.
CI/CD Pipeline ImplementationDevelop and maintain CI/CD pipelines for application and AI model deployment.Automate build, test, and deployment processes using tools like Azure DevOps, GitHub Actions, or Jenkins.Ensure smooth promotion of code and models across environments (dev, test, prod).Troubleshoot pipeline failures and deployment issues.
MLOps & AI Deployment SupportDeploying machine learning models and LLM-based services.Integration of AI components into production systems.Contribute to model versioning, monitoring, and lifecycle management.Work with AI engineers to operationalize RAG pipelines and inference services.
Monitoring, Logging & Issue ResolutionImplement and maintain monitoring and logging solutions (e.g., Prometheus, Grafana, ELK).Track application performance, system health, and availability.Respond to incidents, troubleshoot issues, and escalate when required.Assist in root cause analysis and continuous improvement.
Automation & ScriptingWrite scripts (Python, Bash) to automate repetitive operational tasks.Support Infrastructure as Code (IaC) initiatives using tools like Terraform or ARM templates.Improve operational efficiency through automation and tooling.
Collaboration & SupportWork closely with Senior DevOps/MLOps Engineers, AI Engineers, and Development teams.Support developers in environment setup, debugging, and deployment processes.Follow DevOps and MLOps best practices and continuously improve operational workflows.
Required Skills & Qualifications
Bachelor s degree in Computer Science, Engineering, or related field.7+ years of experience in DevOps or platform engineering roles.Basic to intermediate experience with Linux system administration.Hands-on experience with Docker and containerization.Familiarity with Kubernetes (deployment and basic management).Experience with CI/CD tools (Azure DevOps, GitHub Actions, Jenkins, etc.).Basic understanding of cloud platforms (Azure, AWS, or GCP).Scripting skills in Python, Bash, or similar.Understanding of version control systems (Git).
Preferred Skills
Exposure to AI/ML model deployment and MLOps practices.Familiarity with LLM deployment concepts and tools.Basic knowledge of GPU environments and high-performance computing.Experience with monitoring and logging tools (Prometheus, Grafana, ELK).Knowledge of Infrastructure as Code (Terraform, ARM templates).Understanding of microservices architecture.
Key Performance Indicators (KPIs)
Deployment success rate and pipeline stability.System uptime and availability.Resolution time for incidents and issues.Efficiency of CI/CD processes.Infrastructure utilization and basic cost optimization.Support effectiveness for development and AI teams.
Stakeholders & Reporting
Reports to: Senior DevOps / MLOps Engineer / Platform LeadKey Stakeholders:AI Engineers & Data ScientistsBackend & Frontend DevelopersDevOps / Platform TeamQA & Release Management Teams