AI DevOps Engineer

Gts Consulting999 — Singapore · Posted ~1 week ago

Senior

Skills

DevOps AI/ML Infrastructure CI/CD Cloud Infrastructure Kubernetes Monitoring Incident Response Infrastructure Automation Security Optimization AWS Azure Cloud GPU Observability

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Lead the production deployment and operation of AI platforms across cloud GPU infrastructure, managed services, and self-hosted environments. You will build CI/CD pipelines, operate Kubernetes-based systems, monitor performance, respond to incidents, maintain SLAs, and optimize reliability, security, and cost across multi-region deployments.

Highlights

Own production deployment and operations for AI infrastructure across cloud GPU, managed, and self-hosted environments. The role combines advanced DevOps, Kubernetes, CI/CD, observability, reliability, security, and cost optimization across large-scale multi-environment systems.

Description

Job Responsibilities Own the deployment of the AI platform into production across cloud GPU infrastructure, managed services, and future on-premise or self-hosted environments, maintaining stability, scalability, and the flexibility to adapt as business needs evolve. Build and maintain CI/CD pipelines, monitor system performance, respond to incidents, uphold SLAs, and drive cost and security optimization across all environments. Skilled in cloud platforms (e.g., AWS, Azure) and containerization (e.g., Kubernetes), with the ability to deploy across both cloud and on-premise setups and select the best-fit solution based on business needs. Proven experience operating AI platforms or large-scale systems, including support for multi-region and multi-environment deployments. Job Requirements Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience. Proven experience in DevOps engineering, with a focus on AI/ML workflows and infrastructure. Strong proficiency in cloud platforms such as AWS, Azure, or Google Cloud, including experience with cloud-native services. Hands-on experience with containerization and orchestration tools like Docker and Kubernetes. Proficiency in scripting and programming languages such as Python, Bash, or Go. Experience with CI/CD tools such as Jenkins, GitLab CI/CD, or CircleCI. Familiarity with AI/ML frameworks and tools such as TensorFlow, PyTorch, or Scikit-learn. Strong understanding of networking, security, and system administration principles. Excellent problem-solving skills and the ability to work collaboratively in a fast-paced environment. Strong communication skills to effectively collaborate with cross-functional teams.