Summary
✨ AI‑Generated
A senior infrastructure engineering role focused on building and operating large-scale machine learning platforms using cloud-native technologies and automation practices.
Highlights
Remote opportunity building scalable ML infrastructure platforms with focus on automation, reliability, and cloud technologies.
Description
DevOps Engineer – ML Infrastructure
Location: Canada – Remote
Job Description:
We are seeking a highly skilled Senior DevOps Engineer – ML Infrastructure to build, operate, and evolve large-scale, business-critical infrastructure platforms supporting high-volume production workloads.
The ideal candidate will have strong expertise in Kubernetes, cloud-native infrastructure, distributed systems, MLOps, AWS EKS, and GPU-based environments.
This role will focus on reliability, scalability, automation, infrastructure modernization, and operational excellence across machine learning platforms.
You will collaborate closely with Data Scientists, ML Engineers, Software Engineers, Platform Teams, and Infrastructure Specialists to productionize ML solutions and build scalable platforms for AI/ML workloads.
Key Responsibilities:
Design, build, and maintain scalable MLOps platforms for ML model training, deployment, and monitoring.Develop and manage cloud-native infrastructure supporting large-scale ML workloads on Kubernetes.Implement and operate Kubernetes batch scheduling solutions such as Volcano and Kueue.Optimize utilization of GPU and compute resources across ML workloads.Build and maintain CI/CD pipelines for ML services, infrastructure, and platform components.Automate infrastructure provisioning and lifecycle management using Infrastructure-as-Code (IaC).Manage and optimize Kubernetes environments, including production workloads running on AWS EKS.Support infrastructure modernization, cloud migration, and platform migration initiatives.Partner with Data Scientists, ML Engineers, and Software Engineers to productionize machine learning solutions.Implement robust monitoring, observability, logging, and alerting for distributed systems and GPU clusters.Ensure platform reliability, security, scalability, and operational excellence.Troubleshoot complex distributed-system issues, infrastructure failures, and performance bottlenecks.Drive automation and AI-assisted operational improvements across infrastructure platforms.
Required Qualifications:
5+ years of experience in DevOps, Platform Engineering, SRE, or Infrastructure Engineering.3+ years of experience supporting production ML/AI platforms.Strong hands-on experience with Kubernetes and large-scale workload orchestration.Hands-on experience with Kubernetes batch schedulers such as Volcano and/or Kueue.Strong understanding of:Distributed systemsContainerization technologiesCloud-native architecturesMicroservices-based platformsExperience managing production workloads on AWS EKS or similar Kubernetes platforms.Proven experience delivering and supporting cloud/infrastructure migration projects.Strong programming skills in Python or Golang.Experience with Terraform and Helm.Strong experience developing and maintaining CI/CD pipelines using:GitHub ActionsJenkinsGitLab CI/CDAzure DevOpsStrong Linux administration and troubleshooting skills.Experience with observability and monitoring tools such as:PrometheusGrafanaCloud-native monitoring solutionsUnderstanding of ML lifecycle management, model deployment, and production operations.Experience working with ML workflows and Kubeflow.
Preferred Qualifications:
Experience supporting GPU-intensive ML/AI training platforms.Hands-on experience with MLflow and/or Kubeflow.Experience with distributed training frameworks and GPU resource management.Familiarity with LLMOps, Generative AI, RAG architectures, and vector databases.Knowledge of GitOps tools such as ArgoCD or Flux.Experience managing multi-cluster Kubernetes environments.Experience with large-scale ML infrastructure and AI platform engineering.
Mandatory Skills:
Distributed SystemsContainerization TechnologiesCloud-Native ArchitecturesMicroservicesCI/CDGitHub Actions / Jenkins / GitLab CI/CDAWS EKSKubernetesML WorkflowsKubeflowTerraform / HelmPython or GolangPrometheus / GrafanaLinuxInfrastructure Automation
kashish.kulbhaje@aptino.com / kashish.kulbhaje@aptino.biz
Contact No - 817 330 7081