Summary
✨ AI‑Generated
A senior infrastructure engineering role focused on designing, operating, and improving scalable platforms for AI and high-volume workloads. The position involves automation, reliability engineering, and collaboration with software teams.
Highlights
Remote opportunity to build large-scale infrastructure, improve engineering productivity, and work on advanced AI and automation challenges.
Description
Role name: DevOps Engineer (ML Infrastructure)
Work Location: Canada (Remote)
DevOps Engineer (ML Infrastructure)
We are seeking a highly skilled Senior DevOps Engineer to help build, operate, and evolve large-scale, business-critical infrastructure platforms.
This role focuses on reliability, scalability, automation, and operational excellence across distributed systems supporting high-volume production workloads.
You will work closely with software engineers, platform teams, and infrastructure specialists to deliver highly available services, improve developer productivity, modernize infrastructure, and drive innovation through automation and AI-assisted operations.
This is an opportunity to solve complex technical challenges at scale while influencing the future direction of platform engineering and infrastructure management.
To design, build, and operate scalable machine learning infrastructure that enables efficient training, deployment, and monitoring of AI/ML workloads.
The ideal candidate combines deep expertise in cloud-native technologies, Kubernetes, distributed systems, and software engineering to support large-scale machine learning platforms and GPU-based environments.
Key Responsibilities
Design, build, and maintain scalable MLOps platforms for training, deploying, and monitoring machine learning models.Develop and manage cloud-native infrastructure supporting large-scale ML workloads on Kubernetes.Implement and operate batch scheduling solutions such as Volcano and Kueue to optimize utilization of GPU and compute resources.Build and maintain CI/CD pipelines for ML services, infrastructure, and platform components.Automate infrastructure provisioning and lifecycle management using Infrastructure-as-Code practices.Manage and optimize Kubernetes environments, including production workloads running on AWS EKS and other cloud platforms.Lead and support infrastructure modernization and migration initiatives across cloud and platform ecosystems.Partner with Data Scientists, ML Engineers, and Software Engineers to productionize machine learning solutions.Implement robust observability, monitoring, and alerting for distributed systems and GPU clusters.Ensure platform reliability, security, scalability, and operational excellence.Troubleshoot complex distributed systems and performance bottlenecks across infrastructure and ML workloads.
Required Qualifications
5+ years of experience in DevOps, Platform Engineering, Site Reliability Engineering, or Infrastructure Engineering.3+ years of experience supporting production machine learning or AI platforms.Strong experience with Kubernetes and large-scale workload orchestration.Hands-on experience with Kubernetes batch schedulers such as Volcano / KueueSolid understanding of:Distributed systemsContainerization technologiesCloud-native architecturesMicroservices-based platformsExperience managing workloads on Kubernetes platforms such as AWS EKS.Proven track record delivering and supporting infrastructure migration projects.Strong programming skills in Python or GolangExperience with Infrastructure-as-Code and deployment tools including Terraform / HelmExperience designing and maintaining CI/CD pipelines using GitHub Actions, Jenkins, GitLab CI/CD, or Azure DevOps.Strong Linux systems administration and troubleshooting skills.Experience building observability solutions using:PrometheusGrafanaCloud-native monitoring toolsUnderstanding of ML lifecycle management, model deployment, and production operations.Preferred QualificationsExperience supporting GPU-intensive machine learning or AI training platforms.Hands-on experience with MLflow / KubeflowExperience with distributed training frameworks and GPU resource management.Familiarity with LLMOps, Generative AI, RAG architectures, and vector databases is a plusKnowledge of GitOps tools such as ArgoCD or Flux.Experience with multi-cluster Kubernetes environments.Parquet
Mandatory skills:
Distributed systemsContainerization technologiesCloud-native architecturesMicroservices-based platforms CI/CD/Gitubs/Jenkkins,AWS EKS ML workflow Kubflow