Summary
✨ AI‑Generated
Build and operate the infrastructure that takes machine learning models from experimentation into secure, measurable production services. The role covers training and deployment pipelines, model and data workflows, Kubernetes-based inference, cloud infrastructure, observability, and governance while partnering closely with ML and platform specialists.
Highlights
Work at the intersection of machine learning and platform engineering, building reliable model delivery infrastructure, improving deployment velocity and inference reliability, and partnering with multidisciplinary technical teams.
Description
About The Role
The MLOps Engineer builds and operates the infrastructure that moves machine learning models from experimentation into reliable production services.
The role spans training pipelines, model registries, feature and data workflows, deployment automation, observability, and the cloud systems that support real-time and batch inference.
You will partner with ML engineers, data scientists, and platform engineers to make model delivery repeatable, secure, and measurable.
The work directly affects deployment velocity, inference reliability, model governance, and the ability to scale AI products across teams.
Key Responsibilities
Build and maintain CI/CD pipelines for model training, validation, packaging, and deployment using tools such as GitHub Actions, Argo CD, Kubeflow, or MLflowDeploy and operate batch and real-time inference services on Kubernetes and cloud platforms such as AWS, GCP, or AzureImplement model and data versioning workflows, experiment tracking, approval gates, and reproducible training pipelinesDevelop monitoring for service health, latency, throughput, resource utilization, data drift, model quality, and prediction performance using Prometheus, Grafana, or equivalent toolsAutomate infrastructure provisioning and environment management with Terraform, Helm, Docker, and KubernetesImprove the reliability and cost efficiency of GPU and CPU workloads through capacity planning, autoscaling, performance testing, and operational runbooksPartner with security and engineering teams to enforce access controls, secrets management, auditability, and governance across the ML platform
What We Are Looking For
3–8 years of experience in MLOps, machine learning infrastructure, DevOps, site reliability engineering, or a closely related engineering disciplineStrong Python and production software engineering skills, with experience building APIs, automation services, and reliable data or ML workflowsHands-on experience with Kubernetes, Docker, infrastructure as code, and cloud services on AWS, GCP, or AzurePractical experience with ML lifecycle tools such as MLflow, Kubeflow, SageMaker, Vertex AI, or equivalent platformsWorking knowledge of CI/CD, GitOps, distributed systems, observability, and production incident responseBachelor’s degree in computer science, engineering, mathematics, or a related technical field, or equivalent professional experienceBonus: experience operating GPU workloads, feature stores, Spark, Airflow, Ray, model serving frameworks such as KServe or Seldon, and familiarity with model governance or security controls