MLOps Engineer

Evlo Ai — United States · Posted ~11 hours ago

Mid Full-time

Skills

MLOps CI/CD machine learning deployment Kubernetes cloud infrastructure model versioning data versioning experiment tracking model governance observability GitHub Actions Argo CD Kubeflow MLflow AWS GCP Azure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Build and operate the infrastructure that takes machine learning models from experimentation into secure, measurable production services. The role covers training and deployment pipelines, model and data workflows, Kubernetes-based inference, cloud infrastructure, observability, and governance while partnering closely with ML and platform specialists.

Highlights

Work at the intersection of machine learning and platform engineering, building reliable model delivery infrastructure, improving deployment velocity and inference reliability, and partnering with multidisciplinary technical teams.

Description

About The Role The MLOps Engineer builds and operates the infrastructure that moves machine learning models from experimentation into reliable production services. The role spans training pipelines, model registries, feature and data workflows, deployment automation, observability, and the cloud systems that support real-time and batch inference. You will partner with ML engineers, data scientists, and platform engineers to make model delivery repeatable, secure, and measurable. The work directly affects deployment velocity, inference reliability, model governance, and the ability to scale AI products across teams. Key Responsibilities Build and maintain CI/CD pipelines for model training, validation, packaging, and deployment using tools such as GitHub Actions, Argo CD, Kubeflow, or MLflowDeploy and operate batch and real-time inference services on Kubernetes and cloud platforms such as AWS, GCP, or AzureImplement model and data versioning workflows, experiment tracking, approval gates, and reproducible training pipelinesDevelop monitoring for service health, latency, throughput, resource utilization, data drift, model quality, and prediction performance using Prometheus, Grafana, or equivalent toolsAutomate infrastructure provisioning and environment management with Terraform, Helm, Docker, and KubernetesImprove the reliability and cost efficiency of GPU and CPU workloads through capacity planning, autoscaling, performance testing, and operational runbooksPartner with security and engineering teams to enforce access controls, secrets management, auditability, and governance across the ML platform What We Are Looking For 3–8 years of experience in MLOps, machine learning infrastructure, DevOps, site reliability engineering, or a closely related engineering disciplineStrong Python and production software engineering skills, with experience building APIs, automation services, and reliable data or ML workflowsHands-on experience with Kubernetes, Docker, infrastructure as code, and cloud services on AWS, GCP, or AzurePractical experience with ML lifecycle tools such as MLflow, Kubeflow, SageMaker, Vertex AI, or equivalent platformsWorking knowledge of CI/CD, GitOps, distributed systems, observability, and production incident responseBachelor’s degree in computer science, engineering, mathematics, or a related technical field, or equivalent professional experienceBonus: experience operating GPU workloads, feature stores, Spark, Airflow, Ray, model serving frameworks such as KServe or Seldon, and familiarity with model governance or security controls