MLOps Engineer

Evlo Ai — United States · Posted ~2 hours ago

Senior Full-time

Skills

MLOps ML CI/CD Machine learning infrastructure Kubernetes Model serving Training pipelines Batch inference ML observability Cloud infrastructure GitHub Actions Kubeflow MLflow KServe Seldon Triton Airflow Kubeflow Pipelines Ray AWS GCP Evidently

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking an MLOps engineer to own the infrastructure layer that reliably runs machine-learning models in production. Responsibilities include ML CI/CD, scalable training and inference pipelines, Kubernetes-based model serving, autoscaling, canary deployments, rollbacks, and lifecycle observability. The role works closely with ML engineers and data scientists and provides significant ownership of platform architecture.

Highlights

High-impact infrastructure role owning the ML production platform, including CI/CD, scalable training, model serving, autoscaling, observability, and cloud infrastructure, with direct influence on engineering delivery speed.

Description

About The Role The role owns the ML infrastructure layer that keeps models running reliably in production - CI/CD for ML, scalable training pipelines, model serving infrastructure, and observability across the entire ML lifecycle. You will work directly with ML engineers and data scientists to turn experimental models into production systems, and own the platform decisions that determine how fast the team ships. Key Responsibilities Build and maintain CI/CD pipelines for ML workflows using tools like GitHub Actions, Kubeflow, or MLflow, enabling rapid and safe model releasesDesign and operate model serving infrastructure on Kubernetes (KServe, Seldon, or Triton), handling autoscaling, canary deployments, and rollbacksOrchestrate large-scale training and batch inference pipelines with Airflow, Kubeflow Pipelines, or Ray, on AWS or GCPImplement model monitoring and observability: data drift detection, latency/throughput dashboards, and automated alerting with tools like Evidently, Grafana, or DatadogEstablish experiment tracking and model registry standards (MLflow, Weights & Biases) to ensure reproducibility across teamsOptimize infrastructure cost and performance - GPU utilization, spot instance strategies, and inference quantization/batchingPartner with data scientists to debug production issues, improve deployment velocity, and codify MLOps best practices across the team What We Are Looking For 3–6 years of experience in MLOps, ML platform engineering, or backend/DevOps engineering with significant ML infrastructure exposureStrong Python and Go (or similar) skills, with solid software engineering fundamentals in production systemsHands-on experience with Kubernetes in production, including deploying and scaling stateful and GPU-accelerated workloadsDeep familiarity with at least one major cloud provider (AWS, GCP, or Azure) and their ML-specific servicesExperience with workflow orchestration (Airflow, Kubeflow, Prefect) and model serving frameworks (TorchServe, Triton, KServe, Seldon)BS in Computer Science, Engineering, or equivalent practical experience; MS a plusBonus: Experience with feature stores (Feast, Tecton), LLM inference optimization (vLLM, TensorRT-LLM), or contributing to open-source MLOps tooling