MLOps Engineer

Evlo Ai — United States · Posted ~2 hours ago

Full-time

Skills

MLOps Kubernetes Docker Terraform CI/CD Machine learning model deployment Model monitoring Model serving Triton Inference Server TorchServe vLLM

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An AI engineering team is seeking an MLOps engineer to own the infrastructure, orchestration, and deployment pipelines that move machine learning models from research into production. You will build Kubernetes- and container-based infrastructure, create ML-specific CI/CD workflows, optimize model serving, and implement monitoring for model and system performance. The role sits at the intersection of software engineering and machine learning.

Highlights

Build machine learning infrastructure at scale, own end-to-end model deployment pipelines, work with cloud-native technologies, optimize high-throughput model serving, and collaborate closely with data scientists and ML engineers.

Description

About The Role The role owns the infrastructure, orchestration, and deployment pipelines that power machine learning at scale, ensuring models move reliably from research to production. The team works at the intersection of software engineering and machine learning to build resilient platforms that handle continuous integration, monitoring, and scaling for complex AI workloads. Key Responsibilities Design and build robust MLOps infrastructure using Kubernetes, Docker, Terraform, and cloud-native toolsImplement end-to-end CI/CD pipelines specifically tailored for machine learning model training, testing, and deploymentManage and optimize model serving layers using platforms like Triton Inference Server, TorchServe, or vLLM to maintain low latency and high throughputConfigure automated monitoring systems to detect data drift, concept drift, and system performance anomalies in production modelsCollaborate with data scientists and ML engineers to containerize models and streamline feature store integrationsEstablish security, compliance, and governance best practices for model artifacts, datasets, and infrastructure What We Are Looking For 3–6 years of experience in MLOps, DevOps, or reliability engineering, with a focus on machine learning infrastructureDeep proficiency in Python, containerization (Docker), and container orchestration (Kubernetes)Hands-on experience with cloud platforms (AWS, GCP, or Azure) and infrastructure-as-code tools like TerraformFamiliarity with ML lifecycle tools such as MLflow, Kubeflow, Weights & Biases, or SageMakerSolid understanding of CI/CD principles, logging, monitoring, and distributed systemsBonus: Experience managing LLM serving infrastructure or vector databases at scale