Summary
✨ AI‑Generated
A hands-on engineering role focused on owning the production lifecycle of machine learning systems at scale. You will build deployment and retraining pipelines, manage cloud infrastructure and costs, implement observability and drift detection, respond to production incidents, and collaborate with ML specialists to deliver reliable systems. The role also includes mentoring and supporting technical demonstrations.
Highlights
Own end-to-end production ML infrastructure and reliability, with strong exposure to cloud optimization, automation, observability, incident response, and cross-functional collaboration. The role also offers opportunities to mentor engineers and contribute to technical demonstrations.
Description
Own the production lifecycle of machine learning models, ensuring validated models are reliably deployed, monitored, optimized, and maintained at scale.
The role focuses on MLOps, cloud infrastructure, automation, observability, cost optimization, and production incident management.
Responsibilities:
Own production serving, CI/CD, deployment, and monitoring of ML models.Build and maintain model retraining, versioning, and deployment pipelines.Manage ML infrastructure and optimize cloud costs through FinOps practices.Implement observability, alerting, model drift detection, and performance monitoring.Own production incident response, troubleshooting, and on-call responsibilities.Collaborate with Data Scientists and ML Engineers to design scalable, production-ready systems.Support presales and PoCs by demonstrating production readiness and scalability.Mentor engineers on MLOps and production-readiness best practices.Communicate infrastructure cost, performance, and reliability trade-offs to non-technical stakeholders.
Qualifications:
6+ years of experience in MLOps, ML Platform Engineering, or related roles with proven production ownership.Strong expertise in CI/CD, containerization, cloud infrastructure, and ML observability.Deep understanding of the ML model lifecycle, including retraining, model versioning, and drift detection.Experience with Infrastructure as Code (IaC) and automated deployment pipelines.Strong knowledge of major cloud platforms such as AWS, Azure, or GCP.Experience with cloud cost monitoring, optimization, and FinOps practices.Strong understanding of monitoring, alerting, SLA management, and production incident response.Ability to troubleshoot and communicate technical incidents clearly to business stakeholders.Strong collaboration and mentoring skills.Willingness to participate in on-call and off-hours production support