Senior ML Platform Engineer

Incedo Inc — United States · Posted ~3 hours ago

Senior Full-time

Skills

ML platform engineering MLOps Machine learning infrastructure CI/CD Model deployment Model serving Feature stores Observability Infrastructure architecture Inference infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior ML platform engineer will own the architecture, roadmap, tooling, and operational foundations for production machine learning systems. The role covers training, serving, feature management, inference infrastructure, automated deployment pipelines, testing, versioning, rollback, observability, and technical mentoring.

Highlights

High-impact senior role owning ML platform architecture and roadmap, with strong technical direction, hands-on engineering, mentoring, and responsibility for production-grade AI operations.

Description

About the Role We're looking for a Senior Platform Engineer to lead the design and operation of the infrastructure that powers our machine learning and AI systems in production. You'll sit at the intersection of platform engineering and ML operations, owning the architecture, tooling, pipelines, and observability that let data scientists and ML engineers ship models reliably and at scale. This is a senior, hands-on role for someone who has already built ML/AI platforms and wants to set technical direction — turning experimental ML workflows into repeatable, automated, production-grade systems, mentoring other engineers, and being the go-to owner for how AI runs in production across the company. What You'll Do Own the architecture and roadmap for the ML platform: model training, serving, feature stores, and inference infrastructure.Set standards and build CI/CD pipelines for model deployment (MLOps) with automated testing, versioning, rollback, and promotion across environments.Design observability for models and infrastructure — monitoring for latency, drift, data quality, and cost — and drive teams to adopt it.Automate provisioning and scaling of GPU/CPU compute using infrastructure-as-code, and lead cost-optimization efforts.Partner with data science and ML engineering leads to streamline the path from notebook to production and shape platform priorities.Own reliability of AI services end to end: on-call leadership, incident response, capacity planning, and performance tuning.Evaluate, select, and integrate AI Ops tooling to reduce operational toil and improve system self-healing.Establish security, governance, and cost controls across the ML lifecycle.Mentor engineers, review designs, and raise the technical bar of the team. Required 7+ years in platform, infrastructure, DevOps, or SRE roles, including hands-on experience building and operating ML/AI systems in production.Track record of owning infrastructure or platform initiatives end to end and setting technical direction.Deep experience with a major cloud provider (AWS, GCP, or Azure).Strong expertise with containers and orchestration (Docker, Kubernetes) at scale.Production infrastructure-as-code experience (Terraform, Pulumi, or similar).Strong programming skills in Python and/or Go.Proven experience designing CI/CD pipelines and automation.Solid command of monitoring and observability stacks (Prometheus, Grafana, Datadog, OpenTelemetry). Nice to Have Hands-on MLOps tooling: Kubeflow, MLflow, SageMaker, Vertex AI, Ray, or similar.Model serving frameworks (KServe, Seldon, Triton, BentoML).Feature store experience (Feast, Tecton).Experience operating GPU workloads at scale and optimizing inference cost/performance.Exposure to LLMOps — serving, evaluating, or monitoring large language models.Data pipeline experience (Airflow, Spark, dbt).Experience mentoring engineers or leading a technical workstream.