Summary
✨ AI‑Generated
A senior ML platform engineer will own the architecture, roadmap, tooling, and operational foundations for production machine learning systems. The role covers training, serving, feature management, inference infrastructure, automated deployment pipelines, testing, versioning, rollback, observability, and technical mentoring.
Highlights
High-impact senior role owning ML platform architecture and roadmap, with strong technical direction, hands-on engineering, mentoring, and responsibility for production-grade AI operations.
Description
About the Role
We're looking for a Senior Platform Engineer to lead the design and operation of the infrastructure that powers our machine learning and AI systems in production.
You'll sit at the intersection of platform engineering and ML operations, owning the architecture, tooling, pipelines, and observability that let data scientists and ML engineers ship models reliably and at scale.
This is a senior, hands-on role for someone who has already built ML/AI platforms and wants to set technical direction — turning experimental ML workflows into repeatable, automated, production-grade systems, mentoring other engineers, and being the go-to owner for how AI runs in production across the company.
What You'll Do
Own the architecture and roadmap for the ML platform: model training, serving, feature stores, and inference infrastructure.Set standards and build CI/CD pipelines for model deployment (MLOps) with automated testing, versioning, rollback, and promotion across environments.Design observability for models and infrastructure — monitoring for latency, drift, data quality, and cost — and drive teams to adopt it.Automate provisioning and scaling of GPU/CPU compute using infrastructure-as-code, and lead cost-optimization efforts.Partner with data science and ML engineering leads to streamline the path from notebook to production and shape platform priorities.Own reliability of AI services end to end: on-call leadership, incident response, capacity planning, and performance tuning.Evaluate, select, and integrate AI Ops tooling to reduce operational toil and improve system self-healing.Establish security, governance, and cost controls across the ML lifecycle.Mentor engineers, review designs, and raise the technical bar of the team.
Required
7+ years in platform, infrastructure, DevOps, or SRE roles, including hands-on experience building and operating ML/AI systems in production.Track record of owning infrastructure or platform initiatives end to end and setting technical direction.Deep experience with a major cloud provider (AWS, GCP, or Azure).Strong expertise with containers and orchestration (Docker, Kubernetes) at scale.Production infrastructure-as-code experience (Terraform, Pulumi, or similar).Strong programming skills in Python and/or Go.Proven experience designing CI/CD pipelines and automation.Solid command of monitoring and observability stacks (Prometheus, Grafana, Datadog, OpenTelemetry).
Nice to Have
Hands-on MLOps tooling: Kubeflow, MLflow, SageMaker, Vertex AI, Ray, or similar.Model serving frameworks (KServe, Seldon, Triton, BentoML).Feature store experience (Feast, Tecton).Experience operating GPU workloads at scale and optimizing inference cost/performance.Exposure to LLMOps — serving, evaluating, or monitoring large language models.Data pipeline experience (Airflow, Spark, dbt).Experience mentoring engineers or leading a technical workstream.