Summary
✨ AI‑Generated
A robotics-focused engineering role for building the production backbone behind machine-learning systems. You will create reproducible training and evaluation pipelines, manage GPU infrastructure, automate model releases, and improve reliability across cloud, simulation, laboratory, and robotic environments.
Highlights
Build production-grade ML infrastructure for robotics, with a strong focus on reproducibility, efficient GPU usage, safe deployment, observability, and improving researcher velocity.
Description
← Back to Careers
ML Systems / MLOps Engineer: Robotics
Seoul Gangnam (On-site)
Full-time·Talent Pool / Year-round Interest
About The Role
Robot models improve only as fast as the systems that train, evaluate, release, and observe them.
We are building a talent pool of ML systems and MLOps engineers to create the production backbone for Holiday Lab, our platform for turning real-world data into reusable robot skills.
You will make experiments reproducible, GPU resources efficient, models traceable, and deployments safe across cloud, lab, simulation, and robot environments.
The role is measured by researcher velocity and production reliability—not infrastructure for its own sake.
Key Areas
Distributed trainingExperiment managementModel registryEvaluation automationGPU orchestrationEdge deploymentObservability
What You'll Do
Build reproducible training, evaluation, packaging, and release pipelines for perception, robot learning, VLA, and world models.Operate GPU clusters and schedulers; improve utilization, throughput, queueing, checkpointing, and failure recovery.Create experiment tracking, dataset/model lineage, registry, artifact, and configuration systems.Develop deployment and rollback workflows for simulators, lab machines, and robot compute.Implement monitoring for data drift, model regressions, latency, resource use, and task-level robot outcomes.
Required Qualifications
Professional experience building ML platforms, distributed systems, data infrastructure, or production ML services.Strong Python and Linux skills, including debugging processes, networking, storage, and GPU workloads.Experience with containers, CI/CD, workflow orchestration, experiment tracking, and model or artifact registries.Experience operating cloud or on-prem compute and handling reliability, security, observability, and cost tradeoffs.Ability to work closely with researchers and convert recurring friction into simple, durable platform capabilities.
Preferred Qualifications
Experience with Kubernetes, Slurm, Ray, Kubeflow, Airflow, Argo, MLflow, Weights & Biases, or equivalent systems.Experience with multi-node GPU training, high-throughput storage, data caching, and checkpoint optimization.Experience deploying models to NVIDIA edge devices, robot computers, or intermittent-connectivity environments.Experience with ROS/ROS2, simulation farms, hardware-in-the-loop, or robotics release workflows.Experience supporting foundation-model training and evaluation at meaningful scale.
What We Offer
Competitive compensation based on experience, level, and technical impact.The tools, compute, robots, and equipment needed to do your best work.The opportunity to build, test, and deploy directly on FRIDAY and the Holiday Robotics full stack.Close collaboration with researchers and engineers across hardware, control, simulation, learning, perception, data, and operations.
This is an evergreen talent-pool posting.
We review profiles continuously, and the timing and level of a formal hiring process may depend on team priorities and available openings.
If your work matches our direction, we would still like to hear from you.
Apply