ML Infrastructure Engineer

Inventure Recruitment — United States · Posted ~2 days ago

Senior Full-time Onsite $220K-$350K base

Skills

ML infrastructure Training infrastructure Distributed systems Robotics Embodied AI Distributed training systems

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

We are seeking an ML Infrastructure Engineer to join a robotics company in Redwood City, CA (fully onsite). This role involves owning training infrastructure end-to-end for general-purpose robotic arms powered by an embodied AI foundation model. The robots are already deployed at real customer sites doing commercial-grade work. The company has raised significant funding from top-tier investors and is building a team with deep expertise in AI. The near-term focus is shipping next-generation robots with whole-body control, with a long-term vision of physical AGI. Base salary range is $220K–$350K.

Highlights

Rare opportunity to own training infrastructure end-to-end at a well-funded robotics company. Competitive base salary of $220K–$350K with work on cutting-edge embodied AI and physical AGI.

Description

ML Infrastructure Engineer | Redwood City, CA (Onsite) | $220K–$350K base I've partnered with a robotics company building general-purpose robotic arms powered by their own embodied AI foundation model, and unlike most of the field they are already out of the lab: their robots are deployed at real customer sites doing commercial-grade work, starting in the hospitality and restaurant world. They have raised over $140M, including a $120M Series A, from a top-tier investor group spanning NVIDIA's venture arm, Samsung NEXT, Salesforce Ventures, First Round Capital and CRV. The founders are repeat founders who previously built and sold their last company to a major consumer marketplace, and the team is stacked with people from Google DeepMind, Meta and Cruise. The near-term focus is shipping the next generation of robots with whole-body control; the long-term vision is nothing less than physical AGI. This is a rare opportunity to own training infrastructure end to end and become the connective tissue between researchers and compute. There is no infra layer above you: you'll turn a multi-cloud GPU fleet into a world-class training engine for massive multimodal models, and drive the architecture yourself. It is a genuinely high-ownership seat where every optimisation you ship shortens the path from model to deployed robot. If you want your systems work to move real machines in the physical world rather than sit behind an API, this is it. The Role: As an ML Infrastructure Engineer, you will: Architect and scale distributed training across large GPU clusters, implementing sharding, activation checkpointing and memory optimisation such as ZeRO and FSDP for multimodal models.Build researcher-friendly tooling and job scheduling on Kubernetes and SLURM, with fast iteration, automated retries and seamless failure recovery.Design high-throughput data pipelines that ingest and transform terabytes of multimodal robot data, from video to proprioception to 3D signals, so the dataloaders never starve the GPUs.Build low-latency inference pipelines for real-time robot control, applying quantisation, distillation and model compilation with tools like TensorRT and Triton.Profile deep into the stack, chasing GPU utilisation, I/O bottlenecks and memory fragmentation to squeeze maximum performance out of an expanding compute fleet. About You: You have at least 5 years of infrastructure engineering experience, ideally 7 or more, and you've built and maintained ML or data infrastructure on a team with a high talent bar.You have deep, hands-on PyTorch experience and real command of distributed training frameworks such as DeepSpeed or Accelerate.You are an expert on GPU bottlenecks, model serving optimisation and monitoring, and you enjoy the systems-profiling side of the work.You are genuinely excited about robotics and physical AI, and you thrive in a fast-paced, high-ownership startup environment.One of these backgrounds fits you: an ML or HPC infrastructure engineer from a top AI lab or research team, an early or founding infrastructure hire at a startup, or an infrastructure engineer coming from robotics or autonomous vehicles.Nice to have: robotics experience, or having built multimodal systems for video, audio or other rich media models.You are based in or willing to relocate to the Bay Area and happy to work onsite five days a week in Redwood City.Visa sponsorship is available, including OPT and H1B transfers. The role comes with competitive equity on top of base. The team is moving quickly and hiring several engineers across ML and data infrastructure. If you're interested in owning the training engine behind robots already working in the real world, apply now or send your CV directly to will@inventurerecruitment.com.