Machine Learning Infrastructure Engineer
David Joseph Inc — United States · Posted ~2 hours ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
San Francisco, CA
On-site (5 days/week) Full-time Compensation: $200K–$400K + competitive early-stage equity
About The Company
Our client is a Series A AI research lab building large-scale foundation models for scientific and physical-AI domains.
Backed by top-tier investors, they are pursuing a deliberately non-consensus technical thesis and are among the best-funded teams in their space.
The founding team comes from self-driving, robotics, and scientific research, and they are scaling their research and engineering org significantly this year.
Founded 2024
Small, fast-growing team Industry: AI / foundation models / physical AI
The Role
You would own the distributed training and inference backbone for a foundation model trained from scratch — standing up clusters, building data and training pipelines at petabyte scale, and squeezing performance out of GPUs at a low level across model scales.
What You'll Be Doing
Design, deploy, and maintain large distributed ML training and inference clustersBuild efficient, scalable end-to-end pipelines to manage petabyte-scale datasets and training across the full ML lifecycleResearch and test training approaches, including parallelization techniques and numerical-precision trade-offs across model scalesProfile and debug low-level GPU operations to optimize performanceTrack new research and bring fresh ideas into the work
Tech stack: Distributed training frameworks (FSDP, DeepSpeed), NVIDIA GPUs, Linux, Python, C++, Kubernetes/Docker, and a major cloud platform (GCP, AWS, or Azure).
Requirements
2–10 years building large-scale ML infrastructure for core foundation modelsHands-on experience building infrastructure for foundation models trained from scratch, rather than fine-tuning existing modelsA background at a science-focused or physical-AI company (for example self-driving, robotics, or biology)Deep, demonstrable expertise optimizing large-scale training and inference workloadsWorking proficiency with distributed training frameworks such as FSDP or DeepSpeedA clear pattern of intentional, mission-driven career decisionsAble to work on-site 5 days/week in San Francisco (relocation supported)
Nice to Haves
Generalist experience spanning the full ML lifecycleLow-level GPU performance optimization and debugging (CUDA, JAX)
Why Join
Take a bet on a distinctive, non-consensus approach to building intelligenceJoin early, with real ownership of the training and inference backboneWork in a domain with fast, objective ground-truth feedback and data at a scale beyond typical LLM trainingWell-funded and building a strong, senior research and engineering team
Details
Location: San Francisco, CAWork policy: In-person 5 days/week (relocation supported)Compensation: $200K–$400K + competitive early-stage equityVisa sponsorship: Open to supporting work authorization for the right candidateEmployment type: Full-time
We have 109,996 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 109,996 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume