Skills
Machine learning infrastructure
GPU compute platform management
GPU cluster design and deployment
Infrastructure engineering
Large-scale model training
Reinforcement learning infrastructure
Distributed computing
GPU
Machine Learning
Reinforcement Learning
Foundation Models
Simulation
Distributed Computing
High-Performance Computing
Summary
✨ AI‑Generated
An ambitious robotics technology company is seeking a senior machine learning infrastructure engineer to own and develop GPU compute platforms supporting large-scale foundation model training. You will contribute to infrastructure design, deployment, and scaling for reinforcement learning and simulation workloads, working with an experienced technical team in a high-impact, on-site role.
Highlights
Join a rapidly growing team developing advanced robotics technology, with significant ownership of GPU infrastructure and large-scale machine learning systems. Work on technically ambitious projects backed by international investors and collaborate with experienced researchers and engineers.
Description
About Flexion
At Flexion, we are building the autonomy stack for humanoid robots.
Our mission is to drive the transition from fragile prototypes to real-world deployments of humanoids.
We were founded by leading scientists in robot reinforcement learning (ex-Nvidia, ex-ETH Zürich) and backed by leading international VC firms.
In just months, we went from our first line of code to deploying real humanoid capabilities with our customers, leveraging simulation and reinforcement learning.
Today, we are rapidly expanding the capabilities of our autonomy stack, our customer base, and our team.
The role
We are looking for an experienced ML engineer to join Flexion's experienced infrastructure team and take ownership of Flexion's GPU compute platforms.
This is a senior, on-site role with significant scope.
At Flexion, we are building the brain for humanoid robots, which involves training foundation models with vast amounts of data on large GPU clusters.
You will own the design, bring-up, operation and optimization of performant clusters.
You will work with AI engineers to help them optimize their training speed and hardware utilization.
You will also influence strategic compute planning and contribute to new tools and platforms for iterating on our AI models efficiently.
This will put you at the heart of Flexion's AI development and allow you to directly impact the execution of our ambitious roadmap.
You will closely collaborate with the company's leadership, engineers of the infrastructure team and AI engineers across the company.
Key Responsibilities
Architect, run and continuously improve existing and future cloud-based GPU clusters.
Select the best frameworks and tooling to run our clusters efficiently.
Work on cluster provisioning, job schedulers and monitoring systemsHelp AI engineers optimize their training workloads and maximize hardware utilization using profilers, contributing to our core ML librariesContribute to short- and long-term GPU compute strategies in collaboration with our AI engineering teams and help execute on them.
Optimize capacity and cost by exploring multi-cloud strategies and evaluating trade-offsRaise the bar on engineering practices, including testing, code quality, documentation, and system reliability
Requirements
Degree in Computer Science, Electrical Engineering or Software Engineering (or equivalent practical experience) plus significant industry experienceHands-on experience with the training or inference of large models (billions of parameters) on distributed multi-node GPU hardware.
This can include bringing up and running the cluster, writing and optimizing training/inference code, building ML pipelines, etcProficiency in Python and working knowledge of PyTorchDeep understanding of distributed training concepts (DDP, FSDP, NCCL)Experience with at least one cloud platform (AWS, GCP, Azure or neoclouds) or large-scale on-premises GPU infrastructureExperience with job scheduling and orchestration tools: Slurm and/or Kubernetes/KubeRay
Nice-to-haves
Familiarity with profilers (e.g., PyTorch Profiler, Dynolog, HTA, Nsight)Experience with high-performance or parallel file systems (e.g., Lustre)Experience provisioning compute on multiple cloud providersExperience with infrastructure-as-code and configuration management (Terraform, Ansible)
Benefits
Competitive CompensationJoining a leading robotics team & exposure to never-done-before researchEnergetic, collaborative culture with a bias for action and regular community events
Zurich
Enhanced pension planRelocation & permit sponsorshipEnhanced holiday & paid leave perksCentral Zürich office with top-tier robotics testing facilities and infrastructure
San Franciso
401(k) with company contributionsHealth, dental & vision coverage with the flexibility to choose your own planOpen PTO policy & paid company holidays