Founding GPU Infrastructure Engineer

Sentiropartners — Unknown · Posted ~1 day ago

Lead Full-time Onsite

Skills

GPU infrastructure HPC AI systems Kubernetes Slurm Distributed systems Distributed object storage NVMe High-performance networking Distributed training MPI InfiniBand RDMA Systems software Production operations Observability Reliability engineering NVIDIA GPUs

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary

Join a seed-stage AI infrastructure company as a founding engineer and build a production GPU cloud platform from bare metal through customer-ready inference. Commission GPU clusters, operate Kubernetes and Slurm environments, develop orchestration systems, support distributed training and batch workloads, and design high-performance storage and networking. You will own observability, reliability, cluster lifecycle tooling, and deployment patterns while helping establish the infrastructure team. The role is primarily in person and includes relocation support for exceptional candidates.

Highlights

Founding infrastructure role with meaningful equity, competitive compensation, performance bonus, benefits, direct founder collaboration, relocation support, and the opportunity to build GPU infrastructure and an engineering team from the ground up.

Description

FOUNDING GPU INFRASTRUCTURE ENGINEER [HPC & AI SYSTEMS] - MEANINGFUL EQUITY REPEAT FOUNDERS WITH SERIOUSLY IMPRESSIVE EXITS. Build a GPU neocloud from bare metal to inference San Francisco | On-site (5 days with some flex) | Relocation support Most infrastructure roles ask you to keep somebody else’s platform alive. This one asks you to build the platform. Sentiro Partners has been retained to find the Founding GPU Infrastructure Engineer for my client, a seed-stage GPU neocloud operating in stealth in San Francisco. The company is building managed GPU infrastructure and high-performance inference for demanding AI workloads. It has secured backing from well-known VCs and is pursuing deployments at hundreds-of-GPUs scale. THE FOUNDERS HAVE DONE THIS BEFORE One founder built an AI infrastructure company operating at the HW layer, bringing low-latency AI computation to resource-constrained devices. Acquired. The other founded & scaled an enterprise technology company to 400+ customers. Their previous companies were backed by major VCs. Ivy league & top tier finance backgrounds. Seriously down to earth. Seriously hard workers. They are in the SoMa office right now, as you read this. Now they’re building again. WHAT SUCCESS LOOKS LIKE Success means taking new GPU capacity from delivered racks to reliable, customer-ready infrastructure and building the systems and team required to repeat that process at increasing scale. You will be the company’s founding infrastructure engineer and one of its first technical hires. When the racks arrive, you will help turn them into a production platform: - Commission and burn in NVIDIA GPU clusters. - Build and operate Kubernetes and Slurm infrastructure. - Develop the orchestration layer above the underlying clusters. - Enable distributed training, MPI and batch workloads. - Architect distributed object storage and high-performance NVMe systems. - Build around InfiniBand, RDMA and other high-bandwidth networking. - Own telemetry, observability, reliability and cluster lifecycle tooling. - Help develop low-latency, high-throughput inference infrastructure. - Create the interfaces through which customers access and operate the platform. - Establish the patterns that make each deployment faster and more reliable than the last. THIS IS GPU INFRASTRUCTURE BUILT WITH HPC DISCIPLINE You should understand what happens between a rack arriving at a facility and a customer successfully running a distributed GPU workload. You don’t need to be a data-centre real-estate or cooling specialist. You do need enough hardware and systems depth to reason about GPU topology, storage, networking, workload behaviour and production reliability as one connected system. WE SHOULD TALK IF YOU HAVE - c. 3-7 years in GPU infrastructure, HPC, AI systems or distributed systems (we're not limited by exp. but you need to be able to go deep). - Meaningful production experience with both Kubernetes and Slurm (either tbh) - Built infrastructure rather than only deployed platforms created by other teams. - Strong knowledge of distributed object storage, NVMe and high-performance networking. - Experience with distributed training, MPI, InfiniBand, RDMA or similar technologies. - Worked across hardware, systems software and production operations. - The ability to move quickly and make sound decisions with incomplete information. - The appetite to own outcomes rather than defend a narrow area of responsibility. Particularly relevant backgrounds include AI labs, GPU neoclouds, quantitative trading firms, hyperscale infrastructure teams and serious HPC environments. Experience with GPU power capping, power-aware scheduling, energy telemetry or grid-flexible compute would be unusually valuable, but it isn’t essential. WHY NOW? Because the upside is real. You will work directly with the founders who have already raised capital, scaled teams and delivered successful exits. You will influence the architecture before the boundaries harden, help deliver the company’s first major GPU clusters and have the opportunity to build the infrastructure team around you. This is a founding role with a competitive San Francisco salary, performance bonus, benefits and meaningful equity. The team is building in person in San Francisco. Relocation support is available for exceptional candidates ready to move. If your ideal role begins where the racks arrive—and ends with customers running serious AI workload... contact me directly. Adrian Clarke Founder, MD, Executive Search Partner Sentiro Partners Frontier AI Search & Advisory (https://sentiropartners.com) ABOUT SENTIRO PARTNERS Sentiro Partners is a global executive search firm specialising in frontier AI, deep technology and high-performance technical leadership.We partner with ambitious founders, AI labs and technology companies to identify the engineers, researchers and leaders building the next generation of intelligent infrastructure.Headquartered in Dublin, Sentiro Partners conducts searches across North America, Europe and Asia-Pacific.