Senior Software Engineer - Infrastructure Platform

Thomas To — United States · Posted ~2 hours ago

Senior Full-time

Skills

Kubernetes cloud infrastructure software engineering distributed systems cloud GPU infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior engineering role focused on building Kubernetes-native infrastructure platforms, optimizing large-scale computing environments, and supporting modern AI workloads.

Highlights

Develop advanced cloud-native infrastructure solutions supporting large-scale machine learning workloads.

Description

Are you passionate about Kubernetes and AI and want to help build the best platform for ML/AI infrastructure? Do you thrive when your work directly empowers teams to push the boundaries of what's possible? We're a collaborative group of engineers, architects, and SREs who are passionate about building and nurturing the declarative, Kubernetes-native control plane that powers GPU-accelerated infrastructure across multiple cloud providers. We are building a platform that gathers topology related information from multiple sources and systems, aggregates and normalizes that data, and makes it available to provisioning systems and workload schedulers. We are looking for Senior Software Engineer who will be directly involved in not only helping maintain this critical open-source project for the community, but interfacing with bleeding edge NVIDIA hardware to ensure GPU to GPU communication is optimized for large-scale workloads across multiple providers. What You'll Be Doing Building a system that gathers topology related information from multiple sourcesTaking data collected to aggregate and normalize the data to make it available for provisioning systems and workload schedulersDirect contributor in a critical open-source project, TopographInteracting with the latest and greatest hardware to ensure new product launches have the most efficient scheduling capabilities What We Need To See At least 8 years of relevant experienceBachelor’s degree in Computer Science, Software Engineering, Computer Engineering, or a related technical field, or equivalent experience.Strong production engineering experience in Go or another systems language.Experience with distributed systems, Kubernetes, Slurm/Slinky, Linux, containers, APIs, and CI.Ability to design clean interfaces between discovery logic, data models, and scheduler output.Familiarity with networking, cluster topology, cloud infrastructure, or large-scale compute systems.Excellent testing, debugging, documentation, and code review habits: Ways To Stand Out From The Crowd Experience with GPU clusters, NVLink, InfiniBand, Ethernet fabrics, or HPC.Hands-on work with Kubernetes scheduling, Slurm/Slinky topology, DRA, Kueue, Slinky, or device plugins.Experience integrating with cloud provider topology APIs or cluster metadata systems. NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for great people like you to help us accelerate the next wave of artificial intelligence. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're a creative, curious, and driven technical leader, we want to hear from you! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until August 23, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.