Senior Platform Engineer - GPU Cloud Infrastructure

Stratitechservices Llc — United States · Posted ~3 hours ago

Senior Full-time Onsite No Visa $210000-$260000 base + equity

Skills

cloud infrastructure platform engineering GPU computing DevOps Kubernetes Python Terraform Go GPU

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior platform engineering role focused on creating scalable GPU cloud infrastructure and solving complex infrastructure challenges for advanced computing workloads.

Highlights

Build foundational infrastructure for advanced computing platforms with competitive compensation and significant technical ownership.

Description

Senior Platform Engineer — GPU Cloud Infrastructure Location: San Francisco, CA — Onsite 5 days/week Employment Type: Full-Time Compensation: $210,000–$260,000 base + equity Work Authorization: Candidates must be currently authorized to work in the United States. This position is not eligible for new or future employer-sponsored work authorization. Note: No C2C arrangements will be considered. Any attempt to use personal or household contact information for solicitation, candidate submission, or vendor outreach is strictly prohibited and will be reported to LinkedIn. About the Role What would it be like to help build a cloud before all of the important infrastructure decisions have already been made? We are partnering with a highly technical, well-funded AI company launching a new GPU-focused neocloud designed to provide high-performance accelerated computing infrastructure. We are looking for a Senior Platform Engineer to help build the foundational systems behind this new GPU cloud. This is a true 0→1 infrastructure role where you will help determine what should be built, make architectural decisions, build the systems, and own them through production. This is not a traditional production-operations SRE role focused primarily on maintaining an established stack. The environment spans Kubernetes, Linux, networking, storage, Infrastructure as Code, distributed systems, observability, automation, and GPU-accelerated compute. What You'll Do Design and build foundational infrastructure for a new GPU cloud platform.Own major infrastructure projects from architecture through implementation and production.Design and evolve Kubernetes-based infrastructure for distributed, compute-intensive workloads.Build platform software, automation, controllers, and internal tooling using Python and/or Go.Develop reusable infrastructure using Terraform / Infrastructure as Code.Build deployment, provisioning, orchestration, and lifecycle-management capabilities.Develop self-service infrastructure that makes complex compute resources easier to consume.Work across Linux, networking, storage, containers, Kubernetes, and distributed systems.Build observability and diagnostic capabilities for large distributed environments.Troubleshoot complex issues across application, OS, network, storage, container, and infrastructure boundaries.Help establish technical patterns and architecture while the platform is still being created. What We're Looking For Approximately 8+ years of relevant engineering experience.Demonstrated experience building infrastructure or platforms, not primarily operating systems built by others.Proven ownership of meaningful projects from architecture through production.Strong Linux and systems fundamentals.Deep hands-on experience with production Kubernetes.Strong Terraform / Infrastructure as Code experience.Production programming or substantial infrastructure automation experience with Python, Go, or another systems-oriented language.Experience building or significantly extending CI/CD, deployment systems, developer infrastructure, or platform tooling.Strong understanding of distributed systems and cross-layer troubleshooting.Strong architectural judgment and ability to explain technical tradeoffs.Comfort working in a 0→1 environment with incomplete requirements and significant individual ownership.Desire to remain deeply hands-on. Especially Valuable Experience We are particularly interested in engineers who have: Built a platform or infrastructure capability from scratch.Served as an early, founding, or first infrastructure/platform engineer.Built internal developer platforms or self-service infrastructure.Written Kubernetes operators, controllers, schedulers, or provisioning systems.Built distributed compute, storage, networking, or data infrastructure.Independently turned ambiguous technical objectives into production systems. Nice to Have Experience with any of the following is valuable but not required: GPU infrastructure • NVIDIA ecosystems • GPU scheduling/orchestration • NVIDIA GPU Operator • CUDA • DGX • InfiniBand • RDMA/RoCE • High-performance networking • Distributed storage • Bare metal • AI/ML infrastructure • Ray • Slurm • ArgoCD/GitOps • Helm Direct GPU or HPC experience is a plus, not a requirement. Exceptional platform and systems engineers who can quickly learn new infrastructure domains are strongly encouraged to apply. Why This Opportunity Most senior infrastructure roles ask you to inherit a platform. This one gives you the opportunity to help create one. You will join while fundamental decisions are still being made about compute provisioning, GPU orchestration, workload scheduling, networking, storage, automation, observability, and developer experience. If you enjoy 0→1 infrastructure engineering, substantial technical ownership, and building systems rather than simply maintaining them, this is an opportunity to work unusually close to the foundation of a new cloud platform.