Summary
An infrastructure engineering role responsible for designing and operating global-scale cloud and GPU-based systems. The position involves reliability, performance optimization, deployment automation, and ownership of critical production infrastructure.
Highlights
High-ownership infrastructure role with end-to-end responsibility, significant equity opportunity, and the chance to build large-scale AI infrastructure from the ground up.
Description
San Francisco, CA
On-site Full-time Compensation: $150,000โ$250,000 + up to 1% equity
About The Company
A seed-stage LLM interpretability and context-optimization company building custom machine learning models that analyze and compress token contexts before they reach the underlying model โ delivering roughly 50% inference cost reduction, lower latency, and higher accuracy for the enterprises and scale-ups integrating LLMs into their products.
Venture-backed, with strong early traction (:1,000 customers within its first seven months).
Founded 2025
1โ10 people Industry: AI Tools / LLM infrastructure
The Role
Own the full multi-region GPU infrastructure stack end to end as the sole infra hire โ global low-latency serving, multi-cloud and on-premise deployments, reliability, and cost efficiency.
Your work sits directly in the critical path of live customer traffic.
This is a high-ownership, in-person role at a fast-moving early-stage team, working a 996 pace (9amโ9pm, six days a week) in San Francisco.
What You'll Be Doing
Own the full infrastructure stack end to end across multi-region GPU deployments, major public clouds, and on-premise enterprise environments.Build and maintain super low-latency GPU serving infrastructure that sits in the critical path of live customer traffic.Manage multi-cloud deployments, including cloud marketplace integrations and provider relationships.Design and iterate on deployment, scaling, reliability, and cost-efficiency systems as the sole infra owner.Support on-premise deployments for enterprise clients and ensure performance and reliability at each site.Research and adopt new infrastructure solutions continuously as the stack and customer base grow.
Tech stack: AWS, GCP, Terraform, Docker, CI/CD, GPU/ML inference infrastructure
Requirements
Own cloud systems serving compression API end-to-endBuild and operate global low-latency high-throughput GPU ML inference infrastructureWork with AWS, Terraform, Docker and CI/CDHave built and operated production infrastructure at a startup or larger companyLearn new solutions and technologies quicklyImprove and research infrastructure solutions continuouslyBased in or willing to relocate to San Francisco to work in person at the hacker houseWillingness to work startup hours in a 996-style environment (9amโ9pm, six days a week)
Green Flags
Quick learner who grasps products and systems fastExperience building for performance and reliability at scaleResearch and product focus mindsetHigh ownership mentalityStartup-minded operator who prioritizes learning and growth over work-life balanceGPU infrastructure experience in productionFirst infra hire at a startupBackground at an infrastructure company
Red Flags
Infra scope limited to model training pipelines only20+ years of experience with a slow-moving, process-heavy backgroundPrioritizes work-life balance as a primary requirementNo production infra ownership
Why Join
Sole infra owner with full-stack ownership from day one, directly in the critical path of live customer traffic.Well-funded seed-stage company with strong early traction and an experienced backer base.Significant equity, housing and food provided at the SF hacker house, visa sponsorship, laundry and cleaning, company off-sites, infinite DoorDash, and health & dental.
Details
Location: San Francisco, CAWork policy: In person (SF hacker house); 996 pace โ 9amโ9pm, six days a weekCompensation: $150,000โ$250,000 + up to 1% equityVisa sponsorship: Available (H-1B, O-1, OPT)Employment type: Full-time