Platform Engineer

Reactor World — United States · Posted ~2 days ago

Senior Full-time

Skills

Kubernetes Multi-cloud Infrastructure GPU Orchestration GitOps Helm Kustomize Infrastructure-as-Code Observability Real-time Networking AWS GPU Cloud Providers

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An AI-focused technology company in San Francisco is hiring a Platform Engineer to own their infrastructure platform. This is not a CI/CD-focused DevOps role — you will work across GPU orchestration, multi-cloud Kubernetes, real-time networking, and observability. You will be the person who knows why a model pod took four minutes to schedule, why cross-region latency spiked, or why a media relay is dropping packets. The company runs production today across multiple Kubernetes clusters, regions, and GPU types, and is actively expanding to additional cloud providers. Responsibilities include provisioning and managing multi-region Kubernetes clusters using infrastructure-as-code, owning the GitOps deployment lifecycle, managing GPU node infrastructure, and leading cloud expansion efforts.

Highlights

Own the infrastructure platform powering AI models in production. Work on cutting-edge GPU orchestration, multi-cloud Kubernetes, and real-time networking. Lead the expansion to additional cloud providers in a high-impact engineering role based in San Francisco.

Description

Department: Engineering Location: San Francisco Description You'll own the infrastructure platform that our AI models run on. This isn't a CI/CD-focused DevOps role. You'll work across GPU orchestration, multi-cloud Kubernetes, real-time networking, and observability. You'll be the person who knows why a model pod took 4 minutes to schedule, why cross-region latency spiked, or why a media relay is dropping packets. We run production today across multiple Kubernetes clusters, regions, and GPU types, and we're actively expanding to additional cloud providers. You'll lead that expansion and keep everything running. What You'll Do Provision and manage multi-region Kubernetes clusters across AWS and GPU cloud providers using infrastructure-as-code.Own the GitOps deployment lifecycle (Helm charts, Kustomize overlays, image automation, and continuous delivery.)Manage GPU node infrastructure: scheduling, model weight caching, image prefetching for fast cold starts, and GPU observability.Operate and improve our networking layer: ingress and gateway management, load balancing, media relay infrastructure, and cross-region connectivity.Build and maintain our observability stack: metrics, logs, traces, and profiling across all services and GPU workloads.Maintain infrastructure security: IAM, secret management, certificate automation, and encryption at rest.Own CI/CD pipelines for monorepo builds spanning Go services, Python model containers, and Helm chart releases.Partner with ML engineers on model serving: container optimization, health checks and startup tuning, media pipeline performance, and multi-GPU configuration. What We're Looking For You've operated Kubernetes in production at scale, not just deployed to it, but debugged node-level scheduling issues, tuned autoscalers, and managed cluster upgrades.Strong infrastructure-as-code experience (Terraform, Pulumi, or similar) across multiple environments and regions.You've worked with GPU workloads on Kubernetes: device plugins, node taints/tolerations, GPU-aware scheduling. You understand why bin-packing matters for expensive hardware.Experience with GitOps tooling (FluxCD, ArgoCD, or similar) and Helm chart authoring.Comfortable with Redis or similar in-memory data stores (replication, persistence, pub/sub or streaming patterns)Familiarity with modern observability stacks (Prometheus, Grafana, OpenTelemetry, or equivalent) and knowing when to reach for metrics vs. logs vs. traces.Solid networking fundamentals: load balancers, TLS, DNS, NAT. Real-time or low-latency networking experience is a strong plus.You've worked in a startup where you owned infrastructure end-to-end, not just one slice of it.Nice to HaveExperience with GPU cloud providers beyond AWS (Crusoe, CoreWeave, Lambda Labs, Nebius)Real-time media or streaming infrastructureGo or Python proficiencyFamiliarity with ML model serving (container image optimization, weight loading, GPU driver and runtime management)FinOps and GPU cost optimizationWhat We're Not Looking ForPure CI/CD pipeline engineers who haven't operated Kubernetes clusters directlyCandidates whose infrastructure experience is limited to managed PaaS (Heroku, Vercel, Railway)People who need a fully defined scope, this role requires figuring out what to build next, not just executing tickets Benefits Competitive San Francisco salary and meaningful equityWe sponsor visas and support relocation to the USGenerous health, dental, and vision coverage