Summary
✨ AI‑Generated
A Senior DevOps / Platform Engineer opportunity at the intersection of software engineering and infrastructure. You will own and evolve scalable cloud and Kubernetes environments, improve production reliability and performance, build CI/CD and infrastructure automation, and develop internal tooling to support increasingly demanding AI workloads.
Highlights
High-ownership senior platform role on a small technical team, with broad influence over cloud and Kubernetes infrastructure, reliability, automation, observability, and infrastructure tooling for demanding AI workloads.
Description
We’re partnering with an early-stage AI company in San Francisco that is building technically complex, real-time AI systems.
They’re looking for a Senior DevOps / Platform Engineer who sits at the intersection of software engineering and infrastructure.
You’ll have significant ownership over the infrastructure, helping design and build the systems required to support increasingly demanding AI workloads as the company scales.
You’ll join a small, highly technical engineering team where you’ll work across Kubernetes, cloud infrastructure, CI/CD, observability, reliability and internal tooling.
Key Responsibilities
Design, build and operate scalable cloud and Kubernetes infrastructure.Improve the reliability, scalability and performance of production systems.Build and evolve CI/CD, deployment and infrastructure automation.Develop internal tooling and software to automate infrastructure and operational workflows.Own infrastructure as code and help establish scalable infrastructure patterns.Design and improve monitoring, observability, alerting and incident-response practices.Troubleshoot complex production issues across applications, infrastructure, networking and distributed systems.Support compute-intensive AI workloads and help evolve the infrastructure as requirements and scale change.Make pragmatic engineering decisions around when to use existing tooling versus configuration, automation or custom software.
Key Requirements
Strong experience with Kubernetes in production environments.Hands-on experience with Terraform or similar Infrastructure-as-Code tooling.Strong understanding of CI/CD and modern deployment practices.Experience building and operating production observability, monitoring and alerting systems.Strong cloud infrastructure experience across AWS, GCP or Azure.Software engineering ability in Go, Python or a similar language.Strong understanding of distributed systems, networking and production reliability.Experience troubleshooting complex infrastructure and production incidents from first principles.Comfortable working in a small, fast-moving environment with significant individual ownership.Experience building or scaling infrastructure rather than solely operating mature, established platforms.
Experience with GPU infrastructure, AI/ML workloads, Kubernetes scheduling/controllers, high-performance compute or large-scale distributed systems would be highly valuable, but isn’t required.