Summary
An executive technology leader is sought to own reliability, uptime, incident response, and operational excellence for a large-scale SaaS environment. The role combines leadership of senior global engineering teams with hands-on development of AI-assisted incident investigation, remediation, and change-validation capabilities, while partnering with enterprise stakeholders during critical events.
Highlights
Executive-level opportunity to lead a global infrastructure and reliability organization, shape AI-enabled operational practices, remain technically hands-on, and influence reliability strategy for enterprise-scale systems.
Description
We're building an AI-native Infrastructure & Reliability organization where autonomous agents investigate incidents, generate RCAs, validate changes, and safely remediate production issues.
We're looking for an SVP of Infrastructure & Reliability to lead this transformation.
You'll own platform reliability for enterprise customers while building the AI operating model that defines how modern infrastructure teams work.
What You'll Do
Own uptime, reliability, MTTR, and customer experience across a large-scale SaaS platform.Lead a senior global team of Infrastructure, Platform, and Reliability engineers.Design and continuously improve AI agents for incident response, auto-remediation, change validation, and operational automation.Stay hands-on by writing code, improving AI agents, and leading major customer incidents.Partner directly with enterprise customers, engineering leadership, and the CEO during critical situations.
What We're Looking For
10+ years leading Infrastructure, Platform Engineering, DevOps, Cloud Operations, or Reliability organizations.3+ years as an SVP, VP, or Head of Engineering.Proven experience operating enterprise SaaS platforms on AWS.Hands-on technical leader who still writes production code.Experience building or leading AI-powered operations, AIOps, or autonomous infrastructure.Daily user of modern AI engineering tools such as Claude Code, Cursor, or similar.Strong executive communication skills and advanced English.
Why Join Us
You'll help build one of the first AI-native Infrastructure & Reliability organizations at enterprise scale.
Instead of managing growing operations teams, you'll create AI systems that eliminate operational work while improving reliability for some of the world's largest brands.
Fully remoteAI-first cultureNo limits on AI toolingGlobal teamEnterprise scale