Summary
✨ AI‑Generated
Lead a small, hands-on SRE team responsible for keeping a large-scale platform reliable and secure around the clock. You will define technical direction, own reliability and security strategy, mentor engineers and remain deeply involved with infrastructure. The environment is remote-first and emphasizes root-cause fixes, transparency and continuous improvement.
Highlights
Lead a small SRE team while remaining hands-on, set technical direction and own reliability and security strategy for a 24/7 platform. The remote-first environment values impact, transparency and continuous improvement, with a strong focus on solving root causes.
Description
Lead Site Reliability Engineer (SRE)
Ocho are working with a client to find a Lead Site Reliability Engineer (SRE) to lead the team responsible for keeping their platform running reliably and securely, 24/7.
Our client helps thousands of teams monitor and improve their applications, and is remote-first, valuing impact, transparency and continuous improvement.
The role
This is a player-coach position.
You'll set the technical direction and own reliability and security strategy for the platform, while staying hands-on with the systems your team runs.
It's a small team with a long-standing habit of fixing root causes, not just alerts, and they're now growing it as the business scales.
Their stack
Mostly bare-metal infrastructure, managed by AnsibleData ingestion and processing in Rust, running on KafkaA Rails app serving the customer-facing UIMongoDB, ClickHouse and ElasticSearch
Responsibilities
Lead the SRE team: set priorities, mentor engineers, grow the teamOwn reliability strategy and the long-term infrastructure roadmapBe part of the on-call rotation, and keep improving itAct as incident coordinator, and lead blameless postmortemsGuide strategic projects, including new AWS infrastructureStay hands-on: tune the Rust codebase and infrastructure automationHandle security researcher reports, coordinate penetration tests, support ISO renewals
What you bring
8+ years keeping large Linux systems reliable, with experience leading an SRE, platform or infrastructure team (formally or as a technical lead).Competent developer across multiple languages, ideally with Rust and Ansible experience.Strong incident response and postmortem experience, comfortable translating business growth into infrastructure strategy.Bonus: AWS, Kubernetes and Docker.
What's on offer
Competitive salary, remote-first culture, stock options, flexible PTO, and a personal development budget.
Please apply now if you are meeting the above criteria or contact Andrew Harrison directly.