Lead Site Reliability Engineer

Ocho People — United Kingdom · Posted ~2 hours ago

Lead Remote

Skills

Site reliability engineering Infrastructure management Ansible Reliability engineering Security Team leadership Mentoring Root-cause analysis Rust Kafka Ruby on Rails MongoDB ClickHouse Elasticsearch Bare-metal infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Lead a small, hands-on SRE team responsible for keeping a large-scale platform reliable and secure around the clock. You will define technical direction, own reliability and security strategy, mentor engineers and remain deeply involved with infrastructure. The environment is remote-first and emphasizes root-cause fixes, transparency and continuous improvement.

Highlights

Lead a small SRE team while remaining hands-on, set technical direction and own reliability and security strategy for a 24/7 platform. The remote-first environment values impact, transparency and continuous improvement, with a strong focus on solving root causes.

Description

Lead Site Reliability Engineer (SRE) Ocho are working with a client to find a Lead Site Reliability Engineer (SRE) to lead the team responsible for keeping their platform running reliably and securely, 24/7. Our client helps thousands of teams monitor and improve their applications, and is remote-first, valuing impact, transparency and continuous improvement. The role This is a player-coach position. You'll set the technical direction and own reliability and security strategy for the platform, while staying hands-on with the systems your team runs. It's a small team with a long-standing habit of fixing root causes, not just alerts, and they're now growing it as the business scales. Their stack Mostly bare-metal infrastructure, managed by AnsibleData ingestion and processing in Rust, running on KafkaA Rails app serving the customer-facing UIMongoDB, ClickHouse and ElasticSearch Responsibilities Lead the SRE team: set priorities, mentor engineers, grow the teamOwn reliability strategy and the long-term infrastructure roadmapBe part of the on-call rotation, and keep improving itAct as incident coordinator, and lead blameless postmortemsGuide strategic projects, including new AWS infrastructureStay hands-on: tune the Rust codebase and infrastructure automationHandle security researcher reports, coordinate penetration tests, support ISO renewals What you bring 8+ years keeping large Linux systems reliable, with experience leading an SRE, platform or infrastructure team (formally or as a technical lead).Competent developer across multiple languages, ideally with Rust and Ansible experience.Strong incident response and postmortem experience, comfortable translating business growth into infrastructure strategy.Bonus: AWS, Kubernetes and Docker. What's on offer Competitive salary, remote-first culture, stock options, flexible PTO, and a personal development budget. Please apply now if you are meeting the above criteria or contact Andrew Harrison directly.