Site Reliability Engineer

Darktrace — United Kingdom · Posted ~1 day ago

Senior Full-time Visa History ✓

Skills

SRE platform engineering DevSecOps cloud infrastructure reliability engineering cloud

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a technology team focused on building reliable platforms, improving operational processes, and applying engineering practices to maintain secure and scalable systems.

Highlights

Work on large-scale reliability challenges with opportunities to influence platform strategy and improve operational excellence.

Description

Darktrace is a global leader in AI for cybersecurity that keeps organizations ahead of the changing threat landscape every day. Founded in 2013, Darktrace provides the essential cybersecurity platform protecting nearly 10,000 organizations from unknown threats using its proprietary AI. The Darktrace Active AI Security Platform™ delivers a proactive approach to cyber resilience to secure the business across the entire digital estate – from network to cloud to email. Breakthrough innovations from our R&D teams have resulted in over 200 patent applications filed. Darktrace’s platform and services are supported by over 2,400 employees around the world. To learn more, visit http://www.darktrace.com. Job Description: About The Role We’re looking for a Site Reliability Engineer (SRE) to bring deep expertise in a key reliability domain and help shape the future of our platform reliability strategy. SRE sits at the heart of our operational trifecta alongside Platform Engineering and DevSecOps. In this role, you’ll act as the go-to authority in your area of specialism, working across teams to embed best practices, solve complex reliability challenges, and improve system resilience at scale. Unlike a generalist SRE, this role focuses on a core domain of expertise—such as observability, performance engineering, data infrastructure reliability, security-focused SRE, or network reliability—while influencing reliability standards across the wider engineering organisation. Key Responsibilities Domain Expertise & Strategy Act as the subject matter expert in your chosen reliability domainDefine and implement standards, frameworks, and best practices across SRE, Platform Engineering, and DevSecOpsStay current with industry trends and bring innovative ideas into the organisation Engineering & Delivery Design and implement solutions to complex, cross-cutting reliability challengesBuild tooling, automation, and frameworks to improve system resilience and scalabilityLead deep-dive investigations into systemic issues and drive long-term fixes Collaboration & Platform Integration Partner with Platform Engineering to ensure your domain is embedded within the internal developer platformCollaborate with DevSecOps to integrate security, compliance, and resilience practicesContribute to cross-team initiatives that improve reliability across the stack Incident & Operational Excellence Play a key role in incident response, particularly within your specialismContribute to on-call rotations and continuous improvement of operational processesDevelop runbooks, documentation, and training materials to support teams Essential What You’ll Bring Proven experience in Site Reliability Engineering, DevOps, or infrastructure engineeringDeep expertise in at least one of the following areas:Observability & monitoring (metrics, logging, distributed tracing)Performance engineering & capacity planningData infrastructure reliability (databases, streaming, pipelines)Security-focused SRE (hardening, compliance automation, secrets management)Network reliability & traffic managementStrong programming skills (e.g. Go, Python, or similar)Experience with cloud platforms (AWS, GCP, Azure) and KubernetesStrong communication skills, with the ability to explain complex technical concepts clearlySelf-driven with the ability to identify and prioritise high-impact work independently Desirable Experience building internal developer platforms or toolingContributions to open-source, technical blogs, or public speakingExperience working in regulated environmentsFamiliarity with SLO frameworks and error budget managementRelevant certifications in your specialist domain Success Measures Improved reliability and performance within your domain of specialismAdoption of best practices across SRE, Platform Engineering, and DevSecOpsReduction in incidents and faster resolution timesScalable, well-integrated solutions within the internal platformStrong collaboration across teams and measurable improvements in operational maturity Why Join Us? Shape reliability strategy in a modern, cloud-native engineering environmentWork on complex, high-impact systems at scaleCollaborate with expert teams across Platform Engineering and DevSecOpsTake ownership of a domain and drive meaningful, organisation-wide impact Benefits: 23 days’ holiday + all public holidays, rising to 25 days after 2 years of service,Additional day off for your birthday,Private medical insurance which covers you, your cohabiting partner and children,Life insurance of 4 times your base salary,Salary sacrifice pension scheme,Enhanced family leave,Confidential Employee Assistance Program,Cycle to work scheme. Darktrace is an Equal Opportunity Employer. We consider all qualified applicants for employment without regard to race, color, religion, sex (including pregnancy, childbirth, and related medical conditions), sexual orientation, gender identity or expression, national origin, age, disability, genetic information, marital status, veteran or military status, or any other characteristic protected by applicable federal, state, or local law. Darktrace is committed to providing reasonable accommodations to qualified individuals with disabilities in accordance with applicable laws. If you require a reasonable accommodation to participate in the application or interview process, please contact your Talent Partner.