Principal Site Reliability Engineer

Hunter Bond — United Kingdom · Posted ~1 day ago

Lead Full-time Hybrid

Skills

Site Reliability Engineering DevOps Linux Python Go Distributed systems Production infrastructure Kubernetes Networking CI/CD Monitoring Observability Infrastructure automation Problem solving On-premises infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary

Take a principal-level role in a high-performing infrastructure engineering team responsible for highly available trading systems. You will automate operations, improve monitoring and observability, troubleshoot demanding production and performance issues, and partner with software engineers to build resilient distributed infrastructure at scale. The environment emphasizes technical ownership, rapid innovation, and meaningful engineering impact.

Highlights

Principal-level SRE opportunity with exceptional compensation, flexible working options, significant technical ownership, and career progression. The role provides exposure to large-scale, business-critical infrastructure, complex reliability challenges, greenfield projects, and an engineering-led environment with minimal bureaucracy.

Description

Job Title: Prinicpal Site Reliability Engineer Sector: Fintech Location: London (Hybrid) Salary: Up to £250k Base + market-leading bonus + outstanding benefits My client is looking for a hands-on Site Reliability Engineer to join a high-performing infrastructure engineering team responsible for the reliability, automation and performance of critical trading platforms. You'll work across Linux, distributed systems and production infrastructure, partnering with software engineers to improve resilience, automate operations and solve complex technical challenges at scale. Role Build and improve highly available production infrastructureAutomate operational processes and eliminate manual tasksWorking on complex production and performance issuesImprove monitoring, observability and platform reliabilityWork closely with engineering teams to deliver resilient applications Skills / Experience 5+ years of experience in an SRE/DevOps/Production roleStrong Linux systems experienceGood Python or Go development skills for automation and toolingExperience supporting distributed systems in productionKnowledge of Kubernetes, networking and CI/CD in an on prem environmentStrong understanding of monitoring and observability platformsExcellent problem-solving skills with an engineering-first mindset Why Apply? Work on large-scale, business-critical infrastructure at a scale thought often impossible!Join an engineering-led environment with genuine technical ownershipSolve complex reliability and performance challenges every dayOutstanding compensation and career progression Sells Flexible hours/work optionsThe chance to work with the most cutting-edge technology in the IndustryTechnologists only report to technologistsTechies treated as top commodity so ‘spoiled’Small team size in a growing company so an ability to make a significant impactConstantly exciting greenfield projects in an ever-evolving environmentNo red tapeBeautiful offices