Site Reliability Engineer

Selby Jennings — United Kingdom · Posted ~1 hour ago

Mid Full-time Visa History ✓

Skills

Site reliability engineering Cloud technologies Containerization Automation Infrastructure Cloud Containers

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A site reliability engineering role focused on maintaining scalable infrastructure, improving system performance, automating operations, and collaborating with technical teams.

Highlights

Work on mission-critical infrastructure with high ownership, modern technologies, and direct business impact.

Description

Site Reliability Engineer Our client is a world-renowned quantitative hedge fund where technology sits at the core of the business. They are seeking a Site Reliability Engineer to join a highly visible platform team responsible for the reliability, performance, and scalability of the infrastructure that powers cutting-edge research and trading. This is a unique opportunity for a Site Reliability Engineer to work directly alongside traders, quantitative researchers, and engineering teams in an environment where your work has immediate business impact. As a Site Reliability Engineer, you'll gain exceptional ownership from day one, helping to support and evolve a sophisticated technology estate built on modern cloud and containerisation technologies. The role combines troubleshooting, automation, platform engineering, and stakeholder interaction, making it ideal for someone looking to accelerate their career within a top 1% investment firm. The fund offers top-of-market compensation, strong bonus potential, and annual compensation increases, alongside the opportunity to work with some of the industry's brightest engineers and researchers. Key Responsibilities Ensure the reliability, availability, and performance of critical research and trading platformsAct as the first point of contact for infrastructure and platform-related issuesInvestigate and resolve production incidents, minimising business impactPartner closely with traders, quants, and engineers to support business-critical systemsDrive improvements to monitoring, automation, operational processes, and platform toolingManage onboarding, permissions, access requests, and platform support activities Key Skills & Experience 3-8 years' experience in Site Reliability Engineering, Platform Engineering, DevOps, Infrastructure Engineering, or Production EngineeringStrong experience with AWS and Kubernetes in production environmentsExcellent Linux administration and troubleshooting skillsExperience with Docker, CI/CD pipelines, infrastructure automation, and cloud-native technologiesKnowledge of observability and monitoring tools such as Datadog, Prometheus, Grafana, CloudWatch, or ELKScripting experience with Python, Bash, or similar languagesStrong communication skills, a proactive mindset, and a genuine sense of ownershipExperience within financial services, trading, or other high-performance environments is advantageous