Site Reliability Engineer

Mthree — United Kingdom · Posted ~21 hours ago

Skills

Site reliability engineering Observability Monitoring Telemetry Incident response Production support Performance troubleshooting Automation Platform reliability Scalability Dashboards SRE

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a global technology environment as a Site Reliability Engineer responsible for enterprise-scale observability and monitoring platforms. You will investigate production incidents, improve reliability and scalability, deliver controlled releases, build automation and dashboards, and partner with engineering teams to strengthen operational resilience.

Highlights

Opportunity to work on enterprise-scale observability and reliability platforms, influence platform decisions, improve resilience and scalability, handle production incidents, and collaborate with engineering teams across a global technology environment.

Description

Global Investment Bank | Observability SRE | London Join a leading investment bank as an Site Reliability Engineer (SRE), where you'll play a key role in shaping the future of monitoring, telemetry, and platform reliability across a global technology estate. This is a fantastic opportunity to work with cutting-edge observability technologies, influence strategic platform decisions, and collaborate with engineering teams worldwide to deliver highly resilient, scalable, and business-critical services. What You'll Do Own and support enterprise-scale observability and monitoring platformsInvestigate and resolve production incidents, alerts, and performance issuesDrive reliability, scalability, and continuous improvement across critical platformsDeliver production releases and changes in a controlled, low-risk mannerBuild and enhance dashboards, automation, and monitoring capabilitiesPartner with global engineering teams to improve operational efficiency and resilienceChampion observability best practices and help shape the future monitoring strategy What We're Looking For Experience with Grafana or other modern observability and monitoring platformsStrong Linux administration and troubleshooting skillsExperience with Python and/or AnsibleBackground supporting production environments, including incident, problem, and change managementExposure to cloud technologies and CI/CD tooling such as GitLab, Jenkins, or AnsibleStrong communication skills with the ability to engage technical and non-technical stakeholders Nice to Have OpenTelemetry knowledgeKubernetes, Docker, or EKS experienceExperience supporting large-scale enterprise environmentsITIL awarenessSQL/database knowledgeExperience working within global teams