Site Reliability Engineer

Caspianone — Poland · Posted ~9 hours ago

Senior

Skills

Site reliability engineering Production engineering Production application support Incident management Root cause analysis Performance troubleshooting Automation Observability Platform reliability Resilience engineering Problem-solving Site Reliability Engineering Observability Platforms Production Monitoring Incident Management Distributed Systems

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A global financial institution is seeking an experienced Site Reliability Engineer to improve the stability, performance, and resilience of business-critical production platforms. You will investigate incidents, perform root cause analysis, build automation, expand observability capabilities, and implement lasting fixes across distributed production engineering teams.

Highlights

Help maintain resilient, business-critical trading platforms in a globally distributed engineering team. Focus on automation, advanced observability, root cause analysis, and permanent reliability improvements rather than repetitive manual support. Contribute directly to the performance and stability of critical financial systems.

Description

We're supporting a global banking organisation as they continue to invest heavily in their Markets Technology function under a new technology leadership team. As part of this growth, they're looking for an experienced Site Reliability Engineer to join a globally distributed production engineering team responsible for the stability, performance and resilience of business-critical trading and markets platforms. This isn't a traditional support role. The team is looking for engineers who enjoy solving problems at scale, automating repetitive processes and improving platform reliability rather than simply firefighting production issues. What You'll Be Doing Supporting critical production applications used across global market businesses.Building out new capabilities for a Large Observability platformInvestigating and resolving complex incidents and performance issues.Performing root cause analysis and implementing permanent fixes.Building automation to reduce manual operational tasks.Enhancing monitoring, observability and alerting capabilities.Working closely with development and infrastructure teams to improve platform resilience.Driving continuous improvements across the production estate.Contributing to on-call and major incident support when required. What They're Looking For Experience working as an SRE, Production Engineer, Platform Engineer or similar.Strong Linux/Unix administration skills.Good scripting and automation experience, ideally with Python or Shell.Experience with monitoring and observability tooling such as Prometheus/Grafana.Proven experience supporting large-scale, business-critical systems.Strong troubleshooting and problem-solving skills.An engineering-first mindset with a genuine interest in reliability, automation and continuous improvement. Nice to Have Financial Services or Capital Markets experience.Cloud experience (AWS, Azure or GCP).DevOps, CI/CD and infrastructure automation experience. Why Consider It? Join a technology organisation going through significant investment and transformation.Work in a genuinely global environment alongside teams across Europe and Asia.Play a key role in shaping SRE and production engineering best practices.Gain exposure to complex, high-availability systems operating at enterprise scale.Strong opportunity for long-term growth as the team continues to expand. Financial services experience is helpful, but it isn't a prerequisite. We're equally interested in engineers from large-scale technology environments who can demonstrate a strong background in reliability engineering, automation and production support.