Site Reliability Engineer

Teksystems โ€” United Kingdom ยท Posted ~2 days ago

Senior

Skills

Site Reliability Engineering Cloud-native applications Python, Java, or Go AWS Kubernetes Terraform Observability Incident management On-call operations Infrastructure as Code Docker Datadog Python Java Go GitHub Actions Jenkins

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

Join a high-performing engineering team responsible for improving the reliability, scalability, and performance of large-scale cloud-native applications. You will build observability capabilities, automate operations, manage infrastructure as code, support production systems, and collaborate closely with software engineers. The role requires at least four years of relevant experience and strong skills in cloud platforms, containers, infrastructure automation, incident response, and Python, Java, or Go.

Highlights

Own reliability initiatives for large-scale customer-facing systems, work with modern cloud-native technologies, and improve observability, automation, performance, and operational resilience.

Description

We're partnering with a globally recognised financial technology company that powers millions of customer transactions every day. As they continue investing in platform reliability, cloud infrastructure, and engineering excellence, they're looking for a Site Reliability Engineer to join a high-performing team working on large-scale, customer-facing systems. The Role As a Site Reliability Engineer, you'll help improve the reliability, scalability, and performance of cloud-native applications. Working closely with software engineers, you'll drive observability best practices, automate operational processes, and play a key role in maintaining highly available production systems. Key Responsibilities Support production systems through incident response, on-call rotations, and post-incident reviews.Define and improve monitoring, logging, alerting, and distributed tracing capabilities.Drive adoption of SLOs, SLIs, SLAs, and error budgets.Build automation to reduce toil and improve developer productivity.Investigate performance issues, improve platform reliability, and support capacity planning.Manage infrastructure using Terraform and Infrastructure as Code principles. What We're Looking For 4+ years' experience as a Software Engineer, SRE, or Production Engineer.Experience building or supporting cloud-native applications.Strong experience with Python, Java, or Go.Hands-on knowledge of AWS, Kubernetes, Terraform, and observability tooling.Experience participating in on-call rotations and incident management. Technology Stack AWS | Kubernetes | Docker | Terraform | Datadog | Python | Java | Go | GitHub Actions | Jenkins Why Apply? This is an opportunity to join a business where reliability is a core engineering function, not a support team. You'll work on complex, large-scale systems, have genuine ownership of reliability initiatives, and collaborate with talented engineers in a modern cloud-native environment.