DevOps / Site Reliability Engineer

Altrapoint โ€” United Kingdom ยท Posted ~1 day ago

Senior Contract Hybrid No Visa

Skills

DevOps Site Reliability Engineering Kubernetes Public cloud Private cloud Platform engineering Incident management Automation Diagnostic tooling Enterprise API platforms System reliability Performance engineering

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

An experienced platform and reliability engineering position supporting a high-throughput enterprise API environment. The role combines cloud operations with hands-on engineering, including Kubernetes management, incident response, diagnostic tooling, automation, and performance improvements.

Highlights

A hands-on platform engineering role with a six-month contract likely to extend to twelve months. Work on a high-throughput enterprise platform across public and private cloud environments, with strong emphasis on automation, reliability, and engineering improvements.

Description

Position: Platform / DevOps / Site Reliability Engineer Contract Duration: 6 Months (likely extension to 12 months) IR35 Status: Outside IR35 Location: Hybrid (London or Glasgow) Right to Work: Candidates must have full existing UK right to work (no sponsorship offered) About the Role: We are seeking an experienced Platform / DevOps / SRE Engineer to join a financial services consultancy on an exciting, long-term project supporting a high-throughput, enterprise API platform. This is a true engineering-meets-operations role. Beyond maintaining system health and reliability across public and private cloud environments, you will actively build diagnostic tooling, automate manual processes, and deliver practical engineering improvements to ensure the platform remains stable and performant. Key Responsibilities: Platform Operations: Maintain and manage a Kubernetes-based platform operating across public and private cloud environments.Incident Management: Drive end-to-end incident response, including root-cause analysis and post-incident reviews.Automation & Performance: Build automation, Helm charts, and custom diagnostic tooling to eliminate toil and evaluate performance.Observability: Work with observability tools to ensure monitoring alerts are fully actionable through clear, up-to-date runbooks.Capacity & Upgrades: Proactively manage capacity scaling and oversee regular software and infrastructure component upgrades. Key Requirements & Skills: Kubernetes in Production: Proven hands-on experience operating production-grade Kubernetes (Service Mesh experience is a strong plus).Linux Fundamentals: Strong Linux administration and command-line troubleshooting skills.Cloud Infrastructure: Practical experience with major cloud providers, preferably AWS or Azure.Problem Solving: Confident debugging complex systems across all layers, from application code to underlying infrastructure.Observability (Plus): Working knowledge of Grafana, Prometheus, Loki, or Tempo.Automation & Scripting (Plus): Hands-on experience with Infrastructure as Code (Terraform, Helm), CI/CD pipelines, or scripting in Python/Java. Note: Prior financial services experience is not required.