Platform DevOps / Site Reliability Engineer

Altrapoint โ€” United Kingdom ยท Posted ~1 day ago

Senior Contract Hybrid No Visa

Skills

DevOps Site reliability engineering Kubernetes Public cloud Private cloud Incident management Infrastructure automation Platform operations Diagnostic tooling SRE APIs

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

An experienced Platform/DevOps/SRE engineer is needed for a long-term enterprise project supporting a high-throughput API platform. You will operate Kubernetes across public and private clouds while actively building automation, diagnostic tooling, incident-response capabilities, and engineering improvements that keep the platform stable and performant.

Highlights

Six-month contract with likely extension, working on a high-throughput enterprise API platform. The role combines engineering and operations, with opportunities to automate processes, build diagnostic tooling, improve reliability, and work across public and private cloud environments.

Description

Position: Platform / DevOps / Site Reliability Engineer Contract Duration: 6 Months (likely extension to 12 months) IR35 Status: Outside IR35 Location: Hybrid (Glasgow or London) Right to Work: Candidates must have full existing UK right to work (no sponsorship offered) About the Role: We are seeking an experienced Platform / DevOps / SRE Engineer to join a financial services consultancy on an exciting, long-term project supporting a high-throughput, enterprise API platform. This is a true engineering-meets-operations role. Beyond maintaining system health and reliability across public and private cloud environments, you will actively build diagnostic tooling, automate manual processes, and deliver practical engineering improvements to ensure the platform remains stable and performant. Key Responsibilities: Platform Operations: Maintain and manage a Kubernetes-based platform operating across public and private cloud environments.Incident Management: Drive end-to-end incident response, including root-cause analysis and post-incident reviews.Automation & Performance: Build automation, Helm charts, and custom diagnostic tooling to eliminate toil and evaluate performance.Observability: Work with observability tools to ensure monitoring alerts are fully actionable through clear, up-to-date runbooks.Capacity & Upgrades: Proactively manage capacity scaling and oversee regular software and infrastructure component upgrades. Key Requirements & Skills: Kubernetes in Production: Proven hands-on experience operating production-grade Kubernetes (Service Mesh experience is a strong plus).Linux Fundamentals: Strong Linux administration and command-line troubleshooting skills.Cloud Infrastructure: Practical experience with major cloud providers, preferably AWS or Azure.Problem Solving: Confident debugging complex systems across all layers, from application code to underlying infrastructure.Observability (Plus): Working knowledge of Grafana, Prometheus, Loki, or Tempo.Automation & Scripting (Plus): Hands-on experience with Infrastructure as Code (Terraform, Helm), CI/CD pipelines, or scripting in Python/Java. Note: Prior financial services experience is not required.