Infrastructure / SRE Engineer

Cavendish Professionals — United Kingdom · Posted ~7 hours ago

Senior Full-time Hybrid

Skills

Infrastructure Engineering SRE Kubernetes Docker Distributed Systems CI/CD Automation Monitoring Telemetry Production Operations

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A high-caliber Infrastructure / SRE Engineer is sought to design, build, and operate mission-critical distributed production platforms. You will own architecture and operations end to end, build resilient Docker and Kubernetes environments, automate deployment workflows, develop internal tooling, and improve monitoring, telemetry, availability, and incident response. The role offers substantial autonomy from day one.

Highlights

High-autonomy infrastructure role with end-to-end architectural and operational ownership of mission-critical distributed platforms. The position offers the chance to solve complex national-scale engineering problems and work with modern container, automation, CI/CD, and observability technologies.

Description

🚀 Infrastructure / SRE Engineer | Kubernetes & Distributed Systems | Secure National Infrastructure 🇬🇧 Hybrid (London) | 📄 Permanent, Full-Time | We are seeking high-caliber Infrastructure / SRE Engineers to design, build, and operate mission-critical production infrastructure supporting vital UK government and national security missions. This is an environment operating at startup velocity to solve complex, national-scale engineering problems with true architectural ownership and autonomy from day one. What You’ll Be Doing: Taking end-to-end operational and architectural ownership of high-security, distributed production platforms.Architecting, deploying, and maintaining resilient container runtimes using Docker and Kubernetes.Designing automated CI/CD deployment workflows and building internal tooling to optimize developer velocity and platform resilience.Developing automation, monitoring, and telemetry infrastructure to ensure ultra-low downtime and rapid incident response across complex distributed systems.Driving platform innovation and automation initiatives with modern tooling, cutting through operational bureaucracy. What We’re Looking For: 1–5 years of commercial experience in Infrastructure Engineering, Site Reliability Engineering (SRE), or Platform Operations.Software-minded engineering approach: strong hands-on experience building, automating, and deploying production architectures (not just operational support/debugging).Solid programming background in Python, Go, Java, or JavaScript.Demonstrated ownership mindset: comfort working in rapid iterations, driving solutions independently, and operating in secure, high-trust environments.Degree in Computer Science, Computer Engineering, or related discipline from a top-50 university. Must-Haves: Active UK Developed Vetting (DV) security clearance.Sole British nationality (dual nationals cannot be considered due to national security compliance).Hands-on expertise with Docker, Kubernetes, and container orchestration in production.Solid coding/scripting capability in Python, Go, Java, or JavaScript.Production Linux systems engineering and cloud/distributed systems delivery (AWS, GCP, Azure).Residency or willingness to relocate within 30 minutes of the secure London site. Location & Working Setup: Hybrid model based out of the London If you are an infrastructure specialist ready to take direct ownership of mission-critical systems where your engineering impact truly matters every day, apply now or connect directly to discuss the opportunity.