Site Reliability Engineer

Ixceed Solutions — United Kingdom · Posted ~4 hours ago

Mid Full-time

Skills

Kubernetes OpenShift Linux incident response automation monitoring monitoring tools

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A reliability engineering position responsible for operating and improving container platforms, automating infrastructure tasks, responding to incidents, and enhancing production stability.

Highlights

Role focused on improving platform reliability, automation, observability, and operational excellence in a collaborative engineering environment.

Description

Key Responsibilities Own the reliability, availability, performance, and operational health of Kubernetes/OpenShift platforms.Participate in incident response, on-call support, troubleshooting, RCA, and post-incident improvements.Support platform upgrades, component releases, infrastructure changes, and production deployments.Drive automation to reduce operational toil and improve scalability and resilience.Develop and maintain scripts, tooling, dashboards, runbooks, and operational documentation.Manage cluster lifecycle activities, including upgrades, capacity planning, performance optimisation, and production readiness.Lead vulnerability management, CVE assessment, prioritisation, patching, and remediation across Kubernetes, OpenShift, Linux, and supporting components.Improve platform observability, monitoring, metrics, logging, and alerting.Collaborate with global SRE, Infrastructure, Security, and L2 Operations teams.Contribute to platform modernisation, self-service, automation, reliability, and customer experience initiatives. Required Skills & Experience Experience in SRE, Platform Engineering, DevOps, Systems Engineering, or Infrastructure Operations.Strong hands-on experience with Red Hat OpenShift or enterprise Kubernetes in production.Strong Linux administration and troubleshooting skills.Deep understanding of Kubernetes architecture, networking, scheduling, storage, container orchestration, and cluster lifecycle management.Strong understanding of SRE principles, incident management, observability, automation, capacity planning, reliability, and continuous improvement.Excellent communication, collaboration, and problem-solving skills