Site Reliability Engineer

Asobbi — United Kingdom · Posted ~2 hours ago

Mid Full-time Remote

Skills

site reliability engineering Python automation incident management monitoring infrastructure APIs monitoring tools ITSM systems

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking an SRE focused on automation and platform reliability. The role involves building internal tools, improving monitoring and incident response processes, and reducing operational effort through software engineering practices.

Highlights

Remote engineering role focused on automation, reliability improvement, and building scalable operational tooling for advanced infrastructure environments.

Description

Site Reliability Engineer (Platform / Automation) Remote, UK About the Role A UK-based AI Sovergeign Cloud is looking for SRE and Platform engineers who default to automation. The successful candidate will help build the operations function from the ground up; turning runbooks, alerts, and manual workflows into safe, auditable automation and internal tooling that improves reliability across AI infrastructure and data-centre platforms. This is a platform and SRE role with a strong software engineering focus. It is not an AI model-building position. The engineer will work closely with operations, platform, and engineering teams to cut manual toil, improve alert quality, and shorten incident response times. Key Responsibilities Write Python-based automation for incident triage, runbook execution, and day-to-day operational tasks.Connect observability, ITSM, and infrastructure APIs to enrich alerts and automate workflows.Improve monitoring signal quality through correlation, enrichment, suppression, and deduplication.Build internal tools and self-service capabilities—CLI utilities, ChatOps integrations, and dashboards.Maintain version-controlled runbook-as-code and reusable automation libraries.Turn post-incident learnings into better tooling, automation, and operational standards.Support safe, auditable automation for higher-risk actions with appropriate approval controls. Essential Experience Experience in SRE, Platform Engineering, or production infrastructure operations.GPU, data-centre, or colocation infrastructure experience.Hands-on work with observability and monitoring tools (e.g., Prometheus, Grafana, or similar).Exposure to incident management, on-call, and converting manual runbooks into automation.Strong Python skills for automation, APIs, and integrations. Nice to Have ITSM integrations (ServiceNow, Halo, Jira Service Management, or similar).ChatOps tooling (Slack or Microsoft Teams bots).OpenTelemetry, logging, or distributed tracing experience.DCIM, IPAM, or hypervisor-control-plane integrations.Experience with LLM-assisted or agent-based operational automation. Why Join The organisation is a mission-driven start-up building critical national infrastructure, where operational excellence directly fuels growth. This role offers high visibility with leadership, real autonomy, and the opportunity to shape how a next-generation company operates at scale. Diversity & Inclusion The organisation is an equal opportunity employer. It celebrates diversity and is committed to creating an inclusive environment for all employees.