Site Reliability Engineer

Asobbi โ€” United States ยท Posted ~2 days ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

SENIOR SITE RELIABILITY ENGINEER | FULLY REMOTE (US) Location: Fully remote - US-based, working a US timezone Package: $170,000 - $220,000 Overview We're supporting a specialist AI infrastructure company - a UK sovereign AI cloud powered by renewable energy - that builds and operates large-scale compute on regenerated industrial and energy sites. Having recently secured a major US customer and taken on their entire cluster, the business is standing up a US-based operations team and is looking for Senior Site Reliability Engineers to build that capability from the groundup. This is a Platform/SRE role with a strong automation and software-engineering bias โ€” not an AI model-building role. You'll turn runbooks, alerts and operational workflows into safe, auditable automation that improves reliability across the platform. Why join? Get in at the ground floor of a brand-new US operations function and shape how it runs at scaleWork on critical infrastructure powering the next generation of AIAutomation-first culture - reduce toil and build tooling, rather than fire fightReal autonomy and high visibility with leadershipFully remote, on a US timezone (West-coast preferred) What you'll be doing Building Python-based automation for incident triage, runbook execution and routine operational tasksIntegrating observability, ITSM and infrastructure APIs to enrich alerts and automate workflowsImproving monitoring signal quality through correlation, enrichment, suppression and deduplicationBuilding internal tools and self-service capabilities - CLI utilities, ChatOps integrations and dashboardsMaintaining version-controlled runbook-as-code and automation librariesTurning post-incident learnings into better tooling, automation and operational standards We're keen to speak with candidates who have Essential: Experience in SRE, Platform Engineering or production infrastructure operationsHands-on experience with observability/monitoring tooling (Prometheus, Grafana or similar)Exposure to incident management / on-call, and converting manual runbooks into automationStrong Python for automation, APIs and integrations Nice to have: GPU, datacentre or colocation infrastructure experienceITSM integrations (ServiceNow, Halo, Jira Service Management or similar)ChatOps tooling (Slack or Microsoft Teams bots)OpenTelemetry, logging or distributed tracing experienceDCIM, IPAM or hypervisor-control-plane integrationsExperience with LLM-assisted or agent-based operational automation Next steps Interested? Apply directly or message me for a confidential discussion.