Site Reliability Engineer
Asobbi โ United States ยท Posted ~2 days ago
๐ Log in to save this job, tailor your resume & track your apply process โ 7 days free, no card needed.
Log in to add to target listDescription
SENIOR SITE RELIABILITY ENGINEER | FULLY REMOTE (US)
Location: Fully remote - US-based, working a US timezone
Package: $170,000 - $220,000
Overview
We're supporting a specialist AI infrastructure company - a UK sovereign AI cloud powered by renewable energy - that builds and operates large-scale compute on regenerated industrial and energy sites.
Having recently secured a major US customer and taken on their entire cluster, the business is standing up a US-based operations team and is looking for Senior Site Reliability Engineers to build that capability from the groundup.
This is a Platform/SRE role with a strong automation and software-engineering bias โ not an AI model-building role.
You'll turn runbooks, alerts and operational workflows into safe, auditable automation that improves reliability across the platform.
Why join?
Get in at the ground floor of a brand-new US operations function and shape how it runs at scaleWork on critical infrastructure powering the next generation of AIAutomation-first culture - reduce toil and build tooling, rather than fire fightReal autonomy and high visibility with leadershipFully remote, on a US timezone (West-coast preferred)
What you'll be doing
Building Python-based automation for incident triage, runbook execution and routine operational tasksIntegrating observability, ITSM and infrastructure APIs to enrich alerts and automate workflowsImproving monitoring signal quality through correlation, enrichment, suppression and deduplicationBuilding internal tools and self-service capabilities - CLI utilities, ChatOps integrations and dashboardsMaintaining version-controlled runbook-as-code and automation librariesTurning post-incident learnings into better tooling, automation and operational standards
We're keen to speak with candidates who have
Essential:
Experience in SRE, Platform Engineering or production infrastructure operationsHands-on experience with observability/monitoring tooling (Prometheus, Grafana or similar)Exposure to incident management / on-call, and converting manual runbooks into automationStrong Python for automation, APIs and integrations
Nice to have:
GPU, datacentre or colocation infrastructure experienceITSM integrations (ServiceNow, Halo, Jira Service Management or similar)ChatOps tooling (Slack or Microsoft Teams bots)OpenTelemetry, logging or distributed tracing experienceDCIM, IPAM or hypervisor-control-plane integrationsExperience with LLM-assisted or agent-based operational automation
Next steps
Interested? Apply directly or message me for a confidential discussion.
We have 61,299 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume โ in under a minute we'll analyze all 61,299 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume