Site Reliability Engineer

Hirefeedd — United Kingdom · Posted ~4 hours ago

Senior Full-time Remote $130000-$200000 USD equivalent

Skills

Site Reliability Engineering DevOps Infrastructure engineering High availability Performance engineering Observability Logging Metrics Tracing SLOs SLIs Error budgets Incident response Post-mortem analysis Automation Scripting Software engineering SRE SLO SLI

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A fully remote Site Reliability Engineer is sought to keep production systems highly available, observable, and performant. You will build observability capabilities across logs, metrics, and traces; define SLOs, SLIs, and error budgets; lead incident response and post-mortems; automate operational toil; and partner with application teams on production readiness. Candidates should have 4+ years of SRE, DevOps, or infrastructure engineering experience and strong scripting and software-development skills.

Highlights

Fully remote SRE work focused on reliability, observability, performance, and automation. The role combines software engineering with operations, includes ownership of SLOs and incident processes, and provides a competitive salary with benefits.

Description

Role: Site Reliability EngineerLocation: RemoteEmployment Type: Full-time Compensation: Competitive salary commensurate with experience, qualifications, and location. Indicative range: $130,000 – $200,000 (USD equivalent), plus benefits. Role Overview We are hiring a Site Reliability Engineer to keep production reliable, observable, and fast. The role combines software engineering with operations — automating away toil and treating reliability as a first-class engineering problem. Key Responsibilities - Design and operate systems for high availability and performance - Build and maintain observability tooling (logging, metrics, tracing) - Define and track SLOs, SLIs, and error budgets - Lead incident response and post-mortem reviews - Automate operational toil through tooling and platform improvements - Partner with application teams on production readiness Required Skills and Qualifications - 4+ years in SRE, DevOps, or infrastructure engineering - Strong scripting and software engineering skills (Python, Go, or similar) - Deep experience with cloud platforms (AWS, GCP, Azure) - Hands-on with Kubernetes, Terraform, and observability platforms - Experience leading incident response in production environments - Strong understanding of distributed systems What You'll Bring - Curiosity to dig into systems and turn findings into shipped improvements - Strong written communication and ability to explain technical decisions - A test-and-learn mindset; you ship fast, measure, and iterate - Comfort working asynchronously across time zones What We Offer - Fully remote, flexible work hours - Performance-based bonus structure - Annual learning & development stipend - Health and wellness benefits (varies by location) - Opportunity to work on high-scale, real-world impact projects Equal Opportunity Statement This is an equal opportunity role. Applications are welcomed from all qualified individuals regardless of race, color, ethnicity, nationality, gender, gender identity or expression, sexual orientation, age, religion, disability, marital status, or any other characteristic protected by applicable law. All hiring decisions are based solely on qualifications, skills, and demonstrated ability.