Site Reliability Engineer

Hirefeedd — New Zealand · Posted ~1 hour ago

Mid Full-time Remote $130000-$200000 (USD equivalent)

Skills

Scripting Software Engineering Observability Tools SLO/SLI/Error Budget Management Incident Response Automation

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A company is hiring a Site Reliability Engineer to keep production reliable, observable, and fast. The role combines software engineering with operations—automating away toil and treating reliability as a first-class engineering problem. Responsibilities include designing and operating systems for high availability and performance, building and maintaining observability tooling, defining and tracking SLOs, SLIs, and error budgets, leading incident response and post-mortem reviews, automating operational toil, and partnering with application teams on production readiness. The role is fully remote with a competitive salary commensurate with experience. Requires 4+ years in SRE, DevOps, or infrastructure engineering with strong scripting and software skills.

Highlights

Fully remote position with competitive salary range of $130,000 - $200,000 USD equivalent. Combines software engineering with operations, treating reliability as a first-class engineering problem. Opportunity to partner with application teams on production readiness.

Description

Role: Site Reliability EngineerLocation: RemoteEmployment Type: Full-time Compensation: Competitive salary commensurate with experience, qualifications, and location. Indicative range: $130,000 – $200,000 (USD equivalent), plus benefits. Role Overview We are hiring a Site Reliability Engineer to keep production reliable, observable, and fast. The role combines software engineering with operations — automating away toil and treating reliability as a first-class engineering problem. Key Responsibilities - Design and operate systems for high availability and performance - Build and maintain observability tooling (logging, metrics, tracing) - Define and track SLOs, SLIs, and error budgets - Lead incident response and post-mortem reviews - Automate operational toil through tooling and platform improvements - Partner with application teams on production readiness Required Skills and Qualifications - 4+ years in SRE, DevOps, or infrastructure engineering - Strong scripting and software engineering skills (Python, Go, or similar) - Deep experience with cloud platforms (AWS, GCP, Azure) - Hands-on with Kubernetes, Terraform, and observability platforms - Experience leading incident response in production environments - Strong understanding of distributed systems What You'll Bring - Curiosity to dig into systems and turn findings into shipped improvements - Strong written communication and ability to explain technical decisions - A test-and-learn mindset; you ship fast, measure, and iterate - Comfort working asynchronously across time zones What We Offer - Fully remote, flexible work hours - Performance-based bonus structure - Annual learning & development stipend - Health and wellness benefits (varies by location) - Opportunity to work on high-scale, real-world impact projects Equal Opportunity Statement This is an equal opportunity role. Applications are welcomed from all qualified individuals regardless of race, color, ethnicity, nationality, gender, gender identity or expression, sexual orientation, age, religion, disability, marital status, or any other characteristic protected by applicable law. All hiring decisions are based solely on qualifications, skills, and demonstrated ability.