Skills
Site reliability engineering
Incident management
Production incident response
On-call engineering
Incident Commander experience
Blameless postmortems
Root-cause analysis
On-call runbooks
Severity and escalation documentation
SLOs
Error budgets
Reliability reviews
SRE
Production incident management
On-call systems
Summary
✨ AI‑Generated
A senior site reliability engineer is sought for a flexible remote opportunity focused on creating, evaluating, and refining AI-generated reliability and incident-management materials. The ideal candidate has 4+ years of experience running or reviewing production incidents, with direct ownership of blameless postmortems and root-cause analyses plus strong knowledge of runbooks, escalation, SLOs, error budgets, and reliability reviews.
Highlights
Fully remote senior-level opportunity with flexible scheduling and strong hourly compensation, allowing experienced reliability engineers to apply their production expertise to advanced AI model development.
Description
About This Opportunity
We're working with a leading foundational AI lab to find experienced senior site reliability engineers who can help train their latest language model on professional document, spreadsheet, and slide deck tasks.
We're looking for senior site reliability engineers with 4+ years running or reviewing production incidents to create, evaluate, and refine AI-generated documents, spreadsheets, and slide decks across core workflows: blameless postmortems and root-cause analyses, on-call runbooks, severity and escalation write-ups, SLO and error-budget reports, remediation action-item trackers, and reliability review decks.
Compensation: $80/hour
Commitment: Flexible, 5-20 hours per week (or more if desired)
Location: Fully remote, work on your own schedule
Start date: ASAP
Qualifications
4+ years as a Site Reliability Engineer, Incident Commander, or production/on-call engineering lead Direct ownership of writing or reviewing blameless postmortems and root-cause analyses Expert-level document, spreadsheet, and slide craftsmanship, with excellent written communication and attention to detail
About Ethos
Ethos is a new expert network built by a McKinsey/SoftBank/DeepMind team and backed by world-leading investors like General Catalyst.
We connect experts with investors and consultancies for paid expert calls, speaking engagements, and advisory opportunities.
Key Requirements
4+ years as a Site Reliability Engineer, Incident Commander, or production/on-call engineering leadDirect ownership of writing or reviewing blameless postmortems and root-cause analysesExpert-level document, spreadsheet, and slide craftsmanship, with excellent written communication and attention to detail