Skills
Site reliability engineering
Incident management
Production incident response
On-call engineering
Incident Commander experience
Blameless postmortems
Root-cause analysis
On-call runbooks
Severity and escalation documentation
SLOs
Error budgets
Reliability reviews
SRE
Production incident management
On-call systems
Summary
✨ AI‑Generated
An experienced platform or reliability engineer is sought for a flexible remote opportunity focused on creating, evaluating, and refining AI-generated professional reliability documentation. Candidates should have 4+ years of experience handling or reviewing production incidents and strong expertise in blameless postmortems, root-cause analysis, runbooks, escalation procedures, SLOs, error budgets, and reliability reviews.
Highlights
Fully remote expert opportunity offering flexible hours and strong hourly compensation while applying production reliability and incident-management expertise to advanced AI model training.
Description
About This Opportunity
We're working with a leading foundational AI lab to find experienced platform engineers who can help train their latest language model on professional document, spreadsheet, and slide deck tasks.
We're looking for platform engineers with 4+ years running or reviewing production incidents to create, evaluate, and refine AI-generated documents, spreadsheets, and slide decks across core workflows: blameless postmortems and root-cause analyses, on-call runbooks, severity and escalation write-ups, SLO and error-budget reports, remediation action-item trackers, and reliability review decks.
Compensation: $80/hour
Commitment: Flexible, 5-20 hours per week (or more if desired)
Location: Fully remote, work on your own schedule
Start date: ASAP
Qualifications
4+ years as a Site Reliability Engineer, Incident Commander, or production/on-call engineering lead Direct ownership of writing or reviewing blameless postmortems and root-cause analyses Expert-level document, spreadsheet, and slide craftsmanship, with excellent written communication and attention to detail
About Ethos
Ethos is a new expert network built by a McKinsey/SoftBank/DeepMind team and backed by world-leading investors like General Catalyst.
We connect experts with investors and consultancies for paid expert calls, speaking engagements, and advisory opportunities.
Key Requirements
4+ years as a Site Reliability Engineer, Incident Commander, or production/on-call engineering leadDirect ownership of writing or reviewing blameless postmortems and root-cause analysesExpert-level document, spreadsheet, and slide craftsmanship, with excellent written communication and attention to detail