Observability Engineer – Expert Opportunity

Askethos — Canada · Posted ~3 hours ago

Senior Contract Remote $80/hour, up to $1600/weeK

Skills

observability site reliability engineering incident management production operations root-cause analysis postmortem writing on-call operations SRE SLOs error budgets on-call systems

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A flexible remote opportunity for an experienced observability or reliability engineer with hands-on production incident experience. You will create, evaluate, and refine AI-generated operational materials such as postmortems, runbooks, SLO reports, escalation documentation, and reliability reviews. Work 5–20 hours per week or more on your own schedule.

Highlights

Flexible fully remote expert work at $80/hour, with a self-directed schedule and the opportunity to apply production reliability expertise to training and evaluating advanced AI systems.

Description

About This Opportunity We're working with a leading foundational AI lab to find experienced observability engineers who can help train their latest language model on professional document, spreadsheet, and slide deck tasks. We're looking for observability engineers with 4+ years running or reviewing production incidents to create, evaluate, and refine AI-generated documents, spreadsheets, and slide decks across core workflows: blameless postmortems and root-cause analyses, on-call runbooks, severity and escalation write-ups, SLO and error-budget reports, remediation action-item trackers, and reliability review decks. Compensation: $80/hour Commitment: Flexible, 5-20 hours per week (or more if desired) Location: Fully remote, work on your own schedule Start date: ASAP Qualifications 4+ years as a Site Reliability Engineer, Incident Commander, or production/on-call engineering lead Direct ownership of writing or reviewing blameless postmortems and root-cause analyses Expert-level document, spreadsheet, and slide craftsmanship, with excellent written communication and attention to detail About Ethos Ethos is a new expert network built by a McKinsey/SoftBank/DeepMind team and backed by world-leading investors like General Catalyst. We connect experts with investors and consultancies for paid expert calls, speaking engagements, and advisory opportunities. Key Requirements 4+ years as a Site Reliability Engineer, Incident Commander, or production/on-call engineering leadDirect ownership of writing or reviewing blameless postmortems and root-cause analysesExpert-level document, spreadsheet, and slide craftsmanship, with excellent written communication and attention to detail