Observability Engineer

Askethos — Netherlands · Posted ~3 hours ago

Senior Contract Remote $80/hour, up to $1600/weeK

Skills

observability site reliability engineering incident management production incident response blameless postmortems root-cause analysis on-call operations SLOs error budgets SRE SLO

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Use your production reliability expertise to help evaluate and refine AI-generated professional documents, spreadsheets, and presentations. You will work with incident-management workflows such as blameless postmortems, root-cause analyses, on-call runbooks, severity reports, SLO and error-budget reviews, and remediation tracking.

Highlights

Fully remote, flexible scheduling with 5–20 hours per week or more, paid at $80 per hour. The role offers an opportunity to apply real-world reliability and incident-management expertise to improve AI-generated professional documents and workflows.

Description

About This Opportunity We're working with a leading foundational AI lab to find experienced observability engineers who can help train their latest language model on professional document, spreadsheet, and slide deck tasks. We're looking for observability engineers with 4+ years running or reviewing production incidents to create, evaluate, and refine AI-generated documents, spreadsheets, and slide decks across core workflows: blameless postmortems and root-cause analyses, on-call runbooks, severity and escalation write-ups, SLO and error-budget reports, remediation action-item trackers, and reliability review decks. Compensation: $80/hour Commitment: Flexible, 5-20 hours per week (or more if desired) Location: Fully remote, work on your own schedule Start date: ASAP Qualifications 4+ years as a Site Reliability Engineer, Incident Commander, or production/on-call engineering lead Direct ownership of writing or reviewing blameless postmortems and root-cause analyses Expert-level document, spreadsheet, and slide craftsmanship, with excellent written communication and attention to detail About Ethos Ethos is a new expert network built by a McKinsey/SoftBank/DeepMind team and backed by world-leading investors like General Catalyst. We connect experts with investors and consultancies for paid expert calls, speaking engagements, and advisory opportunities. Key Requirements 4+ years as a Site Reliability Engineer, Incident Commander, or production/on-call engineering leadDirect ownership of writing or reviewing blameless postmortems and root-cause analysesExpert-level document, spreadsheet, and slide craftsmanship, with excellent written communication and attention to detail