Summary
✨ AI‑Generated
A senior reliability engineering contract focused on bringing critical data systems into robust production environments. The role involves observability, incident prevention, resilience improvements, and operational excellence.
Highlights
High-impact SRE role improving reliability, monitoring, and resilience of large-scale production data systems.
Description
Site Reliability Engineer (SRE) – Travel – Amsterdam
Hourly rate: €90 - €130
Duration: 6 months
Hybrid: 2 days per week
Start: ASAP
We’re looking for an SRE to help our data engineering team bring two important event-data pipelines fully into production.
The work focuses on making the pipelines reliable, observable, and resilient at scale—not just getting them running, but ensuring they meet production requirements.
The pipelines support company analysis and experiment safety.
You’ll help the team understand how they perform in production, detect issues quickly, prevent data loss, and recover from faults.
What you’ll do
Help take data pipelines from development to reliable, fully monitored production systems.Design and improve monitoring, alerting, and operational visibility.Identify ways to detect missing or delayed data quickly, so issues can be addressed before they cause downstream problems.Build resilience into systems, including fault handling and self-healing behavior.Assess system capacity and reliability as pipeline scale and usage grow.Work across the boundaries of SRE, software development, and data engineering to improve production readiness.Support the migration of existing services and libraries, including systems running in AWS Lambda and on-premises environments.Help address architectural issues, including reducing reliance on sidecars where appropriate.Provide operational support during the workday and help design systems that avoid unnecessary pager-based support.
What we’re looking for
Experience taking systems from a greenfield stage into production.Experience operating or improving high-scale systems, ideally event-driven systems or data pipelines.Strong understanding of monitoring, fault detection, reliability, and recovery practices.Ability to identify production risks and turn them into practical engineering improvements.Comfort working across software, infrastructure, and data engineering concerns.A proactive, collaborative approach to solving reliability problems.
Helpful experience
Experience with AWS Lambda or on-premises systems.Familiarity with migrations, event pipelines, or systems with strict data-integrity needs.Experience with native-library integration techniques such as JNI or FFI.
This is useful, but not a core requirement.