Senior Site Reliability Engineer / Platform Engineer

Niahitrecruitment — Netherlands · Posted ~3 hours ago

Senior Full-time Onsite

Skills

Site reliability engineering Platform engineering Cloud infrastructure Observability Metrics Logging Tracing Alerting CI/CD Progressive delivery Incident response

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A fast-growing technology scale-up is seeking a Senior SRE / Platform Engineer to become a central authority for reliable production operations. You will establish reliability targets, improve metrics, logs, traces and alerting, make deployments safer through progressive delivery, scale cloud workloads, and improve recovery when failures occur. The role is full-time and on-site.

Highlights

Senior full-time on-site role with ownership of production reliability, observability, safer deployments, cloud scalability, recovery processes, and engineering standards for a growing technology platform.

Description

Senior SRE / Platform Engineer | Zuid-Holland | Full-time | On-site About the opportunity Our client is a fast-growing, tech-driven scale-up with a large international customer base. Engineering is rebuilding key parts of the platform and bringing machine learning and LLM-based features into production. They're looking for a senior engineer to become the go-to person for how software runs in production: how stable it is, how safely it ships, and how efficiently it uses the cloud. This goes well beyond keeping servers alive. You'll set the standards the rest of engineering works by. Your focus Setting reliability targets for key services and making sure teams actually work with them Making systems easier to understand under pressure: better metrics, logs, traces and alerting Making releases boring: safer pipelines, progressive rollouts, fast recovery when something breaks Running and scaling cloud workloads, including services that depend on third-party APIs Keeping an eye on throughput, backlogs, API quotas and dependencies outside your control Finding and cutting wasted cloud spend Taking the lead during outages, then making sure the same failure doesn't happen twice Planning for growth, failover and recovery Replacing manual ops work with tooling so product teams can ship on their own What you bring Years of experience keeping busy, customer-facing systems healthy Deep hands-on cloud experience, with infrastructure defined in code (Terraform or similar) Solid grounding in Linux, containers, databases, caching, messaging and networking Comfort with modern observability tooling The ability to write real code, ideally Python or TypeScript A preference for fixing root causes over applying patches Willingness to own things end to end Happy to work on-site Nice to have: experience running LLM-based features, async job processing, or anything that has to live with external rate limits. What's in it for you Real ownership, a say in how the engineering organisation works, problems that matter to the business, and a team that makes decisions quickly. Curious? Get in touch for a no-strings conversation. We look at how you think, not just your CV.