Site Reliability Engineer II

Brookwood Recruitment Ltd — Netherlands · Posted ~12 hours ago

Mid Full-time

Skills

Site reliability engineering Software engineering Automation Incident response Scalability Observability SRE Cloud infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is hiring an SRE to improve service reliability through automation, software engineering practices, monitoring, incident management, and scalable infrastructure solutions.

Highlights

Role focused on building reliable systems, automation, operational excellence, and improving large-scale service performance.

Description

Site Reliability Engineer II Overview As a Site Reliability Engineer II (SRE II), you’ll treat operations as a software problem—building resilient, scalable services while reducing manual effort through automation. This role is centred on availability, performance, scalability, latency, observability, and efficiency, with a strong emphasis on end-to-end service ownership. You’ll design and implement technical solutions based on business needs, estimate effort and impact with confidence, and deliver high-quality engineering outcomes—working primarily within your team while collaborating with partner teams when it benefits reliability and delivery. You’ll also play an active part in incident response and post-incident learning, helping drive long-term reliability improvements. Required SkillsProven ability to improve system reliability across availability, performance, scalability, and latencyStrong software engineering capability, including:Building software applications using relevant development languagesWriting readable, reusable code using standard libraries and patternsRefactoring and simplifying code (including introducing design patterns where appropriate)Applying standard testing techniques aligned to a test strategySystem and service design skills:Evaluating architecture options with cost, business, and technology trade-offs in mindUnderstanding the implications of changes to existing systems within a broader platform contextEnd-to-end ownership of services in production:Monitoring health and performanceDefining, tracking, and acting on meaningful metricsWriting and maintaining operational documentation (e.g., runbooks / operational documentation)Technical incident management:Mitigating customer impact within SLARoot cause analysis and long-term corrective actionsContributing to postmortems and incident trackingAutomation and toil reduction:Reducing technical debt, identifying bottlenecks, and preparing services for scaleBuilding small software features to improve reliability and operational efficiencyObservability (monitoring & alerting):Capacity planning and tracking service KPIs and observability metricsPartnering with engineering teams to implement effective instrumentation and alertingStrong critical thinking, communication, and collaboration skills, including the ability to explain concepts clearly to different audiences and work towards shared outcomes Nice to Have SkillsExperience advising product or engineering teams on architectural direction and non-functional requirementsExposure to continuous delivery and experimentation frameworks to reduce risk and improve feedback loopsExperience partnering with vendors or evaluating tools/approaches through prototyping or technical spikesDemonstrated contributions to process and quality improvements across systems and team practices Preferred Education and ExperienceBachelor’s degree (preferred)Broad relevant job knowledge: 3–5 years in a related reliability, operations, or software engineering environment Other RequirementsParticipation in incident response for issues affecting the team’s servicesAbility to work across stakeholders (e.g., product stakeholders and peers) with clear, inclusive communicationComfortable independently managing deployment and production operations for owned servicesIf you’re excited to build reliable systems, reduce toil through automation, and improve how services run at scale, apply now and share your experience—your next reliability challenge starts here.