Summary
✨ AI‑Generated
Take ownership of reliability, observability, cloud infrastructure, and AI platform operations within a rapidly evolving technology environment. You will define service objectives, build operational guardrails and automation, and help engineering teams innovate while maintaining stability and performance. The position is fully office-based and suited to an experienced SRE professional.
Highlights
Ownership of reliability, observability, cloud infrastructure, and AI platform operations in a technology-focused environment. The role emphasizes automation, innovation, continuous improvement, and direct business impact.
Description
Senior Site Reliability Engineer (SRE & AI Platform Operations)
Rotterdam, 5 days in office
Competitive Salary
This is an opportunity to take ownership of reliability, observability, cloud infrastructure, and AI platform operations within a fast-moving technology environment.
You will play a key role in building the guardrails, automation, and operational excellence that enable engineering teams to innovate at pace while maintaining stability and performance.
The Company
They are a technology-driven organisation undergoing significant platform transformation, modernising core systems and investing heavily in cloud-native architecture and AI-enabled capabilities.
Their engineering culture is focused on innovation, automation, and continuous improvement.
You will join a collaborative environment where technical expertise is valued and where your work will have a direct impact on business performance and customer experience.
The Role
Define and manage Service Level Objectives (SLOs), SLIs, and error budget policies across critical services.Lead initiatives across observability, distributed tracing, monitoring, and incident response.Improve deployment safety through CI/CD best practices, automated rollbacks, and progressive delivery techniques.Own AI platform operations, including runtime performance, scalability, reliability, and cost optimisation.Drive FinOps initiatives across cloud infrastructure, identifying opportunities to improve efficiency and manage costs.Lead incident management activities and develop automated safeguards to prevent recurring issues.Enhance resilience through capacity planning, load testing, disaster recovery planning, and service reliability improvements.Manage infrastructure through Infrastructure as Code and cloud automation practices.Build tooling, runbooks, and self-service capabilities that improve the developer experience and reduce operational overhead.
Your Skills & Experience
Strong commercial experience operating high-traffic, distributed production systems.Deep knowledge of Site Reliability Engineering principles, monitoring, and platform operations.Hands-on experience with Google Cloud Platform and Infrastructure as Code using Terraform.Strong troubleshooting capabilities across Linux environments, cloud infrastructure, containers, databases, and distributed systems.Experience with observability tooling, monitoring platforms, and distributed tracing.Proficiency in Python or JavaScript/TypeScript for automation and platform tooling.Familiarity with modern software engineering environments and cloud-native architectures.Understanding of AI workloads, LLM integrations, API-driven services, or automated processing pipelines.A proactive, solutions-focused approach with a passion for automation and operational excellence.
What They Offer
Competitive salary and benefits package.The opportunity to shape platform reliability and AI operations at scale.Real ownership and influence across engineering and infrastructure strategy.An international and collaborative working environment.Clear opportunities for professional growth and career progression.The chance to work on cutting-edge cloud and AI technologies.
How to Apply
If you are an experienced Site Reliability Engineer looking to take ownership of reliability, cloud infrastructure, and AI platform operations within a high-impact environment, apply today to learn more.