Summary
✨ AI‑Generated
Join a reliability engineering team responsible for highly available cloud platforms. You will improve resilience, automate operations, define reliability targets, and build self-healing systems.
Highlights
Senior engineering role focused on large-scale cloud reliability, automation, observability, and designing resilient systems.
Description
Are you an experienced Site Reliability Engineer who sees reliability as an engineering challenge rather than a series of incidents to resolve? Do you enjoy building self-healing infrastructure, eliminating repetitive operational work and improving large-scale cloud platforms? Then this could be an interesting opportunity for you!
About the position
As a Senior Site Reliability Engineer, you will be responsible for the reliability of a large international SaaS environment running primarily on Microsoft Azure.
The platform operates across more than 10 global data centres, serves millions of end users and needs to deliver 24/7 availability backed by strict SLAs.
This is not a traditional operations or ticket-driven infrastructure role.
You will approach reliability from an engineering perspective.
You define and manage SLOs and error budgets, improve observability, automate repetitive work and design systems that can recover automatically when something goes wrong.
You will work closely with cloud engineers and software development teams.
Instead of only becoming involved after an incident occurs, you will participate early in the development process and help engineering teams make architectural decisions that improve scalability, performance and reliability.
What will you do?
Define and manage SLOs and error budgets for critical services across the Azure environmentIdentify operational toil and replace repetitive manual work with automation and self-healing solutionsImprove and standardise observability, including metrics, alerting and tracingLead complex production incidents and facilitate blameless postmortemsTranslate lessons from incidents into structural improvements and automationImprove CI/CD and progressive delivery through safe deployments, canary releases and automated rollbackManage capacity, performance and infrastructure cost forecasting across the platformBuild reliability automation using technologies such as Python and TerraformIntroduce AI agents and bounded automation into detection, diagnosis and remediationWork proactively with product and engineering teams to incorporate SRE principles into new solutions from the design stage
Who are you?
You are an experienced engineer who enjoys taking ownership of complex production environments.
You don't just want to keep systems running; you want to understand why problems occur and engineer them out of the environment.
Ideally, you bring:
Around 5+ years of experience as a Site Reliability Engineer, DevOps Engineer, Cloud Engineer or Infrastructure EngineerHands-on experience managing production cloud environments, preferably Microsoft AzureExperience working with SLOs, error budgets and reliability engineering principlesStrong knowledge of observability and monitoring, for example Grafana, Prometheus or similar toolingHands-on Kubernetes experience, including areas such as Helm, ingress, RBAC and persistent volumesExperience with Terraform and Infrastructure as CodeExperience using Python or another language for automationStrong Linux knowledgeExperience working with CI/CD environmentsThe confidence to take ownership during production incidents and participate in an on-call rotationStrong communication skills and the ability to document technical decisions, incidents and architecture clearly
About the organization
You will join an international software company that develops service management software for organizations across sectors such as government, education, healthcare and industry.
The organization employs more than 700 people across eight international offices, while its software is used by more than 10 million users worldwide.
The working environment is characterized by limited hierarchy, significant individual responsibility and close collaboration between engineering teams.
Technology and innovation are central to the organization, with continuous investment in its SaaS platform, automation and the adoption of AI within both its products and engineering processes.
What do we offer?
Salary up to €102.000,-32 to 40-hour working week26 vacation daysHybrid working, including a home office allowanceA pension scheme with employer contributionA personal development budget for conferences, certifications, and coursesA Vitality Budget of €50 per month for sports, gym membership, or other things that contribute to your mental and physical healthTravel expense reimbursement or an NS Business Card, plus the option to lease a bicycle through the companyA senior technical position with significant ownershipThe opportunity to work on an international SaaS platform used by millions of users10% of your working time dedicated to personal developmentA personal development budget equivalent to 10% of your gross annual salaryA modern laptop and phone
Interested?
Are you an experienced Infrastructure Engineer looking for a role that combines infrastructure, software development and automation? And would you like to work on the technology behind a global SaaS platform?
Get in touch and send us your latest CV.
We will be happy to discuss whether your experience matches the position.