Summary
Join a global platform engineering team to improve the reliability, scalability, and resilience of cloud-based production systems. You will prevent incidents through automation and observability, build self-healing capabilities, improve monitoring and alerting, define service-level objectives, investigate root causes, and reduce operational complexity. The role requires at least four years of experience in SRE, platform, DevOps, or cloud engineering, along with strong Azure, Terraform, CI/CD, observability, and scripting skills.
Highlights
Work on a global cloud platform with strong emphasis on automation, resilience, observability, and operational excellence. Benefits include 27 vacation days, holiday pay, a fully employer-paid pension, hybrid work, performance bonuses, home-office support, and wellness allowances.
Description
Keep our global cloud platform reliable, scalable and resilient.
Do you enjoy solving complex production challenges before they become incidents? Are you the kind of engineer who automates repetitive work, improves reliability through engineering, and believes that every outage is an opportunity to build a better system?
We're looking for a Site Reliability Engineer to join the Global Platform Team at HeadFirst x Impellam Group.
In this role, you'll help build and operate the cloud platform behind our global Workforce-as-a-Service (WaaS) ecosystem, working alongside Cloud Engineers, Platform Engineers, Data Engineers and AI specialists to improve platform resilience, reduce operational overhead and ensure our Azure-based engineering landscape remains reliable, scalable and resilient.
Your impact
As a Site Reliability Engineer, your focus is simple: keep our platforms healthy, reliable and easy to operate.
You'll help build the engineering foundations behind our Headless Data Architecture (HDA), running on Azure and Databricks, as well as the Custom Apps Infrastructure (CA) that powers integrations, internal applications and operational workflows across our international organisation.
Rather than spending your days reacting to incidents, you'll focus on preventing them through automation, observability and reliability engineering.
You'll reduce operational toil, improve platform resilience and build systems that scale, recover automatically where possible, and give engineering teams across Cloud, Data and AI the visibility they need to operate production workloads with confidence.
What You Will Do
Improve the reliability, availability and performance of our Azure platform and production environments;Build and improve monitoring, logging and alerting using Grafana, OpenTelemetry, Azure Monitor and Log Analytics;Automate operational tasks and eliminate repetitive manual work using Infrastructure as Code and scripting;Design self-healing capabilities and automated remediation to reduce incidents and improve recovery times;Investigate production incidents, perform root cause analyses and implement long-term improvements;Define, measure and improve Service Level Indicators (SLIs) and Service Level Objectives (SLOs);Optimise platform performance, scalability and operational efficiency;Work closely with Cloud Engineers to improve platform architecture, resilience and security;Support Data and AI teams by improving the reliability of Azure Databricks environments;Drive engineering best practices around observability, automation and operational excellence;Continuously look for opportunities to reduce operational complexity and improve developer productivity.
About The Role
As part of the Global Platform Team, you'll work alongside engineers across Cloud, Data and AI to improve the reliability of our Azure-based platform.
Using technologies such as Kubernetes, Terraform, Databricks, GitHub Actions, Grafana and OpenTelemetry, you'll help ensure our global Workforce-as-a-Service ecosystem remains reliable, scalable and resilient.
About HeadFirst Group x Impellam Group
HeadFirst Group x Impellam Group is one of Europe's leading providers of workforce and talent solutions.
Operating across multiple countries, we're transforming into a cloud-native, AI-powered organisation that connects people, technology and data through a modern digital platform.
The Global Platform Team is at the heart of that transformation, enabling engineering teams across Cloud, Data, AI and Software Engineering to build and operate scalable solutions for the future.
Interested?
If you're excited about building reliable, scalable cloud platforms and enjoy solving complex engineering challenges, we'd love to hear from you.
Apply today and let's discover how you can make an impact as part of our Global Platform Team.
Here’s what we offer you
Salaris dat past bij jouw ervaring
Je ontvangt een range aan salaris dat aansluit op jouw kennis en ervaring.
We belonen je eerlijk voor je inzet en ontwikkeling.
Vakantiegeld en vrije dagen
Je ontvangt 8,33% vakantiegeld en hebt recht op 27 vakantiedagen per jaar op basis van een fulltime dienstverband.
Zo heb je voldoende ruimte om op te laden.
Hybride werken
Wij ondersteunen hybride werken en hanteren hiervoor een 60/40-verdeling, waarbij je bij een fulltime dienstverband drie dagen op kantoor bent en twee dagen vanuit huis kunt werken.
Premievrij pensioen
Wij betalen jouw volledige pensioenpremie.
Je bouwt dus pensioen op zonder dat je daar zelf aan bijdraagt.
Je houdt hierdoor een hoger nettosalaris over.
Prestatiebonus
Afhankelijk van jouw prestaties én die van HeadFirst Group kun je jaarlijks een extra beloning ontvangen van één of twee maandsalarissen.
Maandelijkse extra’s
Je ontvangt een maandelijkse vergoeding voor internet, een vitaliteitsbudget, een bijdrage voor lunch en ondersteuning bij het inrichten van je thuiswerkplek.
What we expect from you
You're passionate about building reliable systems and solving operational challenges through engineering rather than manual intervention.
You enjoy understanding how distributed systems behave, thrive in cloud-native environments and are always looking for ways to improve automation, resilience and observability.
You stay calm under pressure, take ownership of problems and enjoy collaborating with others to continuously improve the reliability of the platform.
Ideally, You Also Bring
4+ years of experience as a Site Reliability Engineer, Platform Engineer, DevOps Engineer or Cloud Engineer;Strong hands-on experience with Microsoft Azure;Experience with Infrastructure as Code using Terraform;Experience building and maintaining CI/CD pipelines using GitHub Actions or Azure DevOps;Experience with observability tooling such as Grafana, OpenTelemetry, Azure Monitor or Log Analytics;Strong scripting skills using Python, Bash or similar languages;Experience supporting distributed cloud platforms in production;Experience with incident management, root cause analysis and post-incident improvements;Familiarity with GitOps principles and modern deployment practices;Experience with Azure Databricks is a strong advantage;Experience with SnapLogic or similar integration platforms is a plus.
We know the perfect candidate doesn't exist.
If this role excites you but you don't meet every single requirement, we'd still love to hear from you.
We're just as interested in your potential, mindset and ambition as we are in your experience.