Summary
A reliability engineer is needed to maintain a large-scale AI infrastructure platform, improve observability, automate operations, and resolve complex production issues.
Highlights
Opportunity to improve reliability of advanced cloud infrastructure, automate operations, work with cutting-edge technologies, and receive sponsorship support.
Description
About Nebul
At Nebul, we’re building Europe’s sovereign AI cloud — trusted, secure, and purpose-built for the next generation of intelligent infrastructure.
Our platform combines Kubernetes, NVIDIA GPU infrastructure, cloud-native services and software written in Go and Python.
As the platform grows, we need to increase our internal reliability capability and reduce the number of operational issues that depend on a small group of senior engineers.
What You’ll Be Doing
As a Site Reliability Engineer, you’ll work approximately 70% on site reliability and operational troubleshooting and 30% on broader DevOps and platform engineering.
Your primary responsibility will be to keep Nebul’s AI cloud platform stable, observable and operationally scalable.
You’ll take ownership of incidents, investigate complex issues and ensure that senior engineers are not pulled into every troubleshooting session.
You’ll work across Kubernetes, NVIDIA GPU infrastructure, services written in Go and Python, networking and cloud-native platform components.
The role is not only about responding to incidents.
You’ll identify recurring problems, automate operational work and improve the platform so that the same issues do not continue to return.
Key Responsibilities
Monitor and improve the reliability, availability and performance of Nebul’s AI cloud platform.Troubleshoot incidents across Kubernetes, NVIDIA GPU infrastructure, Linux, networking and platform services.Take ownership of technical troubleshooting sessions and coordinate issues through to resolution.Investigate issues affecting services written in Go and Python.Perform root-cause analysis and translate findings into structural platform improvements.Build and improve monitoring, metrics, logging, tracing and alerting.Ensure alerts are relevant, actionable and connected to clear operational procedures.Create runbooks, escalation procedures and troubleshooting documentation.Automate repetitive operational tasks using Go, Python or scripting.Support Kubernetes cluster operations, deployments, upgrades and platform changes.Improve the production readiness of new services and infrastructure components.Work with engineering teams to improve resilience, observability and failure handling.Identify reliability risks before they result in platform or customer impact.Support operational improvements across GPU workloads and NVIDIA-based infrastructure.Reduce the operational dependency on senior platform engineers.Contribute to broader DevOps work when additional capacity is needed within the team.
What Your Day Will Not Look Like
Acting as a first-line support engineer who only closes tickets.Escalating every complex issue to Diego or another senior engineer.Spending all your time manually operating Kubernetes.Resolving incidents without addressing their underlying causes.Building isolated automation that is not integrated into the platform.
What You Bring
Strong experience as a Site Reliability Engineer, DevOps Engineer or Platform Engineer.Hands-on production experience with Kubernetes.Strong Linux and infrastructure troubleshooting skills.Experience investigating issues across applications, infrastructure, networking and cloud platforms.Experience with services written in Go or Python.Practical experience with monitoring, logging, metrics and alerting.Experience responding to production incidents and performing root-cause analysis.The ability to independently lead complex troubleshooting sessions.Experience automating operational work using Python, Go or scripting.A solid understanding of cloud-native and distributed systems.A calm, analytical and structured approach to incidents.An ownership mindset and the ability to move from reactive troubleshooting to lasting improvements.
Bonus Points If You Have
Experience with NVIDIA GPU infrastructure.Experience supporting AI, machine-learning or high-performance computing workloads.Knowledge of Kubernetes GPU scheduling and resource management.Experience with multi-tenant cloud environments.Familiarity with Go-based cloud or platform services.Experience with Infrastructure as Code and automated platform deployment.Knowledge of distributed storage, networking or database troubleshooting.Experience defining service-level indicators, objectives and operational reliability targets.Experience working in sovereign, regulated or security-sensitive cloud environments.
Eligibility & Application Information
We welcome non-native Dutch speakers to apply.
However, to be eligible, you must:
Have a valid work permit in the Netherlands.
( wo do offer sponsorship if needeed)Reside in the Netherlands and be able to travel to the office in Leiden (near The Hague).Be fluent in English.
Dutch is not required.
Ready to make Europe’s sovereign AI cloud more reliable and operationally scalable?
Apply now through Frank Poll and help Nebul build a cloud platform that engineering teams and customers can depend on.