Tech Lead - AI Backend, Kubernetes & GPU Platform

Nebul — Netherlands · Posted ~2 hours ago

Lead Full-time Visa History ✓

Skills

Backend engineering Technical leadership Kubernetes GPU orchestration AI infrastructure LLM inference Model serving Distributed systems Scalable architecture GPU Backend Open source

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Lead the backend engineering of a next-generation AI infrastructure platform while staying hands-on with implementation. You will shape scalable architecture for LLM inference and model serving, Kubernetes-based services, GPU orchestration, and distributed AI applications, combining technical leadership with strong backend engineering.

Highlights

Deeply technical leadership role shaping scalable AI platform architecture while remaining hands-on. Work spans LLM inference, model serving, Kubernetes services, GPU orchestration, and distributed AI workloads using open technologies.

Description

About Nebul At Nebul, we’re building Europe’s sovereign AI cloud — secure, high-performance and purpose-built as a European alternative to the American hyperscalers. Our NeoCloud platform runs on Nebul-owned infrastructure and open-source technologies, giving us control from the hardware layer all the way through to the AI and cloud services our customers use. Within our AI engineering teams, we’re building the backend and platform capabilities needed to run modern AI workloads at scale. That means everything from LLM inference and model serving to Kubernetes-based services, GPU orchestration and distributed AI applications. We’re looking for a Tech Lead – AI Backend, Kubernetes & GPU Platform who combines strong backend engineering experience with hands-on technical leadership. What You’ll Be Doing This is a deeply technical role. You’ll help shape the architecture behind our AI platform, while staying close to the software being built. Your focus will be on designing scalable backend systems, APIs and orchestration services that allow AI workloads to run reliably across Kubernetes and GPU infrastructure. You’ll work on systems supporting LLM inference, model serving, AI applications, multi-tenant environments and distributed GPU workloads. You’ll lead architectural decisions around microservices, APIs, asynchronous workflows, state management and distributed systems, primarily using Go. You’ll work closely with AI, platform, infrastructure and networking engineers to make sure the software and infrastructure layers work together as one system. This is not an ML research role and you do not need to be a machine-learning scientist. You do, however, need to understand how modern LLMs and AI applications behave in production and what that means for scalability, availability, networking and GPU utilisation. Alongside building software yourself, you’ll guide engineers, lead design discussions, identify bottlenecks before they become scaling problems and help set the engineering standards for reliability, observability, testing and maintainability. What You Bring You have extensive backend engineering experience and strong production experience with Go / Golang — this is a must-have. You’ve worked as a Senior, Staff, Principal, Lead Engineer or Tech Lead and have designed large-scale microservices and distributed systems in production. You understand APIs, service-to-service communication, asynchronous processing, databases, queues, caching, retries, idempotency and fault tolerance. You’re comfortable working in Kubernetes and cloud-native environments and understand Kubernetes beyond simply deploying applications. Most importantly, you can take ownership of complex technical problems, create clear architectural direction and help other engineers make better technical decisions. Experience with Python, LLM inference, model serving, AI agents, RAG, GPU workloads, Kubernetes controllers, NVIDIA infrastructure, vLLM, Triton, KServe, Ray, Kafka, NATS, Redis or OpenTelemetry would be a strong advantage. This Role Is Not This is not a traditional Engineering Manager, Scrum, DevOps or ML research role. You won’t spend your days managing delivery plans, training foundation models, manually provisioning infrastructure or writing endless Terraform and YAML. Your primary responsibility is technical leadership across backend software and AI platform engineering. Eligibility You must already live and work in the Netherlands and be able to travel to our Leiden office. We offer visa sponsorship, but only for candidates who are already based in the Netherlands. English fluency is required. Dutch is not. Build the Backend Behind Europe’s AI Infrastructure Running AI in production is about much more than the model. It requires scalable backend architecture, intelligent orchestration, reliable distributed systems and infrastructure capable of handling demanding GPU workloads. If that is the kind of engineering challenge you enjoy, we’d like to hear from you. Apply through Frank Poll and help us build the technology powering Europe’s sovereign AI future.