Description
This position is listed on behalf of a partner company, who manages all applications and next steps.
Our partner is looking for a Technical Lead - GPU Infrastructure based in Poland.
This is a hands-on technical leadership role responsible for building and delivering a full-stack GPU infrastructure platform.
You will own architecture and implementation across bare-metal GPU scheduling, Kubernetes, managed inference, and platform observability.
The role combines deep infrastructure engineering with leadership of a distributed team spanning backend, frontend, DevOps, QA, and documentation.
You will help operate GPU compute environments for research, model training, and inference workloads at scale.
A key focus will be reliable GPU fleet operations, high-performance networking, multi-tenant infrastructure, and production-grade platform services.
You will also serve as the primary technical interface for infrastructure partners while translating internal workload requirements into robust platform capabilities.
The position is fully remote and offers the opportunity to shape a critical AI infrastructure platform from architecture through production delivery.
Accountabilities
Own the end-to-end platform architecture, including architecture proposals, high-level and low-level designs, technical reviews, and maintaining the architecture baseline.Lead and line-manage a distributed engineering team across backend, frontend, DevOps, QA, and documentation, setting engineering standards and overseeing code reviews, design reviews, release gates, one-to-ones, and performance development.Design, build, and operate a managed Slurm service for research users, covering controllers, accounting, partitions, login nodes, node onboarding, NVIDIA drivers and CUDA baselines, job and node-health monitoring, autohealing, storage visibility, identity, and workload isolation.Own Kubernetes cluster bootstrap and lifecycle on bare-metal infrastructure, including NVIDIA GPU and Network Operators, KubeVirt and VFIO-based GPU isolation, upgrades, backup and recovery, and node replacement.Define and deliver managed inference architecture covering serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable infrastructure.Establish observability and operational practices across the control plane, GPU fleet, and application tiers, including metrics, logging, alerting, SLOs, incident response, post-incident reviews, and a sustainable on-call model.Act as the primary technical interface with infrastructure partners and vendors, translating requirements into specifications and acceptance tests, managing escalations, and contributing to capacity planning and hardware sourcing.Work directly with research, model-training, and product teams to understand workloads, translate requirements into platform capabilities, and manage capacity constraints.Hire additional members of the platform team and establish the technical standards and expectations for future engineering hires.Maintain a hands-on contribution to architecture, technical reviews, implementation decisions, and infrastructure delivery rather than operating solely in a management capacity.
Requirements
Bring 8+ years of hands-on engineering experience, including at least 3 years leading teams that build and operate infrastructure platforms used by other teams.Hold a Bachelor's or Master's degree in Computer Science, Engineering, or a related field, or demonstrate equivalent practical experience.Have hands-on experience operating Slurm at scale, including slurmctld, slurmdbd, partitions, QoS and priority, accounting, prolog and epilog, node-health scripting, and upgrades while jobs remain active; experience with HPC or GPU training clusters is highly desirable.Demonstrate deep experience operating NVIDIA GPU fleets on bare metal, including driver and CUDA lifecycle management, Fabric Manager, NVSwitch on SXM systems, DCGM-based health and utilization, MIG, node burn-in, and acceptance processes.Have strong knowledge of high-performance interconnects, including InfiniBand fabrics, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance issues.Possess advanced Linux systems knowledge covering kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning for compute-intensive workloads.Have production Kubernetes operations experience beyond deployment, including control planes, upgrades, CNI and CSI, operators, custom controllers, and multi-tenancy architecture.Understand HPC storage and large-scale data movement, including shared filesystems such as VAST, Lustre, or NFS, node-local NVMe caching, and distributing large model weights and datasets across multiple nodes.Be experienced with observability and production operations using Prometheus, Grafana, Loki, or equivalent technologies, together with SLO management, incident response, and post-incident reviews.Have working fluency in JavaScript and Node.js sufficient to review control-plane, CLI, and worker services and make architecture decisions, without requiring feature-development specialization.Have shipped a platform with real users, such as a multi-tenant IaaS/PaaS, research computing service, or comparable infrastructure platform involving resource isolation, quotas, usage metering, and user-facing APIs or CLI surfaces.Demonstrate leadership that remains technically engaged, including people management across time zones, cross-track technical reviews, documented architecture decisions, and the confidence to challenge partners or executives with clear technical reasoning.Have excellent written and spoken English, particularly for technical, partner, and leadership communication.Be based within the UTC to UTC+5:30 time-zone range to provide working-hour overlap with teams and partners in Europe and India, with availability for occasional travel to partner sites and team events.Experience with Slurm operators on Kubernetes, Kubernetes-native schedulers, modern AI serving stacks such as vLLM, SGLang, or TensorRT-LLM, and GPU parallelism or quantization strategies is a plus.Experience with KubeVirt, Kata Containers, QEMU/KVM, Firecracker, confidential computing, Cluster API, kubeadm, Cilium, infrastructure as code, GitOps, or GPU autohealing technologies is desirable.Background working on GPU cloud platforms, university or national HPC centers, AI lab infrastructure teams, peer-to-peer or distributed systems, or hardware-provider relationships is also considered an advantage.
Benefits
Fully remote work within the specified UTC to UTC+5:30 time-zone range.The opportunity to lead architecture and delivery of a full-stack GPU infrastructure platform spanning bare-metal compute, Kubernetes, Slurm, and managed AI inference.A hands-on technical leadership position combining engineering, architecture, people management, and infrastructure-partner engagement.Collaboration with a distributed international team across Europe and India.Exposure to advanced GPU infrastructure, high-performance networking, AI workloads, multi-tenant compute, and production inference systems.Opportunities to shape engineering standards, platform architecture, hiring, and long-term infrastructure capabilities.Occasional travel opportunities to partner sites and team events.
How Jobgether Works
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements.
Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company.
The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer.
This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR).
You may exercise your rights (access, rectification, erasure, objection) at any time.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information.
These tools assist our recruitment team but do not replace human judgment.
Final hiring decisions are ultimately made by humans.
If you would like more information about how your data is processed, please contact us.