Infrastructure / Hardware Engineer

Hays — Germany · Posted ~11 hours ago

Senior Contract Hybrid

Skills

Linux Bare-metal servers Infrastructure automation Kubernetes Data center operations Networking Storage systems Ansible Terraform PXE/iPXE MAAS Metal3 Tinkerbell Ironic BGP L2/L3 networking Python PXE iPXE Redfish IPMI

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An experienced Infrastructure and Hardware Engineer is sought to design, deploy, and operate modern infrastructure across on-premises and isolated environments. The role combines bare-metal servers, Linux, networking, storage, Kubernetes, data-center operations, and infrastructure automation using tools such as Ansible and Terraform.

Highlights

Experienced infrastructure and hardware engineering contract with hybrid work. The role offers hands-on ownership of bare-metal server fleets, modern data-center networking, storage, Kubernetes, automation, and air-gapped infrastructure environments.

Description

For a berlin based defence client we are currently looking for an experienced Infrastructure / Hardware Engineer who builds servers from scratch and supports their deployment on-site. The engineer will be responsible for designing, deploying, and maintaining modern infrastructure solutions, including bare-metal server environments, networking, storage systems, and automation frameworks. The role requires hands-on experience with Linux-based platforms, infrastructure automation, Kubernetes ecosystems, and data center operations. Projekt tasks: - Design, size, provision, and operate bare-metal server fleets across on-prem and air-gapped environments (firmware/BIOS/UEFI, BMC via Redfish/IPMI, OS, RAID, kernel and storage tuning) using zero-touch provisioning (PXE/iPXE, MAAS/Metal3/Tinkerbell/Ironic) and automation (Ansible, Terraform, or equivalent— the tooling matters less than the automation mindset). - Build and run modern data-center networking: L2/L3 design, IP Fabric, BGP and switching, and RDMA fabrics (RoCE/InfiniBand) sized to scale without ripping out the core. - Engineer resilient, highly available storage (Ceph/Rook, NVMe) with capacity planning and encryption at rest. - Operate confidently in air-gapped and on-prem environments: offline mirrors and registries, signed artifacts, firmware/driver lifecycle without internet access, and system hardening. - Support MLOps and inference-serving fundamentals — GPU model serving (Triton/KServe/vLLM), GPU scheduling and sharing, and throughput/latency optimization — in partnership with our SRE and ML teams. - Plan and run on-site build-outs: rack integration, power budgets, thermal/cooling and UPS sizing, commissioning, capacity planning, runbooks, and operator handover, with SWaP awareness for field sites. Your Qualifications// Must-Haves: -Automation & Tooling: Experience with Ansible or Terraform (the key factor is the mindset and way of working rather than the specific tool), as well as bare-metal provisioning solutions such as MAAS, Metal³, Tinkerbell, or Ironic. Strong knowledge of Git, CI/CD, and scripting with Python and Bash. -Infrastructure & Experience: Hands-on experience working in on-premises and air-gapped environments. -Focus Area: We strongly prefer a highly skilled Infrastructure and Network Engineer without GPU expertise over someone with GPU knowledge but limited experience in modern data center and networking infrastructure. GPU experience is a plus, but priority is given to server hardware sizing, modern networking, data center infrastructure, and a strong DevOps mindset. -Networking & Storage: Deep expertise in L2/L3 networking, IP Fabric, BGP, and switching, as well as RDMA (RoCE/InfiniBand) and storage technologies such as Ceph/Rook and NVMe, including capacity planning. Nice To Haves: - experience with Palantir or Defence - NVIDIA GPU stack knowledge (drivers, CUDA, GPU Operator, MIG, DCGM) and cross-node GPU interconnect experience (NVLink, InfiniBand, NCCL). - Kubernetes bare-metal fundamentals — how cluster bring-up and GPU device plugins interact with the underlying hardware, not day-to-day cluster operation. - Inference optimization (vLLM, TensorRT-LLM, quantization) and familiarity with switch NOS (SONiC/Cumulus). - Relevant certifications (NVIDIA, Red Hat, CKA/CKS) or field/forward-deployed engineering experience. - Ü2 / Secret clearance- check Project language: English Project start: asap Project duration: 6 MM - Fulltime, with Option for a Extension Location: onsite in Berlin and occasionally remote and business travel