Summary
✨ AI‑Generated
An experienced Infrastructure and Hardware Engineer is sought to design, deploy, and operate modern infrastructure across on-premises and isolated environments. The role combines bare-metal servers, Linux, networking, storage, Kubernetes, data-center operations, and infrastructure automation using tools such as Ansible and Terraform.
Highlights
Experienced infrastructure and hardware engineering contract with hybrid work. The role offers hands-on ownership of bare-metal server fleets, modern data-center networking, storage, Kubernetes, automation, and air-gapped infrastructure environments.
Description
For a berlin based defence client we are currently looking for an experienced Infrastructure / Hardware Engineer who builds servers from scratch and supports their deployment on-site.
The engineer will be responsible for designing, deploying, and maintaining modern infrastructure solutions, including bare-metal server environments, networking, storage systems, and automation frameworks.
The role requires hands-on experience with Linux-based platforms, infrastructure automation, Kubernetes ecosystems, and data center operations.
Projekt tasks:
- Design, size, provision, and operate bare-metal server fleets across on-prem and air-gapped environments (firmware/BIOS/UEFI, BMC via Redfish/IPMI, OS, RAID, kernel and storage tuning) using zero-touch provisioning (PXE/iPXE, MAAS/Metal3/Tinkerbell/Ironic) and automation (Ansible, Terraform, or equivalent— the tooling matters less than the automation mindset).
- Build and run modern data-center networking: L2/L3 design, IP Fabric, BGP and switching, and RDMA fabrics (RoCE/InfiniBand) sized to scale without ripping out the core.
- Engineer resilient, highly available storage (Ceph/Rook, NVMe) with capacity planning and encryption at rest.
- Operate confidently in air-gapped and on-prem environments: offline mirrors and registries, signed artifacts, firmware/driver lifecycle without internet access, and system hardening.
- Support MLOps and inference-serving fundamentals — GPU model serving (Triton/KServe/vLLM), GPU scheduling and sharing, and throughput/latency optimization — in partnership with our SRE and ML teams.
- Plan and run on-site build-outs: rack integration, power budgets, thermal/cooling and UPS sizing, commissioning, capacity planning, runbooks, and operator handover, with SWaP awareness for field sites.
Your Qualifications// Must-Haves:
-Automation & Tooling: Experience with Ansible or Terraform (the key factor is the mindset and way of working rather than the specific tool), as well as bare-metal provisioning solutions such as MAAS, Metal³, Tinkerbell, or Ironic.
Strong knowledge of Git, CI/CD, and scripting with Python and Bash.
-Infrastructure & Experience: Hands-on experience working in on-premises and air-gapped environments.
-Focus Area: We strongly prefer a highly skilled Infrastructure and Network Engineer without GPU expertise over someone with GPU knowledge but limited experience in modern data center and networking infrastructure.
GPU experience is a plus, but priority is given to server hardware sizing, modern networking, data center infrastructure, and a strong DevOps mindset.
-Networking & Storage: Deep expertise in L2/L3 networking, IP Fabric, BGP, and switching, as well as RDMA (RoCE/InfiniBand) and storage technologies such as Ceph/Rook and NVMe, including capacity planning.
Nice To Haves:
- experience with Palantir or Defence
- NVIDIA GPU stack knowledge (drivers, CUDA, GPU Operator, MIG, DCGM) and cross-node GPU interconnect experience (NVLink, InfiniBand, NCCL).
- Kubernetes bare-metal fundamentals — how cluster bring-up and GPU device plugins interact with the underlying hardware, not day-to-day cluster operation.
- Inference optimization (vLLM, TensorRT-LLM, quantization) and familiarity with switch NOS (SONiC/Cumulus).
- Relevant certifications (NVIDIA, Red Hat, CKA/CKS) or field/forward-deployed engineering experience.
- Ü2 / Secret clearance- check
Project language: English
Project start: asap
Project duration: 6 MM - Fulltime, with Option for a Extension
Location: onsite in Berlin and occasionally remote and business travel