Senior Platform Engineer – AI Cloud

Circle B — Netherlands · Posted ~4 hours ago

Senior Visa History ✓

Skills

platform engineering cloud infrastructure observability SRE reliability engineering Prometheus GPU infrastructure monitoring SLOs cloud

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Own the platform layer above AI compute clusters, making GPU infrastructure observable, measurable, billable, and reliable. Build and operate the observability stack, define and maintain SLOs, support service operations, and develop platform capabilities for demanding cloud workloads.

Highlights

Own reliability and observability for AI cloud infrastructure, build systems that make GPU resources measurable and dependable, and shape platform operations for demanding regulated workloads.

Description

What we do- Circle B builds sustainable IT infrastructure for the AI and cloud era. For over a decade— we have designed and deployed datacenter, edge, and AI/HPC systems on Open Compute Project (OCP) hardware. We are independent, vendor-neutral, and ISO 9001 / 14001 / 27001 certified, with deployments across multiple countries. Our newest initiative is a sovereign EU GPU cloud — operating under full Dutch/EU jurisdiction and beyond the reach of the US CLOUD Act, for regulated European organizations that cannot compromise on where their data lives. The Role- You own everything above the cluster: the observability platform, the GPU metering-to-billing pipeline, the operation of the AI service catalogue, and the SLOs that define reliability. Your job is to make the GPU clusters delivered to you observable, measurable, billable, and reliable. What you will own- Observability & Reliability Build and operate the full observability stack: metrics (Prometheus + long-term store such as VictoriaMetrics/Thanos), logs (Loki + a forwarder), and alerting (Alertmanager)Integrate DCGM GPU telemetry into the pipeline: utilisation, memory, temperature, power, SM activity — per GPU, per tenantSurface GPU and fabric health (DCGM XID errors, NCCL/RDMA, RoCE/link health) from the Infra and Network engineers into a single pane — you own dashboards, alerting, and tenant-impact SLOs; Infra/Network own the fabric remediationBuild the dashboarding capability (Grafana with SSO), ready for tenants at launchDefine SLOs, SLIs, and error budgets; implement burn-rate alertingImplement OpenTelemetry instrumentation across platform services; lay the alerting and runbook foundations for reliable operation laterMetering, Billing & Service Delivery Build the GPU metering pipeline: DCGM + scheduler/namespace state → accurate per-tenant usage events (GPU-hours per tenant) with idempotency and reconciliation → a usage-based billing engine, billing primarily on allocated/reserved GPU-timeDeploy cost-attribution (OpenCost or equivalent) for per-tenant GPU and infrastructure showback/chargebackIntegrate a billing engine (Lago, Stripe Billing, or Metronome): flat fee + storage/egress overages, SEPA B2B direct debitDeploy and operate the AI service catalogue — model-serving runtimes (KServe / NVIDIA NIM Operator, Triton, vLLM), notebooks (JupyterHub), experiment tracking (MLflow), a vector DB (Qdrant), and a distributed-job framework (Ray/KubeRay) — as repeatable, GitOps-driven templates. Operation, not model authoringOperate the air-gapped registry and controlled artifact seeding into the sovereign zone (with Network on the path policy)Expose platform APIs (gateway, rate limiting, integration glue) for the future customer portalBuild the tenant onboarding/offboarding automation — the path from a new tenant to running GPU workloads, end to end What We Are Looking For- Required Skills- 5+ years in platform engineering / SRE, deploying and operating complex service stacks on K8s (operators, CRDs, Helm, scheduling, multi-tenancy / quota)Go and/or Python to a software-engineering standard — control-plane services and automation, not just scriptingProduction observability end-to-end: Prometheus + a long-term store (VictoriaMetrics / Thanos / Mimir), Grafana, Alertmanager, Loki, OpenTelemetry; designing SLIs, SLOs, and error budgetsGPU observability: running the DCGM exporter and surfacing GPU/fabric health (XID, NCCL, RoCE) into dashboards and alerts (surface, not fabric-remediate)Ability to build a GPU metering pipeline: DCGM + scheduler/namespace state → accurate, idempotent per-tenant usage events → a billing engineDeploy and operate (not author) model-serving + ML-platform services on K8s: KServe and/or NIM/Triton/vLLM, plus Ray, JupyterHub, MLflow, Harbor, a vector DBInfrastructure-as-Code + GitOps: Helm, Kustomize, and ArgoCD or Flux in productionStrong Linux fundamentals and production troubleshooting (systemd, container runtimes, REST/gRPC APIs, event-driven pipelines)Self-driven and comfortable with the breadth of a pre-launch platform: context-switching across observability, metering, service delivery, and onboarding Preferred Skills- Prior hands-on with a commercial / OSS usage-based billing engine (Lago, OpenMeter, Metronome, Stripe Billing, Amberflo) and event-driven metering; EU payment integration (SEPA, Mollie, or Stripe EU) a plusNVIDIA AI Enterprise / NIM Operator; Run:ai or KAI Scheduler for GPU quota, fair-share, and multi-tenancyOpenCost / FinOps: GPU cost allocation, chargeback / showbackContainer registry operations in air-gapped environments (Harbor, Quay, or similar)Identity & access integration (OIDC/SAML, SSO, tenant-scoped RBAC) and API gateways (Kong, Envoy, Traefik)EU regulatory awareness: GDPR, AI Act, DORA, NIS2, NEN 7510 in a platform contextTypeScript (billing/admin tooling; interfacing with the future Fullstack portal) Nice to Have Event streaming for telemetry / meteringExperience operating a Kubernetes-native multi-cluster management platformExperience at an AI neocloud, GPU cloud provider, or managed ML platformFamiliarity with NCCL / RDMA fabric troubleshooting Why Join Us Help build a sovereign EU GPU cloud from the ground up.Own a critical platform layer, not just tickets or maintenance.Work on modern AI infrastructure, GPU platforms, Kubernetes, observability, and automation.Join a company with deep experience in OCP, datacenter, AI/HPC, and cloud infrastructure.Build infrastructure for organizations where data location, compliance, and reliability truly matter. Benefits Competitive salary.Pension scheme.Visa sponsorship.Company gym.Modern office in Hoofddorp.Informal and open working culture.Participation in relevant conferences and exhibitions across Europe.Opportunity to develop your skills in a fast-growing technology environment. Our Work Culture Circle B offers an informal working atmosphere with energetic people who enjoy being part of a growing technology company. We have an open management culture and encourage colleagues to contribute to improving our products, services, and processes. If this sounds like a good fit, please send your CV and motivation letter to: surbi@tauruseu.com