Summary
✨ AI‑Generated
Join a small senior engineering team building an enterprise AI platform that operates across cloud, on-premises, and isolated environments. You will own the target architecture and platform, covering Kubernetes, infrastructure as code, packaging and delivery, GPU model serving, observability, security, and cost. You will also lead platform engineering and mentor a junior engineer.
Highlights
High-impact platform ownership in a small senior team, with substantial responsibility for architecture, infrastructure, AI workloads, security, delivery, and mentoring.
Description
We are building an enterprise AI platform in which AI agents operate business applications through their user interface, the way a person does.
It runs on Kubernetes across Microsoft Azure, on-premises, and air-gapped environments, delivered through GitOps.
Because it is deployed inside customers' own perimeters, it must install and run with no dependency on cloud services in the critical path.
We are a small, senior team working in two-week sprints towards a first production release in December 2026.
You will own the target architecture, hosting, and platform: infrastructure as code, cluster management, packaging and delivery, GPU model serving, sandboxed browser workloads, observability, security, and cost.
You will lead platform engineering and mentor a junior DevOps engineer.
What you will do
Define and own the target architecture and hosting.
Produce the architecture and security documentation needed for internal approvals and security clearance.Design and run Kubernetes platforms across Azure (AKS) and on-premises clusters (e.g.
RKE2, OpenShift, Rancher, or upstream Kubernetes), with consistent tooling and policies across both.Package the platform as Helm charts that install into customer environments, including air-gapped installs: offline registries (e.g.
Harbor), image and chart mirroring, and signed release bundles.Build GitOps delivery with Argo CD or Flux: repository structure, environment promotion, progressive delivery, and drift detection.Define all infrastructure as code with Terraform / OpenTofu and/or Bicep.Run GPU workloads: GPU node pools and scheduling (NVIDIA GPU Operator), and self-hosted serving of open-weight multimodal models (e.g.
vLLM, SGLang) behind an OpenAI-compatible gateway.Run fleets of sandboxed headless browsers for AI agents at scale, with strong isolation (e.g.
gVisor, Kata Containers, network policies) and controlled egress.Run stateful services on Kubernetes without managed cloud equivalents: PostgreSQL (e.g.
CloudNativePG), S3-compatible object storage, and secrets management (e.g.
OpenBao / Vault).Own platform security: policy as code (OPA Gatekeeper / Kyverno), image scanning, SBOMs, and supply-chain security.Build observability: metrics, logs, and traces (Prometheus, Grafana, Loki, OpenTelemetry), LLM tracing (e.g.
Langfuse), SLOs, and alerting.Own reliability: backup and disaster recovery, capacity planning, incident response, and blameless post-mortems.Lead and mentor the DevOps Engineer – CI/CD & Environments.What we are looking for
Must have
10+ years in software, infrastructure, DevOps, SRE, or platform engineering, including production Kubernetes at scale.Experience defining target architectures and getting them through security and architecture reviews.Deep hands-on Azure experience (AKS, networking, identity, Key Vault, Monitor, ACR).Experience running on-premises Kubernetes and connecting cloud and on-premises environments.Experience delivering software into air-gapped or highly restricted environments.Production experience with GitOps (Argo CD or Flux), and writing and maintaining Helm charts.Strong infrastructure as code skills with Terraform / OpenTofu and/or Bicep.Solid Linux, networking (TCP/IP, DNS, TLS, load balancing), and scripting (Bash plus Python or Go).A security-first mindset: least privilege, secrets management, policy as code, and supply-chain security.A track record of technical leadership: owning architecture decisions, mentoring engineers, and raising engineering standards across a team.Nice to have
GPU workloads on Kubernetes and LLM serving (vLLM, SGLang, Triton, KServe).Running headless browser fleets (Chromium, Playwright) or other untrusted workloads in sandboxes.Service mesh (Istio, Linkerd, Cilium) and eBPF-based networking.CKA / CKS and/or Azure certifications (AZ-104, AZ-305, AZ-400).Experience in regulated or data-sovereign environments.What we will assess
System design: a platform that runs on Azure and in an air-gapped on-premises site, covering packaging, networking, identity, security, GPU serving, observability, and disaster recovery.Technical exercise: given a service, design its GitOps repository layout, IaC, and promotion flow across environments.Incident scenario: troubleshoot a realistic production failure.Leadership: how you would guide and grow a junior engineer.Why join
Own the architecture of a platform that installs and runs anywhere, from Azure to fully air-gapped sites.Run infrastructure for production AI workloads: GPUs, self-hosted models, and browser sandboxes.Lead platform engineering from day one.