Platform Engineer, Sandboxing and Air-Gapped Infrastructure

Stealth Startup Community โ€” United Arab Emirates ยท Posted ~7 hours ago

Senior Full-time Remote

Skills

Platform engineering Infrastructure automation Networking Linux GPU infrastructure GPU Infrastructure Automation LLM

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary โœจ AIโ€‘Generated

An innovative AI-focused organization is hiring a platform engineer to build secure execution environments for autonomous systems. The role focuses on isolated infrastructure, deployment automation, networking controls, and reliable large-scale platforms.

Highlights

Work on advanced AI infrastructure challenges involving secure environments, automation, and large-scale deployment systems.

Description

๐—ฅ๐—ฒ๐—บ๐—ผ๐˜๐—ฒ (๐—ณ๐—ผ๐—ฟ ๐—จ๐—”๐—˜ ๐—–๐—ผ๐—บ๐—ฝ๐—ฎ๐—ป๐˜†) ยท ๐—”๐—œ ๐—ฃ๐—ฟ๐—ผ๐—ฑ๐˜‚๐—ฐ๐˜ & ๐—ฃ๐—น๐—ฎ๐˜๐—ณ๐—ผ๐—ฟ๐—บ ยท ๐—˜๐—ป๐—ด๐—ถ๐—ป๐—ฒ๐—ฒ๐—ฟ ๐—œ๐—œ๐—œ, ๐—Ÿ๐Ÿฑ ยท ๐— ๐—๐——๐Ÿฌ๐Ÿต.๐Ÿฌ.๐Ÿญ Platform engineer who builds the ๐—ฒ๐—ป๐˜ƒ๐—ถ๐—ฟ๐—ผ๐—ป๐—บ๐—ฒ๐—ป๐˜๐˜€ ๐—ถ๐—ป ๐˜„๐—ต๐—ถ๐—ฐ๐—ต ๐—”๐—œ ๐—ฎ๐—ด๐—ฒ๐—ป๐˜๐˜€ ๐—ฟ๐˜‚๐—ป, made for one task, closed to the network except where allowed and the same every time, and who makes the ๐˜„๐—ต๐—ผ๐—น๐—ฒ ๐—ฝ๐—น๐—ฎ๐˜๐—ณ๐—ผ๐—ฟ๐—บ ๐—ถ๐—ป๐˜€๐˜๐—ฎ๐—น๐—น๐—ฎ๐—ฏ๐—น๐—ฒ ๐—ถ๐—ป๐˜€๐—ถ๐—ฑ๐—ฒ ๐—ฎ ๐—ฑ๐—ฎ๐˜๐—ฎ ๐—ฐ๐—ฒ๐—ป๐˜๐—ฟ๐—ฒ ๐˜„๐—ถ๐˜๐—ต ๐—ป๐—ผ ๐—ฐ๐—ผ๐—ป๐—ป๐—ฒ๐—ฐ๐˜๐—ถ๐—ผ๐—ป ๐˜๐—ผ ๐˜๐—ต๐—ฒ ๐—ถ๐—ป๐˜๐—ฒ๐—ฟ๐—ป๐—ฒ๐˜, GPUs included. Greenfield. THE TECHNICAL CHALLENGE We build agents on open-weight LLMs that operate web applications through the browser. Each agent runs in an environment of its own, where it can do no harm. The environment is made for one task, the same every time and deleted afterwards. It reaches the network only where allowed, holds no credentials and records every action. The whole platform, GPUs included, has to be installable in a data centre with no connection to the internet. The engineering work is a control of outbound traffic that an agent cannot get around. The same work covers the updates after the first install and the sharing of GPUs between workloads. This role builds that environment and keeps it reliable. KEY RESPONSIBILITIES โ€ข Build the environment an agent and its browser run in: made for one task, isolated, torn down afterwards, the same every time, with environments prepared in advance so that a task starts without waiting โ€ข Enforce at the network layer what may leave the environment, where the agent cannot change it; keep credentials outside the environment; record every action in an audit log โ€ข Make the whole platform installable and updatable inside a data centre with no connection to the internet, GPU serving included: registries, signed images, certificates, installation media for updates โ€ข Provide the environment as a platform with stated guarantees and measured reliability DESIRED QUALIFICATIONS โ€ข A sandbox or execution service built on microVMs or a user-space kernel (Firecracker, Kata Containers, gVisor or comparable), with start time, number per host and reset time measured โ€ข A platform delivered into a site with no connection to the internet and kept up to date there: registries, package mirrors for GPU drivers, signed images, certificate rotation, update media โ€ข Open-weight models served on a few GPUs of one's own, with the sharing between workloads chosen deliberately and utilisation measured โ€ข Contributions to Firecracker, Kata Containers, gVisor, Cloud Hypervisor, Cilium, Talos, Zarf or Hauler EXPECTED QUALIFICATIONS โ€ข T-shaped: deep in one domain, with working breadth in a neighbouring one โ€ข Structures a large, incompletely specified problem and drives it to a working result independently; understands that updating an installed platform is harder than the first install โ€ข Kubernetes run in production with network policy, storage and upgrades, incidents handled; isolation technologies (gVisor, Firecracker, Kata or comparable) used deliberately, with the trade-offs understood; uses AI coding tools daily and verifies their output โ€ข Outbound control built at the network layer, with secrets kept outside the workload and a written account of how it was tested โ€”โ€” HOW WE WORK โ€ข Product engineering: we own what we build and run it in production โ€ข Small teams, two-week cycles, working software at every review โ€ข AI coding tools are part of the standard workflow WHAT WE OFFER โ€ข Founding-team scope โ€ข AI-augmented engineering environment โ€ข Access to on-premise Nvidia B200s โ€ข Flexible work environment PROCESS โ€ข Introductory call โ€ข Technical conversation โ€ข Practical session; the format is agreed with you In coding exercises, AI tools are allowed and expected. No LeetCode. WHO WE ARE New product organisation as part of a large semi-government in Abu Dhabi. International, ex-FAANG team. Completely greenfield, with a modern tech stack. โ€”โ€” REQUIREMENTS TO BE CONSIDERED โ€ข Clear written and spoken English โ€ข 5+ years in platform, infrastructure or reliability engineering; own infrastructure operated in production, with on-call; comfortable in Python or TypeScript and shell โ€ข Linux, containers and networking understood deeply enough to debug them during an incident; one system the candidate built and can show โ€ข Bachelor's degree in any field, or self-taught with a track record of open-source contributions RELATED TECHNOLOGIES AND CONCEPTS โ€ข Runtime: Kubernetes, Helm, operators, the Kubernetes agent-sandbox project, gVisor, Firecracker, Kata Containers, Cloud Hypervisor, warm pools, snapshot and restore โ€ข Network and security: network policy, Cilium or comparable, TLS interception, egress proxies, certificate authorities, OpenBao or comparable secrets management, audit logging โ€ข Installation without internet: Zarf, Hauler, RKE2, Talos, image registries and mirrors, signed images, infrastructure as code, GitOps โ€ข Serving: NVIDIA GPU Operator, MIG and time-slicing, vLLM and SGLang, object storage (RustFS or comparable)