Description
๐ฅ๐ฒ๐บ๐ผ๐๐ฒ (๐ณ๐ผ๐ฟ ๐จ๐๐ ๐๐ผ๐บ๐ฝ๐ฎ๐ป๐) ยท ๐๐ ๐ฃ๐ฟ๐ผ๐ฑ๐๐ฐ๐ & ๐ฃ๐น๐ฎ๐๐ณ๐ผ๐ฟ๐บ ยท ๐๐ป๐ด๐ถ๐ป๐ฒ๐ฒ๐ฟ ๐๐๐, ๐๐ฑ ยท ๐ ๐๐๐ฌ๐ต.๐ฌ.๐ญ
Platform engineer who builds the ๐ฒ๐ป๐๐ถ๐ฟ๐ผ๐ป๐บ๐ฒ๐ป๐๐ ๐ถ๐ป ๐๐ต๐ถ๐ฐ๐ต ๐๐ ๐ฎ๐ด๐ฒ๐ป๐๐ ๐ฟ๐๐ป, made for one task, closed to the network except where allowed and the same every time, and who makes the ๐๐ต๐ผ๐น๐ฒ ๐ฝ๐น๐ฎ๐๐ณ๐ผ๐ฟ๐บ ๐ถ๐ป๐๐๐ฎ๐น๐น๐ฎ๐ฏ๐น๐ฒ ๐ถ๐ป๐๐ถ๐ฑ๐ฒ ๐ฎ ๐ฑ๐ฎ๐๐ฎ ๐ฐ๐ฒ๐ป๐๐ฟ๐ฒ ๐๐ถ๐๐ต ๐ป๐ผ ๐ฐ๐ผ๐ป๐ป๐ฒ๐ฐ๐๐ถ๐ผ๐ป ๐๐ผ ๐๐ต๐ฒ ๐ถ๐ป๐๐ฒ๐ฟ๐ป๐ฒ๐, GPUs included.
Greenfield.
THE TECHNICAL CHALLENGE
We build agents on open-weight LLMs that operate web applications through the browser.
Each agent runs in an environment of its own, where it can do no harm.
The environment is made for one task, the same every time and deleted afterwards.
It reaches the network only where allowed, holds no credentials and records every action.
The whole platform, GPUs included, has to be installable in a data centre with no connection to the internet.
The engineering work is a control of outbound traffic that an agent cannot get around.
The same work covers the updates after the first install and the sharing of GPUs between workloads.
This role builds that environment and keeps it reliable.
KEY RESPONSIBILITIES
โข Build the environment an agent and its browser run in: made for one task, isolated, torn down afterwards, the same every time, with environments prepared in advance so that a task starts without waiting
โข Enforce at the network layer what may leave the environment, where the agent cannot change it; keep credentials outside the environment; record every action in an audit log
โข Make the whole platform installable and updatable inside a data centre with no connection to the internet, GPU serving included: registries, signed images, certificates, installation media for updates
โข Provide the environment as a platform with stated guarantees and measured reliability
DESIRED QUALIFICATIONS
โข A sandbox or execution service built on microVMs or a user-space kernel (Firecracker, Kata Containers, gVisor or comparable), with start time, number per host and reset time measured
โข A platform delivered into a site with no connection to the internet and kept up to date there: registries, package mirrors for GPU drivers, signed images, certificate rotation, update media
โข Open-weight models served on a few GPUs of one's own, with the sharing between workloads chosen deliberately and utilisation measured
โข Contributions to Firecracker, Kata Containers, gVisor, Cloud Hypervisor, Cilium, Talos, Zarf or Hauler
EXPECTED QUALIFICATIONS
โข T-shaped: deep in one domain, with working breadth in a neighbouring one
โข Structures a large, incompletely specified problem and drives it to a working result independently; understands that updating an installed platform is harder than the first install
โข Kubernetes run in production with network policy, storage and upgrades, incidents handled; isolation technologies (gVisor, Firecracker, Kata or comparable) used deliberately, with the trade-offs understood; uses AI coding tools daily and verifies their output
โข Outbound control built at the network layer, with secrets kept outside the workload and a written account of how it was tested
โโ
HOW WE WORK
โข Product engineering: we own what we build and run it in production
โข Small teams, two-week cycles, working software at every review
โข AI coding tools are part of the standard workflow
WHAT WE OFFER
โข Founding-team scope
โข AI-augmented engineering environment
โข Access to on-premise Nvidia B200s
โข Flexible work environment
PROCESS
โข Introductory call
โข Technical conversation
โข Practical session; the format is agreed with you
In coding exercises, AI tools are allowed and expected.
No LeetCode.
WHO WE ARE
New product organisation as part of a large semi-government in Abu Dhabi.
International, ex-FAANG team.
Completely greenfield, with a modern tech stack.
โโ
REQUIREMENTS TO BE CONSIDERED
โข Clear written and spoken English
โข 5+ years in platform, infrastructure or reliability engineering; own infrastructure operated in production, with on-call; comfortable in Python or TypeScript and shell
โข Linux, containers and networking understood deeply enough to debug them during an incident; one system the candidate built and can show
โข Bachelor's degree in any field, or self-taught with a track record of open-source contributions
RELATED TECHNOLOGIES AND CONCEPTS
โข Runtime: Kubernetes, Helm, operators, the Kubernetes agent-sandbox project, gVisor, Firecracker, Kata Containers, Cloud Hypervisor, warm pools, snapshot and restore
โข Network and security: network policy, Cilium or comparable, TLS interception, egress proxies, certificate authorities, OpenBao or comparable secrets management, audit logging
โข Installation without internet: Zarf, Hauler, RKE2, Talos, image registries and mirrors, signed images, infrastructure as code, GitOps
โข Serving: NVIDIA GPU Operator, MIG and time-slicing, vLLM and SGLang, object storage (RustFS or comparable)