Infrastructure Engineer
Openrelayinc โ Argentina ยท Posted ~19 hours ago
๐ Log in to save this job, tailor your resume & track your apply process โ 7 days free, no card needed.
Log in to add to target listDescription
Why OpenRelay
OpenRelay is GPU compute and LLM inference for people who ship.
Customers rent GPU machines and call inference APIs; providers plug their hardware into our network and get paid for it.
We run a distributed fleet on real hardware, not a reseller skin over someone else's cloud.
Small team, high bar, and room for A-players to own things that are live in production.
No degree required, show us what you have run
We care about what you can do, not diplomas.
The profile we want most is simple: you have built infrastructure and kept it alive.
Not a tutorial cluster, real machines with real workloads, where a mistake takes something down and you are the one who brings it back.
Show us the systems you have run, what broke, and what you did about it.
Core tech
Terraform, AWS (ECS Fargate, DynamoDB, S3, VPC networking)Nomad on bare-metal GPU hosts (QEMU/KVM, VFIO GPU passthrough)Go (control plane, gateways, node agents)Linux deep enough to argue with: kernel parameters, systemd, networking, driversOverlay networking (Nebula, WireGuard, Tailscale), mTLS PKIVictoriaMetrics + Grafana, GitHub ActionsClaude Code and agentic tooling, dailyWhat you will own
The Terraform that defines both AWS environments, and the discipline of applying it: plan, review, beta, then prodThe provider fleet: enrollment, node identity (mTLS certificates), health reconciliation, and the lifecycle of VMs scheduled onto bare-metal GPU hosts through NomadThe ugly, real parts of GPU hosts: VFIO passthrough, IOMMU groups, kernel flags, why this exact box will not release its GPUThe network paths: customer traffic through our gateways, the control overlay between nodes and the control plane, and the debugging when a path silently degradesMonitoring and the alerts a human actually acts on: dashboards, runbooks, and the discipline of making the next incident boringOps tooling for a fleet you cannot always SSH into: safe remote execution, self-updating agents, automation that fails loudlySkills and qualifications
2+ years running production infrastructure, professional or a serious homelab/fleet you can showTerraform or equivalent infra-as-code on a real cloud accountAdvanced Linux: you have debugged a boot problem, a driver problem, and a network problem, and can tell the storiesComfortable reading and writing Go, or strong in one systems language and willingDocker and at least one orchestrator (Nomad, Kubernetes, or equivalent)Able to own a system end to end: design, rollout, monitoring, incident responseWorking English, written and spokenNice to have: GPU or bare-metal experience, QEMU/KVM or VFIO, PKI/mTLS, Nebula/WireGuard/Tailscale, DynamoDB, on-call experience at a small company.
The video
Required, five minutes or less, in English.
Two parts: your face on camera introducing yourself, and a screen recording walking through the best infrastructure you have built or run and the hardest problem it gave you.
Casual is fine, content matters more than polish.
What to expect
Fast-moving U.S.
startup building GPU compute and inference infrastructureHigh autonomy, high expectations, fair and human cultureDirect influence on product decisionsFully remote, anywhere in ArgentinaCompetitive pay, strong mentorship and growthHow to apply
Apply on our site, not here: https://openrelay.inc/careers/infrastructure-engineer-argentina .
The form takes a few minutes and asks for a short demo video, five minutes or less, in English.
The video is required.
There is an example on the page showing exactly what we are looking for.
We have 144,555 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume โ in under a minute we'll analyze all 144,555 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume