Infrastructure Engineer

Pursuit Talent Advisory — United States · Posted ~8 hours ago

Mid

Skills

Infrastructure engineering Deployment Cloud environments On-premise systems Cloud Infrastructure DevOps

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An infrastructure engineering position focused on deploying, operating, and improving software platforms across constrained hardware and cloud environments.

Highlights

High ownership role working on challenging deployment systems across physical and cloud environments with significant autonomy.

Description

ABOUT THE COMPANY We recently partnered with an early growth-stage company that builds AI systems for industrial environments. The platform runs on customer-owned hardware, on-site, without a cloud dependency at runtime. It takes in live data from plant equipment and business systems, turns it into a structured view of the operation, and puts an AI layer on top of it for the people running the floor. The team is small, flat, and AI-native. ICs own their work end to end: scope, build, ship, verify. Autonomy is the default. THE ROLE You own how the platform gets deployed, runs, and stays up, both on constrained on-premise hardware with unreliable connectivity and in the cloud environments used to develop and validate it. This is not a “keep the CI green” job: the deployment target is physical hardware at a customer site, and getting a release onto it repeatably is a genuine engineering problem. You'll be one of the first few engineers on the team. Expect to write application code too. WHY THIS IS INTERESTING Most infrastructure roles are cloud roles. This one isn't. The software has to run on a box someone can walk up to and unplug, in a building with a spotty VPN, next to machines that cost more than the company. That constraint makes almost every decision (deployment, secrets, observability, rollback) more interesting than the cloud version of the same problem. WHAT YOU'LL DO Own the on-premise deployment path. Lightweight Kubernetes on single-node and small multi-node topologies, templated and layered manifests, staged bring-up of a full stack from bare hardware, and clean recovery after hard power loss.Own GitOps. Declarative reconciliation from a Git source of truth, CI-driven image pinning, and revert-based rollback, with a clear, enforced boundary around what the reconciler is and isn't allowed to manage.Own the cloud dev and staging estate. Infrastructure as code for compute, managed databases, container registry, and storage. Enforce the promotion flow from branch to PR to static checks to a real staging environment before anything reaches a shared environment.Own CI/CD. Multi-service image builds, manifest generation, migration gating, and ephemeral environments that come up and tear down without a human babysitting them.Own the GPU inference layer. The company runs its own models on customer hardware rather than calling a provider API. You'll own that: driver and device-plugin prep, model and quantization choices against whatever hardware you're given, and serving configuration (batching, KV cache, context length, parallelism, memory utilization) tuned so a fixed box meets latency targets. Customers buy one appliance, so utilization is a hard budget, not a cost-optimization exercise.Make the appliance operable. Secrets that don't live in Git, telemetry and model-call tracing a support engineer can actually read, and runbooks for the failure modes that will happen at 2 am on a factory floor.Harden it. Industrial networks are a real trust boundary, and some components need elevated access to do their job. Keep the isolation between privileged components, application services, and the AI runtime intact as the system grows.Reduce toil. If a deployment step needs a person, automate it or delegate it to an agent. WHAT WE'RE LOOKING FOR Required: Strong Kubernetes fundamentals, not just kubectl apply, but StatefulSets, storage, networking, ingress, and debugging a cluster that's misbehaving. Bare-metal or single-node experience (k3s, RKE2, MicroK8s) counts more here than managed EKS/AKS.Infrastructure as code in production (Terraform or equivalent), with real state management and multi-environment discipline.Docker beyond the basics: multi-service builds, image size and layer hygiene, registry workflows.CI/CD ownership: you've built and maintained pipelines, not just consumed them.Hosting AI workloads on GPUs. You've run model inference on your own hardware: an inference server (vLLM, TGI, TensorRT-LLM, or similar) behind a real workload, on GPUs you were responsible for. You can reason about VRAM budgets, batching and concurrency, KV cache, quantization tradeoffs, and where throughput vs. latency actually breaks. You know how to read nvidia-smi and a profile and say why a GPU is underutilized.Comfortable in Linux and Bash, and able to read and write Python. Backend services are written in Python; you'll be in the code.You debug from first principles and you write down what you learn. Nice to have: GitOps at scale (Flux or Argo CD).Edge, air-gapped, or on-prem deployments: anywhere you couldn't assume a cloud control plane.Industrial / OT exposure: PLCs, industrial protocols, MQTT, or manufacturing environments generally.Multi-tenant or multi-model GPU sharing (MIG, time-slicing, MPS), and GPU scheduling in Kubernetes.Operating Postgres and time-series databases: migrations, backup/restore, retention.Observability: OpenTelemetry, or LLM tracing and evaluation tooling. HOW THE TEAM WORKS 1-week sprints. Plan Monday, review end of week.Async-first. Write it down over scheduling a meeting. Daily standup is short and time-boxed.AI-native. The team leans on AI tooling across the whole workflow. Mechanical work gets automated or handed to an agent.Bias to ship. Small, frequent changes over big-bang releases.Multiple hats. Lean startup. Everyone stretches beyond their title. This role reports to the CTO and works alongside the engineering ICs.