Staff Software Engineer, LLM Infrastructure

Harrisonclarke — United States · Posted ~21 hours ago

Full-time

Skills

Distributed Systems Agentic AI Infrastructure Agent Orchestration Fault-Tolerant Systems Multi-Agent Coordination Low-Latency Systems Durable Memory Production-Scale Systems LLM Infrastructure Agentic AI Multi-Agent Systems Scheduling Consensus Persistent Storage

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Help define the architecture of infrastructure that powers production-grade agentic AI. You will build reliable orchestration for multi-step workflows, coordinate autonomous agents sharing state, enable low-latency tool execution, and develop durable memory systems at scale. This is an early-stage opportunity where experienced engineers can have significant influence over foundational technical decisions.

Highlights

Foundational engineering role focused on building core infrastructure for agentic AI. Engineers can shape architecture and technical direction while solving challenging problems in orchestration, distributed systems, coordination, latency, and durable memory.

Description

Software Engineer - Distributed Systems & Agentic AI Infrastructure The Opportunity We're partnering with a startup being incubated inside a top-tier venture firm. They're not building another AI wrapper. They're building the hard infrastructure underneath agentic AI. The runtime that agents actually execute on. The founding team is small, the technical direction is still being shaped, and the engineers who join now will define the architecture. What You'd Be Building Reliable agent orchestration - fault-tolerant scheduling and execution of multi-step, multi-agent workflows at production scaleStateful multi-agent coordination - consensus, conflict resolution, and shared state across autonomous agents operating in parallelLow-latency tool execution - sub-second invocation layers that let agents interact with external systems without becoming the bottleneckDurable memory at scale - persistent, queryable memory backends that give agents long-horizon context without blowing up latency or cost What We're Looking For Deep distributed systems experience - you've built, operated, or scaled systems at growing companies working with vast volumes of data.Production-grade engineering - replication, consistency, failure recovery, and state management aren't abstract concepts to you - you've shipped themInterest in the agentic AI stack - you don't need to be an ML researcher, but you see the infrastructure gaps in how LLM-powered agents are built todayEarly-stage appetite - you want to be in the room when the architecture is still on a whiteboard, with the backing of a world-class VC behind youNice to have: Experience at the intersection of databases/storage systems and AI infrastructure (vector stores, retrieval systems, context management at scale). Why This Role The "agentic AI" wave is real, but the infrastructure is immature. Most agent frameworks today are single-threaded, stateless, and fragile. This company believes the unlock is at the systems layer - and they're hiring the distributed systems engineers who can build it. With Tier-1 VC incubation, you get the resources and network of a top firm with the autonomy and upside of a founding team.