Staff Distributed Systems Engineer

Harrisonclarke — United States · Posted ~21 hours ago

Lead

Skills

Distributed systems Python Go Backend development Low-latency systems High-throughput systems Data orchestration Compute orchestration Reliability engineering Backend services Orchestration APIs

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A Staff-level distributed systems engineering opportunity for a strong hands-on coder who enjoys solving difficult, open-ended infrastructure problems. You will design and operate reliable distributed systems, build core backend services in Go and Python, develop orchestration and control-plane capabilities, and help create the reliability foundations for production AI workloads.

Highlights

Hands-on staff-level role tackling complex distributed-systems challenges, with exposure to low-latency services, high-throughput pipelines, scalable orchestration, and production reliability in an advanced AI engineering environment.

Description

We’re working with a high-growth, tier-1 VC-backed startup in the AI code generation space, hiring a Staff Distributed Systems Engineer to help design and build the core systems underpinning a next-generation AI product. This is a hands-on role for someone who enjoys bleeding-edge tech and thrives on complex, unsolved engineering problems - the kind where you’ll be building new primitives, not just wiring together existing ones. The Role You’ll work on the heart of the platform: low-latency services, high-throughput pipelines, scalable data and compute orchestration, and the reliability foundations required to run production-grade systems at pace. They’re specifically looking for a strong coder with production experience in Python and Go. What You’ll Be Doing Design, build, and operate distributed systems that are reliable under real-world load and failure modes.Develop core backend services in Go and Python (service frameworks, orchestration, control planes, APIs).Solve problems across consistency, concurrency, throughput, latency, resiliency, backpressure, and graceful degradation.Build systems for job scheduling / workload orchestration and efficient compute utilisation (including demanding AI workloads).Improve observability and debugging for complex systems: tracing, metrics, structured logging, and profiling.Lead architectural decisions: data flows, service boundaries, state management, and scaling strategies.Set engineering standards and mentor others, while remaining deeply technical and hands-on. What They’re Looking For Strong experience building production distributed systems.Excellent coding skills in Go and Python.Deep understanding of:Distributed systems fundamentals (consensus concepts, replication, consistency trade-offs)Networking & performance (RPC patterns, load balancing, latency analysis)Reliability engineering (timeouts, retries, idempotency, circuit breaking, chaos/failure testing)Experience scaling services and data flows in cloud environments (AWS/GCP/Azure).Comfortable working in ambiguity and moving quickly without compromising core quality. Nice-to-Haves Experience with high-scale systems: streaming, queues, event-driven architectures, or large-scale caching.Familiarity with Kubernetes and cloud-native infrastructure (helpful, but not the focus).Experience with ML/AI infrastructure or compute-heavy systems (e.g., GPU scheduling, batch/online hybrid workloads).