Forward Deployed Engineer - LLMOps

Systems Limited — Jordan · Posted ~2 hours ago

Mid

Skills

LLMOps MLOps LLM serving scaling load balancing caching inference cost management observability model deployment incident response LLM agentic AI

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Lead production operations for LLM and agentic AI workloads, managing inference infrastructure, scaling, load balancing, caching, observability, and model rollout strategies. Monitor token-related costs, respond to incidents, and partner with AI engineering teams to improve production readiness.

Highlights

Own production operations for LLM and agentic workloads, focusing on scalability, observability, cost control, resilient deployments, and production readiness.

Description

ABOUT Owns production operations for LLM and agentic workloads — serving, cost, and observability for a fundamentally less predictable class of system than classical ML. KEY RESPONSIBILITIES Own production serving and scaling for LLM/agentic workloads (inference infra, load balancing, caching)Monitor and control inference cost — token usage, retry/loop cost, model routing decisionsBuild observability for LLM-specific failure modes: hallucination rate, latency spikes, prompt driftManage model/version rollout strategy (canary releases, fallback models, A/B testing)Own incident response for LLM/agent production issuesPartner with GenAI Engineers and Agentic AI Architects on production-readiness reviewsExplain token-cost dynamics to client finance/business stakeholdersCollaborate closely with GenAI Engineers without needing a hard line between build and runSupport the practice in setting cost governance policy for LLM workloads REQUIREMENTS & SKILLS 4–6 yrs platform/MLOps engineering with hands-on LLM/GenAI production experienceDeep understanding of LLM inference economics — token costs, batching, caching, model routingExperience with LLM observability tooling (tracing, eval pipelines, prompt/version management)Familiarity with multiple model hosting platforms and their cost/performance tradeoffs — Microsoft Azure AI Foundry, AWS Bedrock, and Google Vertex AI, plus self-hosted open-source options (vLLM, TGI) as a good-to-haveExperience building canary/rollback strategies for probabilistic systemsComfortable with the higher unpredictability of agentic workloads vs. classical ML servingCost-conscious communicator — can explain a token-cost blowup to a client's finance stakeholderCollaborates closely with GenAI Engineers without needing a hard line between “build” and “run”Calm under pressure during live incidents affecting client-facing systems