Forward Deployed Engineer – LLMOps

Systems Limited — Jordan · Posted ~2 hours ago

Senior Full-time

Skills

MLOps platform engineering LLMOps LLM inference infrastructure scaling load balancing caching observability cost optimization incident response model deployment LLM AI agents inference infrastructure CI/CD

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Take ownership of production operations for LLM and agentic AI systems, ensuring reliable serving, efficient scaling, strong observability, and disciplined cost management. You will design rollout and fallback strategies, respond to incidents, and partner with AI engineering teams to make workloads production-ready.

Highlights

Own production operations for modern LLM and agentic workloads, covering serving, scalability, observability, cost control, rollout strategy, and incident response. The role offers significant ownership and close collaboration with AI engineering and architecture teams.

Description

ABOUT Owns production operations for LLM and agentic workloads — serving, cost, and observability for a fundamentally less predictable class of system than classical ML. KEY RESPONSIBILITIES Own production serving and scaling for LLM/agentic workloads (inference infra, load balancing, caching)Monitor and control inference cost — token usage, retry/loop cost, model routing decisionsBuild observability for LLM-specific failure modes: hallucination rate, latency spikes, prompt driftManage model/version rollout strategy (canary releases, fallback models, A/B testing)Own incident response for LLM/agent production issuesPartner with GenAI Engineers and Agentic AI Architects on production-readiness reviewsExplain token-cost dynamics to client finance/business stakeholdersCollaborate closely with GenAI Engineers without needing a hard line between build and runSupport the practice in setting cost governance policy for LLM workloads REQUIREMENTS & SKILLS 4–6 yrs platform/MLOps engineering with hands-on LLM/GenAI production experienceDeep understanding of LLM inference economics — token costs, batching, caching, model routingExperience with LLM observability tooling (tracing, eval pipelines, prompt/version management)Familiarity with multiple model hosting platforms and their cost/performance tradeoffs — Microsoft Azure AI Foundry, AWS Bedrock, and Google Vertex AI, plus self-hosted open-source options (vLLM, TGI) as a good-to-haveExperience building canary/rollback strategies for probabilistic systemsComfortable with the higher unpredictability of agentic workloads vs. classical ML servingCost-conscious communicator — can explain a token-cost blowup to a client's finance stakeholderCollaborates closely with GenAI Engineers without needing a hard line between “build” and “run”Calm under pressure during live incidents affecting client-facing systems