Lead AI Engineer, Agentic and RAG Systems

Epam Systems β€” Kyrgyzstan Β· Posted ~2 days ago

Lead

Skills

Generative AI Agentic AI RAG LLM-backed services LangGraph Agent orchestration Production AI systems Cloud architecture Observability Reliability engineering Technical leadership GenAI LLMs

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A Lead AI Engineer is sought to architect, build, and operate production-grade GenAI platforms serving real users and demanding reliability targets. The role covers agent orchestration, RAG pipelines, LLM services, distributed architecture, observability, cost and latency optimization, and safe deployment. You will lead engineers across multiple workstreams while owning technical direction, architecture, delivery quality, and operational readiness.

Highlights

Engineering leadership role with end-to-end ownership of production GenAI platforms, agentic workflows, RAG pipelines, and LLM services. Offers the opportunity to define technical standards, lead multiple workstreams, and optimize reliability, latency, cost, observability, and safe deployment at scale.

Description

We are looking for a seasoned Lead AI Engineer who architects, builds, and operates production GenAI platforms – agentic workflows, RAG pipelines, and LLM-backed services with real users and real SLAs – while leading engineers and setting the technical direction across multiple workstreams. This is an engineering leadership role, not a research role. The bar is reliability, latency, cost, observability, and safe deployment at scale, with end-to-end ownership from architecture through on-call, and accountability for the technical quality and delivery of the team. Typical workloads include enterprise knowledge platforms, conversational analytics, agentic automation, and LLM-augmented data products. Responsibilities Own the end-to-end architecture of GenAI platforms across multiple services and teams, defining standards, patterns, and reference implementationsLead the design of agent orchestration (graph/state, conditional routing, tool calling, memory, checkpointing) in LangGraph / LangChain or equivalent, and set best practices for the teamArchitect production RAG end-to-end: chunking, embeddings, vector stores, hybrid retrieval, reranking, caching, and grounded synthesis – and mentor engineers in building itDrive the design and delivery of Python / FastAPI services – async, SSE streaming, session handling, and structured error contracts – establishing service templates and conventionsDefine the observability and evaluation strategy (MLflow, OpenTelemetry, or equivalent) for accuracy, cost, and regression across the platformOwn the deployment platform on Docker + Kubernetes (EKS/AKS/GKE) with CI/CD, test, eval, and canary gates – setting release standards for AI systemsLead LLM cost engineering strategy – model routing, prompt optimization, caching, token accounting, and build-vs-buy decisions at portfolio levelEstablish GenAI safety & governance practices: hallucination control, prompt-injection defense, PII handling, and HITL where requiredPartner with data engineering leadership on semantic layers and pipelines (PySpark / SQL where applicable), and align roadmaps across teamsMentor and grow senior and mid-level engineers through design reviews, pairing, and technical coaching; conduct hiring and technical interviewsRepresent engineering in conversations with clients, product, and executive stakeholders; translate business goals into technical strategy and delivery plans Requirements 6+ years in software engineering, with 3+ years shipping production LLM / agentic systems (not POCs or research)1+ years of experience leading engineers or technical workstreamsProven track record of owning architecture for multi-service GenAI or distributed systems in productionExpert-level proficiency in Python and FastAPI (async, REST, SSE)Deep production expertise in LangChain and LangGraph (or equivalent serious production experience with LlamaIndex, AutoGen, or MCP stacks)Strong background in production RAG: embeddings, chunking, and hybrid retrieval with reranking and caching – with the ability to define standards across teamsAdvanced skills in vector databases such as Pinecone, Weaviate, pgvector, OpenSearch, or Databricks Vector SearchHands-on production experience with at least one major LLM provider – AWS Bedrock (preferred), OpenAI / Azure OpenAI, or Anthropic – including model selection, routing trade-offs, and multi-provider strategyStrong competency in Kubernetes and Docker in real production environments (EKS/AKS/GKE), including platform-level decisionsDeep expertise in cloud engineering on AWS, including cost, security, and scalability trade-offsSolid command of observability and tracing tools (MLflow, LangSmith, OpenTelemetry), evaluation harnesses, and latency/cost ownership at platform scaleExperience designing and owning CI/CD for AI systems (GitHub Actions, Jenkins, or equivalent) with test/eval gatesDemonstrated experience mentoring engineers, leading design reviews, and driving technical decisions across teamsStrong written and spoken English (B2+ level); able to lead design discussions, present to senior stakeholders, and influence technical direction with clients and executives Nice to have Databricks depth – MLflow (tracking & serving), Vector Search, Unity Catalog / Metric Views, PySpark / SQLExperience with LLM fine-tuning – PEFT, LoRA, QLoRA – and the ability to guide build-vs-fine-tune-vs-prompt decisionsStrong understanding of MCP servers and tool integration patternsExpertise in GenAI governance & FinOps – auditability, prompt-injection hardening, PII, and token cost in regulated environmentsBackground in classical ML / DL – NLP, BERT-family, time-series, and CV We offer We connect like-minded people:Delivering innovative solutions to industry leaders, making a global impactEnjoyable working environment, whether it is the vibrant office or the comfort of your own homeOpportunity to work abroad for up to two months per yearRelocation opportunities within our offices in 55+ countriesCorporate and social eventsWe invest in your growth:Leadership development, career advising, soft skills and well-being programsCertifications, including GCP, Azure and AWSUnlimited access to EPAM's internal learning databaseFree English classes with certified teachersWe cover it all:Monetary bonuses for engaging in the referral programMedical & family care packageSix trust days per year (sick leave without a medical certificate)Coverage of psychology sessions of your choiceDiscounts for fitness clubs and sports programsBenefits package (sports activities, a variety of stores and services) EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments. Experience the freedom of remote work from anywhere in Kyrgyzstan, whether it's the comfort of your home or our modern office in Bishkek.