Senior AI Engineer, Agentic and RAG Systems

Epam Systems β€” Kyrgyzstan Β· Posted ~2 days ago

Senior Full-time

Skills

Generative AI Agentic AI RAG Python FastAPI LLM systems Vector search System architecture Production engineering LangGraph LangChain LLM MLflow OpenTelemetry Vector databases

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking a Senior AI Engineer to design, build, and operate production-grade GenAI systems. You will develop agentic workflows and end-to-end RAG pipelines, build Python-based LLM services, and own architecture through on-call operations. The role emphasizes reliability, latency, cost, observability, evaluation, and safe deployment at scale rather than research.

Highlights

Hands-on senior engineering role building production GenAI systems with real users and SLAs. Strong end-to-end ownership across agent orchestration, RAG, LLM services, observability, reliability, latency, cost optimization, and safe deployment at scale.

Description

We are seeking a hands-on Senior AI Engineer who designs, builds, and operates production GenAI systems – agentic workflows, RAG pipelines, and LLM-backed services with real users and real SLAs. This is an engineering role, not a research role. The bar is reliability, latency, cost, observability, and safe deployment at scale, with end-to-end ownership from architecture through on-call. Typical workloads include enterprise knowledge platforms, conversational analytics, agentic automation, and LLM-augmented data products. Responsibilities Design agent orchestration (graph/state, conditional routing, tool calling, memory, checkpointing) in LangGraph / LangChain or equivalentBuild production RAG end-to-end: chunking, embeddings, vector stores, hybrid retrieval, reranking, caching, and grounded synthesisOwn Python / FastAPI services – async, SSE streaming, session handling, and structured error contractsInstrument with tracing and evaluation harnesses (MLflow, OpenTelemetry, or equivalent) for accuracy, cost, and regressionShip on Docker + Kubernetes (EKS/AKS/GKE) via CI/CD with test, eval, and canary gatesDrive LLM cost engineering – model routing, prompt optimization, caching, token accounting, and build-vs-buy decisionsApply GenAI safety & governance: hallucination control, prompt-injection defense, PII handling, and HITL where requiredPartner with data engineering on semantic layers and pipelines (PySpark / SQL where applicable) Requirements 5+ years in software engineering, with 2+ years shipping production LLM / agentic systems (not POCs or research)Proficiency in Python and FastAPI (async, REST, SSE)Production expertise in LangChain and LangGraph (or equivalent serious production experience with LlamaIndex, AutoGen, or MCP stacks)Background in production RAG: embeddings, chunking, and hybrid retrieval with reranking and cachingSkills in vector databases such as Pinecone, Weaviate, pgvector, OpenSearch, or Databricks Vector SearchKnowledge of at least one major LLM provider in production – AWS Bedrock (preferred), OpenAI / Azure OpenAI, or Anthropic – with model selection and routing trade-offsCompetency in Kubernetes and Docker in real production environments (EKS/AKS/GKE)Expertise in cloud engineering on AWSFamiliarity with observability and tracing tools (MLflow, LangSmith, OpenTelemetry), evaluation harnesses, and latency/cost ownershipCapability to build CI/CD for AI systems (GitHub Actions, Jenkins, or equivalent) with test/eval gatesStrong written and spoken English (B2 level); able to own design discussions with engineering and business stakeholders independently Nice to have Databricks depth – MLflow (tracking & serving), Vector Search, Unity Catalog / Metric Views, PySpark / SQLExperience with LLM fine-tuning – PEFT, LoRA, QLoRAUnderstanding of MCP servers and tool integrationQualifications in GenAI governance & FinOps – auditability, prompt-injection hardening, PII, and token cost in regulated environmentsBackground in classical ML / DL – NLP, BERT-family, time-series, and CV We offer We connect like-minded people:Delivering innovative solutions to industry leaders, making a global impactEnjoyable working environment, whether it is the vibrant office or the comfort of your own homeOpportunity to work abroad for up to two months per yearRelocation opportunities within our offices in 55+ countriesCorporate and social eventsWe invest in your growth:Leadership development, career advising, soft skills and well-being programsCertifications, including GCP, Azure and AWSUnlimited access to EPAM's internal learning databaseFree English classes with certified teachersWe cover it all:Monetary bonuses for engaging in the referral programMedical & family care packageSix trust days per year (sick leave without a medical certificate)Coverage of psychology sessions of your choiceDiscounts for fitness clubs and sports programsBenefits package (sports activities, a variety of stores and services) EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments. Experience the freedom of remote work from anywhere in Kyrgyzstan, whether it's the comfort of your home or our modern office in Bishkek.