Gen AI Engineer (LLM &RAG)

Tat It Technolgies — Kuwait · Posted ~5 days ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

Urgent requirement for Gen AI Engineer (LLM &RAG) in Banking Domain is required for our banking clients in Kuwait Strong in Azure OpenAI Service, Azure AI Search, Azure AI Content Safety, Azure AI Foundry (prompt flow/evaluation), Copilot Studio..-Must Hands on Python with LangChain, LlamaIndex, Semantic Kernel; Hugging Face Transformers; open embedding models..--MustHands on pgvector, Qdrant, Chroma, Milvus; evaluation tools like RAGAS, DeepEval.-MustGuardrails (Azure AI Content Safety, Llama Guard, NeMo Guardrails), retrieval‑time access control, PII/PCI masking, audit trails.-Must Role Overview :The builder of the intelligence behind every GenAI and Agentic AI use case . This role designs and implements the retrieval-augmented generation pipelines — chunking, embeddings, retrieval, re-ranking, prompting and grounding making sure AI return accurate, sourced, role-aware answers. Owns prompt design, guardrails, evaluation harnesses and (where needed) fine-tuning or domain-tuning of models. Ensures answers are traceable to source documents, a hard requirement for compliance and trust in a bank. Role Experience 4+ years in ML/NLP or software engineering, with 1.5+ years hands-on building LLM / RAG applications. Proven delivery of a RAG system: document ingestion, embeddings, vector search, prompt orchestration and evaluation. Experience with hallucination control, grounding, citation of sources, and structured evaluation of GenAI quality. Familiarity with fine-tuning / domain adaptation and with prompt-injection and jailbreak defense. Core Skills & Capabilities :Microsoft stack (reference build): Azure OpenAI Service, Azure AI Search, Azure AI Content Safety, prompt flow and evaluation in Azure AI Foundry. Copilot Studio for conversational experiences on the KIB intranet. Open-source / custom stack:Python with LangChain / LlamaIndex / Semantic Kernel; Hugging Face Transformers and open embedding models.Open vector databases (pgvector, Qdrant, Chroma, Milvus); open evaluation tooling (RAGAS, DeepEval); local serving with Ollama / vLLM.Open and fine-tunable models (Llama, Mistral) with LoRA / PEFT techniques. Security, access & data management:Implements guardrails and content filtering (Azure AI Content Safety or open equivalents such as Llama Guard / NeMo Guardrails) against prompt injection, data leakage and unsafe output.Enforces retrieval-time access control so a user only ever sees content their identity is entitled to (document-level and row-level security passed from Entra ID / the source system).Prevents sensitive data (PII, client, PCI) from being logged or sent to models outside the approved boundary; applies masking and redaction in the pipeline.Builds source-citation and audit trails so every answer can be traced to approved material — essential for regulatory defensibility. Skills: llm,rag,genai