Lead AI Engineer

Compunnel Software Group — United States · Posted ~1 day ago

Lead

Skills

Python REST APIs Generative AI Agentic AI conversational AI chatbot development multi-turn context session memory human handoff multichannel delivery response streaming semantic caching response caching retry handling rate-limit handling autoscaling load scaling Conversational AI

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A lead AI engineering role for an experienced software engineer with deep hands-on expertise in production Python and REST APIs. The role focuses on building and scaling generative and conversational AI applications, including multi-turn interactions, memory, human handoff, streaming, caching, resilience, autoscaling, and high-concurrency production systems.

Highlights

Lead production-grade generative and conversational AI initiatives with substantial ownership, high-scale chatbot experience, and exposure to fast-paced enterprise delivery and stakeholder collaboration.

Description

10+ Years of experience in Software Engineer and 3 years of experience in Gen AI/Agentic AI.Enterprise delivery: built and shipped at least two production-grade GenAI or conversational AI applications, with ownership of significant components, including one delivered on a tight timeline.Embedded, fast-paced delivery: worked as part of consulting or client teams; comfortable with ambiguous scope, weekly demos and stakeholder exposure.Hands-on engineering: strong production Python and REST APIsConversational AI: built production chatbots or virtual assistants with multi-turn context, intent handling, session memory, human handoff and multichannel delivery (web, mobile)Chatbot scale: worked on chatbots running at high concurrency in production. Candidates should state peak concurrent users, p95 latency and their part in scaling it.Scale engineering: response streaming, semantic and response caching, retries and rate-limit handling, provisioned throughput, autoscaling and load testing.RAG: implemented retrieval pipelines with vector stores (OpenSearch, pgvector, Aurora, Pinecone or similar) and measured retrieval quality and groundedness.AWS Bedrock (essential): model invocation and selection, Knowledge Bases, Guardrails, Agents and model evaluation.Bedrock AgentCore (essential): hands-on with Runtime, Memory, Gateway, Identity and Observability to deploy and operate agents.MCP: built MCP servers and clients, including authentication and authorisation for tools.A2A and multi-agent: built multi-agent workflows using A2A and frameworks such as Strands Agents, LangGraph or CrewAI.AWS foundations: Lambda, API Gateway, ECS or EKS, DynamoDB, S3, IAM, VPC networking and CloudWatch.Responsible AI: guardrails, hallucination control, prompt-injection defence and PII handling in regulated environments.LLM observability and evaluation: Langfuse, Ragas, Bedrock evaluations or similar tools to trace, monitor and evaluate LLM applications.