Summary
✨ AI‑Generated
Lead quality assurance for AI-native applications by building automated functional tests and evaluation systems for model and agent outputs. You will detect quality drift, assess adversarial behavior, integrate evaluations into CI/CD, establish release quality gates, and develop reusable AI evaluation practices.
Highlights
Own quality assurance for AI-native applications, combining traditional automated testing with model evaluation for accuracy, consistency, drift, and hallucination. The role provides broad ownership of evaluation frameworks, quality gates, and AI-specific testing practices.
Description
ABOUT:
Owns quality for AI-native applications — functional testing plus the AI-specific evaluation (accuracy, drift, hallucination).
KEY RESPONSIBILITIES
Build test plans and automation for AI-native application features (functional + AI-specific)Design evaluation harnesses for model/agent outputs — accuracy, consistency, hallucination rateRun regression testing across model/prompt/config changes to catch silent quality driftRed-team AI features for edge cases and adversarial inputs where relevantBuild automated eval pipelines integrated into CI/CDPartner with AI Architects to define testability requirements before build startsOwn the quality gate before any AI feature ships to productionCommunicate quality risk to delivery leadership in terms they can act onTrain delivery teams on AI-specific testing practicesOwn the evals framework for the practice — golden datasets, scoring rubrics, LLM-as-judge calibration, and versioned benchmarks per use caseDefine eval acceptance thresholds per engagement and gate releases on themBuild eval engineering tooling — dataset curation, trace capture, offline/online eval runs, and dashboards delivery teams can readInstrument production evals and drift monitoring, feeding failures back into the golden datasets
REQUIREMENTS & SKILLS
5–9 yrs QA/test engineering, with 2+ yrs testing AI/ML-powered features specificallyStrong test automation skills (Python-based frameworks, CI/CD integration)Understands AI-specific failure modes — hallucination, bias, drift, non-determinism — and designs tests for themStatistically literate enough to interpret model evaluation metrics, not just pass/fail resultsFamiliarity with red-teaming methodologies for AI systemsClear, assertive communicator — willing to block a release over a quality concernDetail-oriented and methodical under delivery-timeline pressureCollaborative but independent — doesn't rubber-stamp under delivery pressureExplains quality risk in business-impact terms, not just technical jargonHands-on evals engineering — builds and maintains eval suites with frameworks such as OpenAI Evals, Ragas, DeepEval, LangSmith, Azure AI Foundry evaluationsDesigns golden datasets and rubrics, and calibrates LLM-as-judge scoring against human reviewUnderstands RAG and agent eval metrics — groundedness, retrieval precision/recall, task completion, tool-call correctness, cost/latencyExperience wiring evals and drift monitoring into CI/CD and production observability