Forward Deployed Engineer - AI Assurance

Systems Limited β€” Jordan Β· Posted ~2 hours ago

Senior Full-time

Skills

AI testing Functional testing AI evaluation Test automation Regression testing CI/CD Adversarial testing Evaluation pipelines AI/ML evaluation LLMs

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Own quality assurance for AI-native applications by designing functional and AI-specific test plans, evaluation harnesses, regression suites, and automated CI/CD pipelines. You will assess model and agent outputs, test adversarial edge cases, define acceptance thresholds, build evaluation frameworks, and communicate quality risks to delivery leaders.

Highlights

Own quality assurance for AI-native applications, combining functional testing with evaluation of accuracy, consistency, drift, and hallucinations. The role offers ownership of evaluation frameworks, quality gates, automated pipelines, and AI-specific testing practices.

Description

ABOUT: Owns quality for AI-native applications β€” functional testing plus the AI-specific evaluation (accuracy, drift, hallucination). KEY RESPONSIBILITIES Build test plans and automation for AI-native application features (functional + AI-specific)Design evaluation harnesses for model/agent outputs β€” accuracy, consistency, hallucination rateRun regression testing across model/prompt/config changes to catch silent quality driftRed-team AI features for edge cases and adversarial inputs where relevantBuild automated eval pipelines integrated into CI/CDPartner with AI Architects to define testability requirements before build startsOwn the quality gate before any AI feature ships to productionCommunicate quality risk to delivery leadership in terms they can act onTrain delivery teams on AI-specific testing practicesOwn the evals framework for the practice β€” golden datasets, scoring rubrics, LLM-as-judge calibration, and versioned benchmarks per use caseDefine eval acceptance thresholds per engagement and gate releases on themBuild eval engineering tooling β€” dataset curation, trace capture, offline/online eval runs, and dashboards delivery teams can readInstrument production evals and drift monitoring, feeding failures back into the golden datasets REQUIREMENTS & SKILLS 5–9 yrs QA/test engineering, with 2+ yrs testing AI/ML-powered features specificallyStrong test automation skills (Python-based frameworks, CI/CD integration)Understands AI-specific failure modes β€” hallucination, bias, drift, non-determinism β€” and designs tests for themStatistically literate enough to interpret model evaluation metrics, not just pass/fail resultsFamiliarity with red-teaming methodologies for AI systemsClear, assertive communicator β€” willing to block a release over a quality concernDetail-oriented and methodical under delivery-timeline pressureCollaborative but independent β€” doesn't rubber-stamp under delivery pressureExplains quality risk in business-impact terms, not just technical jargonHands-on evals engineering β€” builds and maintains eval suites with frameworks such as OpenAI Evals, Ragas, DeepEval, LangSmith, Azure AI Foundry evaluationsDesigns golden datasets and rubrics, and calibrates LLM-as-judge scoring against human reviewUnderstands RAG and agent eval metrics β€” groundedness, retrieval precision/recall, task completion, tool-call correctness, cost/latencyExperience wiring evals and drift monitoring into CI/CD and production observability