Description
As a Mid-Level AI Software Quality Engineer, you will design, architect, and execute automated validation and testing strategies for enterprise AI-powered applications, autonomous agents, copilots, retrieval-augmented generation (RAG) platforms, and machine learning systems.
This position bridges core software engineering in test (SDET) principles with state-of-the-art AI evaluation methodologies to ensure generative models and agentic workflows are robust, accurate, scalable, and secure.
Operating in a high-governance enterprise environment, you will take ownership of building end-to-end testing frameworks from scratch.
You will validate tool-calling behaviors, evaluate Model Context Protocol (MCP) integrations, mitigate hallucinations, and construct automated pipelines to assess model reliability across all stages of production deployment.
Responsibilities:
AI Quality Architecture & Test Strategy: Architect comprehensive automated testing frameworks tailored for Large Language Models (LLMs), agentic decision loops, and RAG architectures; evaluate model outputs for precision, consistency, latency, and safety.
Test Automation Engineering: Develop and maintain test automation suites from scratch utilizing Python, Java, JavaScript/TypeScript, or C#; design mocks, synthetic datasets, and test harnesses for non-deterministic AI outputs.
Agent & Protocol Integration Testing: Validate Model Context Protocol (MCP) tool integrations, function calling, external API triggers, and state management within complex multi-agent workflows.
Adversarial & Safety Evaluation: Execute automated adversarial, prompt injection, and boundary testing to uncover edge-case failures, bias, data leakage, and hallucinations while enforcing enterprise guardrails.
CI/CD Pipeline Integration: Embed automated AI evaluation and regression benchmarking suites directly into CI/CD pipelines (e.g., GitHub Actions, Jenkins) to score quality metrics prior to production releases.
Performance & Scalability Validation: Conduct stress, load, concurrency, and multithreading tests on distributed AI orchestration services, vector databases, and real-time inference APIs.
Cross-Functional Collaboration: Partner with AI research engineers, application developers, enterprise architects, and governance teams to drive defect triage, establish quality gates, and maintain regulatory compliance.
Minimum Qualifications:
Bachelor’s degree in Computer Science, Software Engineering, Information Systems, or equivalent practical experience.3–6 years of professional experience in software quality engineering, automated test development, or backend SDET roles.Strong programming proficiency in at least one modern language: Java, Python, TypeScript/JavaScript, or C#, including solid functional programming and multithreading fundamentals.Demonstrated hands-on experience with API automation tools and test frameworks (e.g., PyTest, JUnit, Playwright, Selenium, Postman).Working knowledge of Linux environments, SQL databases, and relational data querying.Practical understanding of distributed web architectures, RESTful APIs, and CI/CD automation systems.Ability to work fully on-site in Ann Arbor, MI, 5 days per week during the initial 6-month ramp period and participate in an on-site interview process.
Preferred Qualifications:
Hands-on experience developing or testing Generative AI applications, autonomous agents, RAG architectures, or Model Context Protocol (MCP) implementations.Direct experience integrating with or evaluating enterprise LLM APIs (OpenAI, Anthropic Claude, Google Gemini, Azure OpenAI).Familiarity with containerization and orchestration tooling (Docker, Kubernetes) and cloud infrastructures (AWS, Azure, or GCP).Background in highly regulated enterprise domains (e.g., Financial Services, FinTech, Healthcare) adhering to strict data privacy, model governance, and compliance standards.Experience implementing automated scoring mechanisms for natural language generation (e.g., semantic similarity, grounding metrics, faithfulness evaluation).Maintenance and Support.