Senior Applied AI Engineer, Agent Runtime

Datasnipper — Netherlands · Posted ~3 hours ago

Senior Full-time Visa History ✓

Skills

applied AI agent systems LLM orchestration prompt engineering context management tool integration AI evaluation observability LLMs MCP retrieval document extraction AI agents orchestration evaluation

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a small team as a Senior Applied AI Engineer building the runtime layer for long-running, multi-step AI agents. You will determine how agents reason, manage context, use tools, coordinate sub-agents, and are evaluated, while optimizing accuracy, latency, and cost in a hands-on engineering role.

Highlights

Own the applied AI layer of an agent runtime in a small team. The role offers substantial technical ownership over agent reasoning, tools, context management, orchestration, evaluation, accuracy, speed, and cost, with direct impact on AI automation for professional workflows.

Description

We are looking for a Senior Applied AI Engineer to join the Agent Runtime team behind Alwin, our new Agentic Automation Platform for audit and finance. Alwin agents run for long times, work through multi-step audit procedures over client evidence, and return finished work papers with every number traced back to source. Humans review and sign off. The runtime is the layer every agent depends on: the model access, the base system prompt and context, the tools (document extraction, retrieval, sandboxes, MCP integrations), the orchestration of agents and sub-agents, and the observability and evals that tell us whether an agent is performing to objective standards. You will own the applied AI half of that layer. Product teams build audit-specific agents on top of it; you decide how the base agent reasons, what tools it gets and at what abstraction, how it manages context over long horizons, and how we measure and hill-climb accuracy, speed and cost. This is a hands-on role in a small team with a large blast radius. Your work goes in front of hundreds of thousands of audit and finance professionals, and the problems are largely open: there is no playbook for production agents in a regulated domain, so you will help write it. About DataSnipper DataSnipper is the Agentic Automation Platform for audit and finance. Known worldwide for our Excel add-in, we are now building Alwin by DataSnipper: purpose-built agents that execute audit and finance workflows end-to-end, with every output traceable back to source evidence and a human signing off. Headquartered in Amsterdam with offices in New York, Tokyo, Kuala Lumpur and Sydney, we are used by hundreds of thousands of professionals at the world's largest firms. Our mission: automate the mundane, unlock the meaningful. What You Will Do Agent Engineering Own the base agents: system prompt, context engineering, memory and state management for tasks that span many turns and hours of executionDesign how agents use the tool layer, including document extraction, retrieval, code and file sandboxes, and MCP integrations, choosing the right level of abstraction so agents handle edge cases without wasting effort on mechanical stepsImplement agentic patterns such as sub-agent composition, planning modes, human-in-the-loop gates, and model routing across multiple providersShip end-to-end: from prototype with our audit domain experts, through evaluation, to production on AlwinDefine and build automatic processes that improve the agent continuously over timeKeep costs under control and implement cost-efficient approaches to AI-centric workflows Evaluation & Quality Extract signal from long agent trajectories: attribute outcomes to specific reasoning steps and tool calls, classify failure modes, and turn them into fixesHill-climb accuracy, latency and token cost, and make the trade-offs explicit for the teams building on the runtime Reliability & Collaboration Instrument agent behaviour with our observability stack and use production traces to drive improvementsPartner with Agent Experience teams, Document Intelligence and the AI Platform team to turn their needs into runtime capabilities that are self-service rather than a request queueStay on the frontier: evaluate new models, techniques and agent patterns, and bring the ones that hold up into production What You Will Bring Must-Have 5+ years of software engineering experience with strong, production-grade PythonExperience shipping and operating an LLM-powered product in production: you have dealt with hallucinations, latency spikes, tool failures and cost explosions at scale, and can explain what broke and how you fixed itHands-on experience building agentic systems: control loops, tool selection, planning versus execution, retries and fallbacks, not only prompt-and-parse pipelinesExperience with evaluation: you have built datasets, run offline and online experiments, and used the results to make an AI system measurably betterFluency with LLM APIs and agent frameworks across more than one model providerProficient with AI-assisted engineering and excited about working with coding agents dailyStrong fundamentals in software architecture and system design, and a track record of reliable, well-tested deliveryExcellent communication; you can work directly with auditors and product partners to define what "correct" means Nice-to-Have Experience with Durable workflow engines or long-running background task systemsExperience with RAG and retrieval pipelines over large, messy document corporaDocument AI: VLMs, OCR, structured extraction and their metricsSandboxed code execution, MCP, or multi-agent orchestration in productionDomain experience in audit, accounting or fintechFamiliarity with OWASP GenAI security practices and working in a regulated, privacy-sensitive environment What We Expect Ownership: You own work end-to-end, anticipate issues, and ensure high-quality delivery with minimal support. We expect you to influence the technical direction of the team.Growth Mindset: You encourage open feedback exchange and provide clear, balanced feedback that helps others growCollaboration: You build strong cross-functional relationships and influence peers through expertise, data, and empathyAdaptability: You navigate ambiguity calmly, model positive behavior, and help peers adjust through clear communicationJudgment: You exercise sound judgment in ambiguous situations, balance speed and accuracy, and adjust priorities proactively What we offer Being part of one of the fastest-growing scale-ups in the NetherlandsMake an impact by disrupting the audit industry with us28 vacation daysExcellent salaryPension planStock participation planHybrid work (Amsterdam-based)International team and environmentDaily lunch 🍽️Mental health support (OpenUp)Social events and team activities 🤩 Recruitment steps Recruiter screenHiring Manager interviewPeer programming sessionSystem design interviewFinal interviews with Engineering leadership