Senior Software Engineer - LLM Operations & Evaluation

Datasnipper — Netherlands · Posted ~1 hour ago

Senior Full-time Visa History ✓

Skills

Software engineering LLM operations AI evaluation Production systems Inference routing Failover Observability Experiment tracking Dataset versioning LLMs AI inference Evaluation systems Versioned datasets

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a small, high-impact engineering team building infrastructure for reliable LLM operations and evaluation. You will own production systems responsible for inference routing, failover, cost control, and quality measurement. You will also build evaluation infrastructure around versioned datasets, experiment tracking, and actionable evaluation runs, with your work influencing AI systems used at significant scale.

Highlights

Own production systems at the intersection of software engineering and AI infrastructure. The role offers substantial influence over inference routing, reliability, cost management, and AI quality evaluation, including versioned datasets, experiment tracking, and evaluation workflows with broad user impact.

Description

Every AI call in DataSnipper goes through us. Product teams do not talk to model providers directly, they talk to our gateway. We are responsible for how inference is routed, how it fails over, what it costs, and how anyone can tell whether the output is any good. The second half of the job is evaluation. We are building the platform teams use to measure AI quality: versioned datasets, experiment tracking, and evaluation runs they can act on. It is a hard problem and largely an open one, so you will have real influence over how we solve it. This is a small team with a large blast radius. You will own real production systems, set the standards other teams build against, and see your work in front of hundreds of thousands of users in audit and finance. Why DataSnipper Audit and finance are still massively manual and we are changing that. DataSnipper is a $1B, bootstrapped unicorn with 600,000+ users across 180+ countries, already embedded in the daily workflows of top audit and accounting firms. Now, we are taking things further with our Excel Agent, bringing AI directly into where the work actually happens. Unlike generic AI tools, we do not sit on the sidelines. Our AI operates inside Excel, with access to real documents and audit evidence, meaning it does not just generate answers, it does the work, with full traceability. We are not just applying AI, we are redefining how audit gets done. If you want to build something category-defining at scale, this is the place. What you will do Technical Delivery Own the LLM gateway: routing, provider failover, rate limits, retries, and cost attribution across multiple model providersDeploy, version and deprecate models across clouds, regions and environments, including quota and capacity planning, managed as infrastructure as codeBuild the shared evaluation platform: versioned datasets, experiment tracking, run and result schemas, reporting, and trace linkage back to the runOwn the infrastructure for async and long-running AI workloads Reliability, Security & On-Call Own observability for AI traffic: latency, retries and fallbacks, token usage, cost and errors, per team and per use caseTake part in the on-call rotation, run incidents, and close the follow-upsImplement the security and compliance controls the platform is held to: retention, access control, RBAC and SSO Collaboration & Impact Define and maintain clean integration contracts between the platform and the teams that consume itPartner with product and ML engineers to turn their requirements into platform capabilities that are self-service rather than a request queue What you will bring Must-Have 5+ years in backend or platform engineering, with strong production PythonExperience building or running LLM inference infrastructure: a gateway or routing layer with multiple providers, failover, rate limiting and cost attributionExperience with cloud at the infrastructure level and infrastructure as code (we run across Azure and GCP with Terraform)Experience running a shared service in production: on-call, incidents, postmortems, SLOsHands-on experience with observability tools (OpenTelemetry, Grafana), including instrumenting services and designing dashboards and alertsComfort with privacy and compliance work: PII handling, anonymisation, retention, access controlExperience building platform or shared-service capabilities consumed by multiple internal teams Nice-to-Have Experience with LLM or agent evaluationTemporal or another durable workflow engineSelf-hosted inference, capacity planning, load testingSynthetic data or document anonymisation pipelinesDocument AI: VLMs, OCR, structured extraction and the metrics that go with itDomain experience in audit, accounting or fintech What We Expect Ownership: You own work end-to-end, anticipate issues, and ensure high-quality delivery without close supervisionGrowth Mindset: You encourage open feedback exchange and provide clear, balanced feedback that helps others growCollaboration: You build strong cross-functional relationships and influence peers through expertise, data, and empathyAdaptability: You navigate ambiguity calmly, model positive behavior, and help peers adjust through clear communicationJudgment: You exercise sound judgment in ambiguous situations, balance speed and accuracy, and adjust priorities proactively Recruitment steps Recruiter screenHiring Manager interviewPeer programming sessionSystem design interviewFinal interviews with Engineering leadership