AI Engineer

Muffintechdotai — Germany · Posted ~1 day ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

AI Engineer (m/f/d) Mid to Senior · Full-time · Hybrid/Remote About the role muffintech is at the forefront of AI-driven solutions for the insurance industry, empowering companies to seamlessly integrate cutting-edge technologies into their operations. We focus on vertical scalability – automating a broad range of insurance processes through a unified, AI-first platform rather than just tackling a single task. We are hiring an AI Engineer to build the AI capabilities at the heart of our product. This is a high-ownership, product-minded role: you will join the engineers already working on our AI layer and take your own features end to end — from retrieval and agent logic through to the backend services and pipelines behind them. Product brings the problem and the domain context; you decide how it gets solved, with tech leadership available on major architectural calls. AI features differ from the rest of the product in one important way: a passing test suite does not tell you whether they work. Much of this role is deciding how a capability should be measured, building the evaluation that answers it, and telling a real improvement from a plausible-looking one. Fluent use of modern AI-assisted tooling in your own work is expected. What you’ll do Own AI features end to end — from a product requirement and its domain context through to a measured capability running in production that you keep healthy over time.Build and evolve our retrieval and agent systems — ingestion and chunking, retrieval quality, tool design, and the orchestration that holds them together.Own the backend your features depend on, including the APIs, data ingestion, and pipelines that feed and serve them.Define how quality is measured: build the eval sets and production measurements that show whether a change improved the product, and use them before you ship.Design for the failure modes language models bring — unsupported answers, prompt injection through retrieved content, partial tool failures, and destructive actions against live systems.Work closely with product — establishing what “correct” means for a feature, surfacing edge cases early, and being straight about what the system can be trusted to do.Treat latency and cost as product constraints alongside quality. Our tech stack Python across the AI and backend layer.Retrieval-augmented generation and agentic workflows over insurance domain documents.Language models from several providers — called directly, with orchestration frameworks where they earn their place.REST and WebSocket for network communication, including streaming responses.Evaluation and observability tooling for model behaviour in production. What we’re looking for Essential Demonstrable experience taking LLM-based features into production — not prototypes or notebooks.Solid Python and backend engineering: you can design, build, and run the services, APIs, and data pipelines your features depend on.Real depth in retrieval: you understand how retrieval systems fail and can diagnose one from evidence rather than by trial and error.Experience designing and running evaluations for non-deterministic systems, and the judgement to recognise when a measured improvement is real.A clear-eyed view of language model failure modes, and the instinct to design around them at the system boundary rather than prompt around them.Sound technical judgment: you can weigh quality against latency and cost and set conventions.Proactive, clear communication with product about what the system can reliably deliver.Fluency with modern AI-assisted development tooling as part of your everyday workflow. Nice to have Experience in insurance, or another regulated domain where being wrong carries real cost.Experience with tool-calling agents against production systems, and the guardrails that makes necessary. Who this role suits This role suits someone comfortable owning a problem whose right answer is not knowable up front and has to be established by measurement. You will join engineers already working on the AI layer, but will be trusted with your own area and expected to set its direction. How we work Location requirement This position must be worked from within Germany. Because we handle sensitive data, the data-protection and regulatory commitments we make to our clients place requirements on where that data is accessed and processed. We are therefore unable to accept applicants who would be working from outside Germany. Remote work is fully supported — your place of work simply needs to be in Germany. Arrangement: Hybrid, or fully remote both possibleLocation: Berlin