Software Engineering Expert for AI Model Evaluation

Huzzle Com — Netherlands · Posted ~3 hours ago

Mid Contract Remote €65-90/hour

Skills

software engineering code review Python technical evaluation

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A remote software engineering contract role focused on testing and evaluating advanced AI systems. You will create technical challenges, review generated solutions, and develop objective evaluation criteria based on real-world engineering practices.

Highlights

Fully remote contract opportunity with competitive hourly compensation and training provided for evaluating advanced AI systems.

Description

Software Engineering Expert - AI Model Evaluation · Contract · Remote · €65–90/h Role Overview Huzzle is seeking software engineers to test how well frontier AI systems handle real engineering work, and to document where they fall short.You will not be building products. You will be finding the point at which an AI system stops being reliable at engineering, and proving that point objectively.Each submission has three parts: a realistic technical request, documented evidence of where the model failed it, and a scoring rubric that lets any reviewer grade a future attempt consistently.Because this is code, most rubric criteria are verified by running it. The grading is unusually objective, which is what makes the work satisfying.Full training and worked examples are provided. Prior AI or data-annotation experience is not required. Key Responsibilities Test technical requests against a frontier AI model and review its output with the scrutiny you would apply to code you were merging.Document exactly where and how it failed, with the line, the failing condition, or the missing behaviour that demonstrates it. Precision matters more than volume.Author scoring rubrics with objective criteria another engineer could apply without repeating your work — ideally by executing the code against a named case.Write the underlying requests: realistic technical tasks requiring live research or real web interaction — actual API behaviour, current documentation, published licences, real datasets — plus implementation and a working file deliverable. Typically a Python script, but also CSVs, config, SQL or technical memos.Escalate task difficulty where the model succeeds, until a genuine gap is exposed. Note that models are strong at conventional coding problems; the gaps are in edge cases, real-world API behaviour and constraint discipline, not in algorithms.Log hours accurately and respond to review feedback within the working day. This is a two-week sprint. Ideal Qualifications 5+ years writing production software, or equivalent demonstrable depth.Strong Python. Comfortable consuming third-party APIs, handling errors and edge cases, and managing dependencies.The instinct to check what an API actually returns rather than trust what the documentation implies.Excellent written English and unusual precision with detail.3–6 hours per day available across the engagement.A public code portfolio (GitHub or equivalent) is helpful but not required. Contract and Payment Terms You will be engaged as an independent contractor.Fully remote, completed on your own schedule within the engagement window.The engagement may be extended, shortened or concluded early depending on programme needs and performance. About Huzzle Huzzle partners with leading AI organisations to evaluate and improve frontier models using deep human expertise. Contributors work directly on the assessment of advanced AI systems in their own field, are paid competitively, and help define the standards by which those systems are measured.