Research Engineer - AI Benchmarks

Meeboss — United States · Posted ~5 hours ago

Full-time Onsite Visa Sponsored $100K-$200K

Skills

Python Benchmark development Software engineering Data analysis Technical writing AI agents Benchmark infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A research-focused engineering team is hiring an engineer to design high-quality benchmarks for evaluating advanced AI agents on domain-specific tasks. You will build benchmark infrastructure, define realistic evaluations with subject-matter experts, develop metrics and analyses, validate real-world relevance, and produce clear technical documentation. The role is onsite in either San Francisco or Singapore, with relocation and visa support for strong candidates.

Highlights

Work on rigorous evaluation systems for advanced AI agents with strong compensation and relocation and visa support. The role combines research, engineering, benchmark design, infrastructure, metrics, and collaboration with technical experts.

Description

Experience: Technical aptitude and learning potential matter more than years Pay: $100k-$200k, full time Location: San Francisco, CA, US or Singapore (on-site) Visa: Relocation and visa support available for strong candidates The role Build high-quality benchmarks for evaluating frontier agents on domain-specific tasks. You will create benchmarks that are technically rigorous, practically useful, and credible to frontier labs. What you'll do Own the design, implementation, and quality of HUD's internal agent benchmarksWork with subject-matter experts to define realistic domain-specific tasksBuild infrastructure to run models and agents reliably against benchmark tasksDevelop metrics and analyses for benchmark difficulty, reliability, and failure modesValidate whether benchmark performance correlates with real-world evaluations, customer needs, and lab expectationsWrite clear documentation and benchmark reports for technical audiences What we're looking for Proficiency in Python, Docker, and Linux environmentsPublished papers or technical blogs about benchmarks, model failure modes, or related topicsA strong understanding of what makes a benchmark realistic, reliable, and usefulExperience with environments and evaluationsCuriosity and the ability to understand how workflows in unfamiliar domains actually workFirst-principles reasoning about task design, scoring, and failure modes Strong signals Detail orientation and an eye for subtle inconsistencies and edge casesIndependent work in unstructured problem spacesEarly-stage startup experienceStrong communication across time zones