Software Engineer - AI Evaluation Systems

Key Talent Solutions1 β€” United States Β· Posted ~2 hours ago

Mid Full-time

Skills

Python machine learning data engineering LLM workflows production software development Machine Learning LLM

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A software engineering role focused on creating datasets, evaluation frameworks, and machine learning tools that improve AI model performance.

Highlights

Opportunity to work on advanced AI systems, building evaluation frameworks and data solutions for machine learning applications.

Description

πŸš€ Software Engineer β€” RL Environments We're partnering with a fast-growing (Series B) AI company that builds the datasets and evaluation systems frontier AI labs use to train their models. You'd be working hand-in-hand with top research teams designing the tasks models practise on, and the scoring that decides whether they're really getting smarter. What you'd own: πŸ”Ή Design data that exposes where models fail across finance, code & enterprise workflows πŸ”Ή Build reward signals and evaluation rubrics for RLHF / RLVR training pipelines πŸ”Ή Develop frameworks to measure dataset quality and its real impact on model performance πŸ”Ή Turn ambiguous research goals into concrete, shippable systems You must have: πŸ”Ή Strong Python + production ML/data experience πŸ”Ή Hands-on with modern ML/LLM workflows training, fine-tuning or evaluation πŸ”Ή Sharp instincts for data quality, edge cases and precision/recall trade-offs πŸ”Ή 1–4 years shipping technical work (RLHF/RLVR or eval experience is a big plus) The details: πŸ“ San Francisco β€” on-site, 5 days πŸ’° starts from $250 base + significant profit share + equity πŸ›‚ Visa sponsorship available (H-1B, O-1, OPT) This is a rare chance to have direct, measurable impact on frontier AI on a small team where your work ships straight into the models defining the field.