Software Engineer - AI Code Evaluation & Benchmarking

Hirefeedd — United States · Posted ~1 day ago

Mid Full-time Remote

Skills

Software engineering AI-generated code evaluation Code benchmarking Debugging Code validation Software quality assessment Programming environments

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking a Software Engineer to evaluate and benchmark the coding capabilities of advanced AI models. The role involves reviewing generated code for correctness, efficiency, maintainability, and requirements compliance; debugging and reproducing issues; validating fixes; and assessing generated explanations and implementations against real-world engineering tasks. This is a full-time remote role restricted to US candidates.

Highlights

Fully remote opportunity for US candidates to evaluate advanced AI coding capabilities, reproduce and debug software issues, benchmark solutions against real engineering tasks, and improve model reliability.

Description

Role: Software Engineer (Remote)Location: Remote (Work from Anywhere)Job Type: Full-TimePayout: Competitive, based on experience Role Overview: We are hiring for one of our clients, seeking a Software Engineer – AI Code Evaluation & Benchmarking (US candidates only) to work on a full-time basis. This role involves evaluating and benchmarking the coding capabilities of advanced AI models to ensure their accuracy, efficiency, and reliability. You will assess AI-generated code, validate solutions against real-world software engineering tasks, and identify correctness and quality issues. Key Responsibilities: • Review and evaluate AI-generated code for correctness, efficiency, maintainability, and adherence to requirements. • Analyze software engineering tasks and validate whether proposed solutions meet expected outcomes. • Debug code, reproduce issues, and verify fixes across different programming environments. • Assess model-generated explanations, reasoning, and implementation approaches for technical accuracy. • Create, refine, and maintain evaluation datasets, benchmarks, and grading rubrics for coding tasks. Required Skills & Qualifications: • Bachelor's or master's degree in computer science, software engineering, or a related technical field. • 3+ years of professional software engineering experience. • Strong proficiency in one or more programming languages, such as Python, Java, C/C++, Go, Swift, Objective-C, PHP, or SQL. • Strong understanding of data structures, algorithms, software design principles, and debugging methodologies. • Experience performing code review and identifying edge cases, failure modes, and areas where AI systems struggle with software engineering problems. More About the Opportunity: This role offers a unique opportunity to work with a global leader in the AI industry, contributing to the advancement of large language models through high-quality human feedback and evaluation. The position is ideal for engineers who enjoy code review, debugging, problem-solving, and applying strong software engineering judgment to complex technical scenarios. Equal Opportunity Employer: We hire based on skills and expertise. All qualified candidates are welcome regardless of background, experience, or prior employment history. Applications are reviewed solely on demonstrated technical ability and qualifications. Apply Now!