AI Evaluation and Automation Software Engineer

Nlb Services — United States · Posted ~3 hours ago

Mid Contract Remote $70-$78/hour

Skills

Python Software engineering Automation Testing Git Docker Java JavaScript TypeScript

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking an engineer to build tools that measure and improve AI-powered software performance. The role involves creating evaluation frameworks, automation workflows, and reproducible testing environments.

Highlights

Remote contract role focused on building evaluation systems, automation tools, and improving AI software quality.

Description

About The Role We’re looking for an AI Evaluation Engineer to build and scale the tooling used to measure how well AI-powered software development tools actually perform. You’ll build evaluation harnesses, automate benchmark runs, create reproducible environments, and ensure our measurements are reliable, meaningful, and validated against real-world engineering outcomes. Location: Remote Type: Contract Pay Range: $70–$78/hour on W2 Roles & Responsibility Build evaluation harnesses, benchmark automation, and reproducible environments using Git, Docker, and isolated worktrees. Run and analyze AI coding benchmarks across quality, productivity, cost, and latency; validate results against human judgment. Identify failure patterns and improve evaluation workflows, collaborating with engineering and data teams. What We’re Looking For Strong software engineering experience in automation, developer tooling, testing, or validation, with proficiency in Python, Java, JavaScript/TypeScript, or similar. Strong knowledge of Git, Docker, APIs, CI/CD, and software development workflows, plus an understanding of LLM/agent evaluation. Hands-on experience with AI coding tools or agentic systems such as Claude Code, Devin, Cursor, or similar. Nice to Have Experience building software benchmarks, execution-based evaluations, or automated grading systems. Familiarity with LLM/agent-as-a-judge approaches, human-rater validation, and evaluation reliability. Experience with Bazel/test selection, reproducible environments, versioned datasets, or technical documentation. ____________________________________ Equal Employment Opportunity NLB Services is an Equal Opportunity Employer. All qualified applicants will receive consideration without regard to race, color, religion, sex (including pregnancy, sexual orientation and gender identity), national origin, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable law. Reasonable Accommodation NLB Services provides reasonable accommodations during the recruitment process as required by law. Applicants requiring an accommodation due to a disability, pregnancy-related limitation, or sincerely held religious belief, practice or observance may contact their recruiter or email notifications@nlbservices.com. For Benefits: Eligible employees may receive medical, dental and vision insurance, and other benefits. Eligibility and benefit terms vary based on employment status and assignment. Other Important Terms Applicants must be legally authorized to work in the United States. Any employment offer may be contingent upon job-related screening, employment verification and other requirements permitted by applicable law.