Summary
✨ AI‑Generated
A part-time remote consulting role for an experienced software engineer focused on evaluating large language models. You will curate high-quality code, develop verification mechanisms, assess AI-generated software against engineering standards, and create rigorous benchmarks. The work spans architecture, prototyping, production, and maintenance while collaborating with technical and research professionals.
Highlights
Flexible part-time consulting opportunity for experienced software engineers working on rigorous evaluation of large language models. The role spans architecture through production, offers remote collaboration, and contributes directly to reproducible benchmarks and higher-quality AI systems.
Description
This position is listed on behalf of a partner company, who manages all applications and next steps.
Our partner is looking for a Senior Software Engineer – LLM Evaluation based in United States.
This part-time consulting opportunity is designed for experienced software engineers interested in advancing the evaluation of large language models.
You will curate high-quality code, develop technical solutions, and evaluate AI-generated software against real-world engineering standards.
The role spans multiple programming languages and covers the complete software-development lifecycle, from architecture and prototyping through production and maintenance.
You will design verification mechanisms and contribute to benchmarks that make AI evaluation more rigorous, consistent, and reproducible.
Your engineering expertise will help research teams identify model strengths, weaknesses, and recurring coding errors.
You will collaborate remotely with technical and research professionals on projects at the intersection of software engineering and AI.
The flexible contractor structure requires a minimum of 10 hours per week, with the possibility of working up to 40 hours depending on project needs.
Accountabilities
Curate high-quality code examples and technical datasets for model training, benchmarking, and evaluation.
Develop accurate solutions to software-engineering tasks and correct or improve implementations across multiple programming languages.
Work with technologies such as Python, JavaScript, ReactJS, C/C++, Java, Rust, and Go as relevant to project assignments.
Evaluate AI-generated code for technical correctness, maintainability, efficiency, scalability, reliability, and adherence to professional engineering standards.
Identify implementation weaknesses, recurring coding errors, and patterns that reveal limitations in AI-generated software.
Provide clear, structured rationales explaining technical evaluation decisions and assessment outcomes.
Build agents and automated mechanisms capable of assessing code quality and verifying software solutions.
Design reliable checks that support consistent and reproducible evaluation across repeated engineering tasks.
Evaluate AI capabilities across the full software-development lifecycle, including prototyping, architecture, API design, production implementation, experimentation, launch, monitoring, and maintenance.
Assess model-generated technical reasoning and decisions against practical software-engineering expectations.
Collaborate with research and cross-functional technical teams to define evaluation strategies and improve coding benchmarks.
Contribute to datasets used for training and benchmarking while maintaining rigorous standards for quality and technical accuracy.
Help improve coding-focused evaluation systems through iterative analysis of model performance.
Complete all work without using confidential, proprietary, unreleased, employer-restricted, client-restricted, or otherwise protected code, datasets, architecture materials, or technical information belonging to any third party.
Requirements
3+ years of professional software-engineering experience.
Strong full-stack development capabilities and experience building scalable, production-grade software.
Strong understanding of software architecture, system design, API design, and production implementation.
Deep knowledge of software development, debugging, code review, and code-quality assessment.
Demonstrated ability to review, troubleshoot, and improve complex software implementations.
Proficiency in one or more relevant programming languages, including Python, JavaScript, Java, C++, Rust, or related technologies.
Experience with ReactJS, C, Go, or additional programming languages is valuable depending on project requirements.
Familiarity with software monitoring, operational maintenance, and production reliability.
Ability to reason across the complete software-engineering lifecycle and evaluate technical decisions from development through ongoing operation.
Strong analytical and problem-solving skills, combined with a rigorous and detail-oriented approach to technical evaluation.
Excellent written and verbal communication skills, including the ability to produce concise and well-structured evaluation rationales.
Ability to distinguish between technically correct implementations and solutions that may introduce scalability, reliability, maintainability, or architectural concerns.
Comfortable working independently and collaborating remotely with research and technical teams.
Must be based in the United States, Canada, or an eligible Western European country.
Willingness to complete a required AI video interview as part of the application process.
Benefits
Fully remote, part-time independent contractor engagement.
Flexible workload starting at 10 hours per week, with the potential to work up to 40 hours per week.
Approximately one-month initial project duration, with potential extension based on performance and project fit.
Opportunity to contribute directly to advanced LLM evaluation, coding benchmarks, and AI-assisted software-engineering research.
Exposure to cutting-edge AI evaluation workflows involving realistic software-development scenarios.
Opportunity to apply professional software-engineering expertise to improve how AI systems are evaluated.
Flexible consulting structure suited to experienced engineers seeking project-based work.
International remote collaboration with research and technical professionals.
Application process expected to take approximately 15–30 minutes, followed by a required AI video interview.
Compensation is project-specific and was not specified in the available job materials.
Medical insurance and paid-leave benefits are not included under the independent contractor arrangement.
How Jobgether Works
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements.
Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company.
The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer.
This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR).
You may exercise your rights (access, rectification, erasure, objection) at any time.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information.
These tools assist our recruitment team but do not replace human judgment.
Final hiring decisions are ultimately made by humans.
If you would like more information about how your data is processed, please contact us.