Freelance Machine Learning Engineer

Devfixr — Germany · Posted ~3 hours ago

Senior Contract Remote $30-100 per hour

Skills

machine learning Python ML engineering production machine learning systems algorithm development Machine Learning GPU

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A remote freelance role for machine learning engineers to design realistic engineering challenges, build evaluation systems, and solve advanced production ML problems. Ideal for engineers experienced with developing and testing robust machine learning solutions.

Highlights

Fully remote freelance opportunity with flexible hours, strong compensation range, and challenging machine learning engineering projects.

Description

Freelance | Fully remote, global | $30–100 per hour | Ongoing, flexible hours About the roleDevFixr is recruiting machine learning engineers for our client, a frontier AI company. You'll build the long-horizon ML engineering tasks used to train and evaluate frontier AI models. These are realistic problems a strong engineer would need hours to solve, each packaged with a grader that fairly scores whatever solution a model comes up with. Most tasks are CPU-based production ML problems, and some involve GPU work. The hard part is the grading. The AI is usually only told the symptom, and your grader has to give full marks to any correct fix while scoring shortcuts, hardcoded answers and tampered tests at close to zero. You prove it works by trying the cheats yourself. What you'll doDesign realistic, multi-step ML engineering tasks set in production systems. Examples include repairing a model that fails in live use, finding point-in-time feature leakage, diagnosing drift in a retrieval pipeline, or fixing a serving pipeline that returns wrong predictions. Some tasks involve GPU work, such as speeding up a slow training or inference job.Write a reference solution for each task, plus alternative correct solutions.Build graders that pass every correct solution and fail every shortcut, and prove it with negative controls.Package each task as a reproducible Docker environment with pinned dependencies, data and workspace.Test tasks against frontier models and calibrate difficulty, so they're hard for current models but clearly solvable and fairly graded.Write clear instructions that describe the problem without giving away the answer.Review other contributors' tasks and give feedback on quality and correctness.What we're looking forStrong Python and solid engineering habits, including testing, Git, Linux and Docker.Hands-on applied ML experience: data pipelines, feature engineering, model evaluation and debugging models in production, using classical methods as well as deep learning.An adversarial mindset. You naturally think about how a solution could game a test, and how to stop it.The ability to write precise, unambiguous technical instructions in English.Good judgement about what makes a problem genuinely hard, rather than just long or obscure.Strongly preferredPrevious work building RL environments, benchmarks, evals or AI training data, for example on Turing, Mercor or similar platforms, or on SWE-Bench or Terminal-Bench-style tasks.Nice to haveTraining and debugging deep learning models in PyTorch or JAX.GPU work such as CUDA or Triton kernels, mixed precision, memory and throughput profiling, or distributed training with FSDP or DeepSpeed.Research publications, open-source ML contributions or strong Kaggle results.TermsFreelance contract directly with our client.$30–100 per hour depending on experience, paid for hours worked.Fully remote and open worldwide, except in countries under sanctions.Ongoing work with flexible hours.How to applyApply with your CV or LinkedIn profile, answer a few short screening questions, and include links to relevant work such as GitHub, benchmarks or papers.Shortlisted candidates have a 30-minute technical interview with DevFixr.Successful candidates start a trial with our client on real tasks, with ongoing work for those who perform well.