Principal ML Researcher
Toloka — Netherlands · Posted ~4 hours ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
About Toloka
At Toloka AI we create data that powers leading GenAI models and innovations.
We work with frontier labs, big tech, renowned AI startups, enterprises and non-profit research organizations worldwide.
We use a combination of Experts + Crowd + Tech Platform to teach AI models to reason and evaluate their efficacy and safety.
We have experts in more than 50 different domains—from doctors and lawyers to physicists and engineers—and boast one of the most diverse global crowds, representing over 100 countries and speaking 40+ languages.
We are a well-funded startup with an enviable portfolio of clients including Anthropic, Amazon, Microsoft, Poolside, Recraft, and Shopify.
Recently, we secured strategic investment led by Bezos Expeditions and Nebius Group with participation from Mikhail Parakhin, CTO of Shopify and board advisor to leading GenAI companies, who now serves as our Chairman of the Board.
Our remote-first team is globally distributed around the world: USA, UK, the Netherlands, Serbia, and more.
About the Team
We are the ML team inside Toloka — we build the machine-learning products that power the platform itself, so every project running on Toloka is faster, cheaper, and more reliable.
A few examples of what we own:
Enterprise post-training- we are helping real companies by providing them small LLMs that beat frontier models in quality at a fraction of the cost.
Off-the-shelf post-training - we are conducting research on datasets we deliver to clients, showing that training on these datasets will improve their performance.
Fine-tuning and RL — adapting frontier and open-source models to Toloka's tasks to hit the right quality at the right cost.
Evaluation, benchmarking, cost modeling, and model selection across providers.
LLM QA — the core technology behind Toloka's automated quality-check mechanism.
Every annotation flowing through Self-Service is reviewed by an LLM agent we design, train, and operate.
We own the full chain.
The same team designs the ML solution, ships it to production, keeps it running 24/7, analyzes the results coming back from real projects, and feeds that signal into the next iteration.
No hand-off between research, engineering, and operations — it's all us.
About the Position
As a Principal ML Researcher, you will define the overarching technical vision, research strategy, and architecture for Toloka’s core ML and post-training stack.
In this high-impact role, you will bridge frontier AI research and large-scale platform engineering.
You will lead technical strategy across greenfield post-training paradigms (such as GRPO, process/outcome reward modeling, and RLAIF), architect resilient automated evaluation ecosystems, and set the standard for how foundation models are adapted and served on our platform.
As a technical authority, you will mentor senior ML engineers, collaborate directly with executive leadership, and represent Toloka in the broader AI research community.
What you’ll do
Drive Technical Strategy: Define the multi-quarter strategy and architecture for Toloka’s post-training stack, ensuring our fine-tuning, RL, and evaluation capabilities stay ahead of industry trends.
Architect Greenfield Post-Training & RL Pipelines: Spearhead the transition beyond basic SFT into advanced RL frameworks (GRPO, PPO/DPO, reward modeling, Process Reward Models, and RLAIF) to power platform-wide LLM alignment and agent behavior.
Pioneer Next-Gen Evaluation Harnesses: Design and calibrate bulletproof automated evaluation ecosystems, including human-aligned LLM-as-a-judge frameworks, dynamic benchmarking suites, and regression control setups.
Advance Distillation & Inference Optimization: Lead research initiatives on model distillation, quantization, context compression (e.g., gisting), and speculative decoding to optimize latency-cost trade-offs at platform scale.
Shape Agentic Platform Capabilities: Architect autonomous guiding agents and tool-use workflows, establishing foundational frameworks for multi-step reasoning, self-correction, and evaluation-driven feedback loops.
Technical Leadership & Mentorship: Elevate the engineering and research bar across the team through architectural design reviews, hands-on mentoring of Senior ML Researchers, and establishing best practices for reproducible ML R&D.
Cross-Functional & Community Influence: Partner with Product and Engineering directors to translate complex client challenges into scalable platform architecture; author high-impact research write-ups, blog posts, and external tech talks.
What we're looking for
6+ years in ML engineering or applied research, with 3+ years of proven track record leading LLM post-training, fine-tuning, or alignment initiatives at scale.
Deep theoretical and practical mastery of LLM alignment: SFT, LoRA/PEFT, RLHF/RLAIF (GRPO, DPO, PPO), reward modeling, and reasoning-oriented post-training.
Mastery of LLM Evaluation & Calibration: Proven history of building evaluation pipelines from scratch, addressing LLM-as-a-judge biases, and tightly calibrating automated evals against human gold standards.
Expert Distributed Training & Systems Engineering: Strong proficiency in Python, PyTorch, distributed training frameworks (Deepspeed, Megatron, FSDP), and high-performance inference engines (vLLM, TensorRT-LLM).
Architectural Vision & Product Leadership: Demonstrated ability to map ambiguous technical challenges into robust platform features, balancing state-of-the-art research with production reliability and cost constraints.
Technical Authority: Proven experience acting as a technical leader, mentoring senior researchers/engineers, and influencing cross-functional roadmaps.
Language: Fluent spoken and written English (C1), with clear communication skills to articulate complex technical concepts to both internal teams and external clients/community.
What we can offer
You will be part of an international, dynamic environment that drives innovation and sets new standards in the AI and technology sector.
Competitive compensation package including base salary, bonus, and ESOP.
Paid PTO and benefits will vary depending on location.
We offer a full remote or hybrid model (if you are based in NL or Serbia).
IT setup and home office allowances.
Equal Opportunity Employer:
Toloka is committed to providing equal opportunity and fostering an inclusive environment.
We welcome applications from all qualified individuals and do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, gender identity, age, marital status, veteran status, disability, or any other characteristic protected by applicable law.
Selection decisions are made based on qualifications, merit, and business need.
[Important Notice] Scam Alert Regarding Fake Job Postings
It has come to our attention that an individual or group is fraudulently impersonating Toloka to post fake jobs and solicit personal information from applicants.
Please be aware:
Official Communication: Our recruiting team will only contact you from an official "toloka.ai" email address.
We will NEVER use Gmail, Yahoo, Tolokainc, toloka.inc, or other personal or seemingly business email accounts.
Our Process: We will never ask for your bank account details, credit card number, or any fees as part of the application or interview process.
Official Listings: All legitimate job openings are posted on our official careers page: https://toloka.ai/careers#job-list
What to do: If you see a suspicious job posting or have been contacted by someone you suspect is a scammer, please do not provide any personal information.
Instead, report the incident to us directly at security@toloka.ai and report the profile/post to LinkedIn.We are taking this matter very seriously and are working with the appropriate parties to resolve it.
Thank you for your vigilance!
To learn how we collect, use, disclose, and store personal data, check out our Privacy Notice.
We have 152,764 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 152,764 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume