Software Engineer, AI Training Data & Evaluation Platform

Atomic Remote Jobs — Canada · Posted ~5 hours ago

Mid Full-time Remote

Skills

software engineering platform development production systems

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A fast-growing AI research organization is seeking a software engineer to build platforms, pipelines, and evaluation infrastructure that improve advanced AI systems. The role combines engineering, experimentation, and production ownership.

Highlights

Fully remote role focused on building impactful AI infrastructure, evaluation systems, and production platforms for advanced research workflows.

Description

Company Overview Our client is a fast-growing applied research lab building the data layer for frontier AI. They partner with leading AI labs and enterprises to deliver two things: proprietary, expert-generated datasets, and rigorous evaluation and benchmarking. The goal is for AI systems to get better at real workflows, not just polished demos. Founded in 2025 and backed by a $30M Series A, the company is fully remote. It works on problems such as turning real-world work into clean training signals, building evaluations for software engineering agents and finance workflows, and proving data quality by measuring actual performance lift. Your Role This is not a "pick up tickets and wait for specs" role. This is a broad builder seat that combines platform engineering with evaluation and experimentation infrastructure. You'll design, build, and run systems in production that researchers and operators depend on every day.The company is scaling the pipelines, evaluation harnesses, and training environments that make its work repeatable, and it needs engineers who can own a system and deliver reliably.In your first 30 to 90 days, success means shipping at least one meaningful production improvement and becoming the go-to owner of a core system. You'll: Build and maintain evaluation harnesses that measure model and agent performance on real tasksImprove eval reliability, coverage, and signal quality through better rubrics, task design support, and scoringShip tools that let researchers and operators run experiments without reinventing the process each timeBuild APIs and backend services that power human-in-the-loop workflows, task routing, and quality checksImprove the pipelines that turn expert work into structured training and evaluation dataMake systems more observable, scalable, and easier to operate through logging, metrics, and debuggingWrite clear, maintainable code, take part in reviews and design discussions, and document decisions so others can build on them You Bring: Strong coding fundamentals in Node.js and TypeScriptStrong coding ability in Python and/or GoExperience building and owning production systems such as APIs, services, and pipelinesA solid understanding of distributed systems and engineering trade-offsComfort with AWS or GCP and modern infrastructure (containers, Kubernetes)A track record of shipping and maintaining systems other people rely on, not only prototypesStrong written communication and comfort working async in a distributed team Bonus Points: Experience with evaluation frameworks, experimentation platforms, or ML toolingExperience with data pipelines or workflow orchestrationExperience building internal platforms for operators or research teamsExperience in early-stage or high-ownership B2B SaaS or platform teams What's Offered: Full-time, fully remote role with a LATAM focus and meaningful overlap with U.S. time zones$7,000–10,000 USD/month, based on experienceReal ownership, with growth into bigger systems, deeper technical leadership, and projects core to how the company scalesA lean, async-first team that values clear writing, sound judgment, and follow-throughResearch-adjacent engineering at the frontier of AI, alongside practical platform workDirect impact on the data and evaluations used by leading AI labs Interview Process: 1️⃣ Take-home assignment covering practical engineering and how you communicate decisions 2️⃣ Application review by the team 3️⃣ Founding engineer screen, a deep dive on system design, trade-offs, and past ownership 4️⃣ Work trial on real-world work, focused on execution, quality, and collaboration 5️⃣ Offer