Summary
✨ AI‑Generated
A specialized AI engineering team is hiring a software engineer to build data systems, evaluation environments, and tools that improve the performance of advanced AI models.
Highlights
Work on challenging AI engineering problems, build evaluation systems, and turn research ideas into practical software solutions.
Description
Location: Yerevan, Armenia
Employment type: Full-time
Experience level: Mid-level
About the roleWe’re a boutique engineering consultancy and data lab helping leading AI model providers solve complex post-training challenges.
We build evaluation environments, synthetic data pipelines, and tools that test and improve how AI models perform on real-world tasks.
We’re looking for a resourceful Software Engineer to join our core delivery team.
You’ll turn open-ended research problems into working systems, explore unfamiliar codebases, and take ownership of projects from initial investigation through delivery.
One project might involve testing whether an AI agent can resolve complex Git merge conflicts in a Linux environment.
Another might involve generating datasets that evaluate multi-step mathematical reasoning.
You don’t need experience training large models from scratch.
You do need strong engineering fundamentals, curiosity, and the ability to learn quickly.
What you’ll doBuild AI evaluation environments: Develop sandboxed operating system environments, Docker-based test setups, and mock APIs to assess how AI agents handle complex tasks.Engineer synthetic data pipelines: Build workflows to generate, filter, transform, and reliably validate high-quality datasets at scale.Develop human-in-the-loop tools: Create internal workflows and annotation interfaces that help domain experts review and label complex data efficiently.Turn research into working software: Translate ambiguous requests into concrete technical approaches, experiments, and reliable pipelines.Investigate and solve unfamiliar problems: Read research papers, explore open-source systems, debug execution logs, and test practical solutions.What you’ll bringExperience building and shipping complex production software.Strong software engineering fundamentals and an understanding of how systems work beyond any single language or framework.The ability to navigate unfamiliar or partially documented codebases and make useful changes quickly.Ownership: you can break down a goal, identify the next steps, communicate challenges, and drive work to completion.Familiarity with asynchronous processing, concurrency, Linux environments, and Docker.An understanding of how to integrate LLM APIs while managing rate limits, reliability, and costs.A practical approach to engineering, with a strong focus on correctness and data quality.Comfort with evolving requirements and projects that require investigation and experimentation.Nice to haveThese are advantages, not prerequisites:
Experience with LLM pipelines using LangChain, LlamaIndex, DSPy, or direct OpenAI/Anthropic API integrations.Familiarity with post-training concepts such as supervised fine-tuning (SFT), RLHF, DPO, and the challenges of synthetic data quality.Experience with AI agent benchmarks such as SWE-bench, OSWorld, or GAIA.A background in competitive programming, complex systems debugging, or technically ambitious side projects.Why join us?You’ll help shape a core engineering team and work on challenging problems in AI evaluation and data infrastructure.
The role offers broad technical ownership, exposure to emerging research, and opportunities to turn new ideas into working systems.
If you enjoy understanding how things work, experimenting with solutions, and seeing difficult problems through to completion, we’d like to hear from you.