Summary
A software engineering role focused on building reliable AI platform infrastructure, backend services, evaluation systems, and scalable data workflows for intelligent applications.
Highlights
Build cutting-edge AI infrastructure with end-to-end ownership, modern engineering practices, and direct impact on emerging intelligent systems.
Description
About the role
Conscium builds an AI agent evaluation platform - systems that test, grade, and stress AI agents at scale, and turn the results into actionable signal about how those agents behave.
We're looking for a strong Software Engineer to build the systems behind the platform.
This is a software engineering role first - designing services, data models, and pipelines that have to be correct, tested, and maintainable - applied to an AI problem domain.
You'll work with LLMs regularly, but the core of the job is engineering, not prompt-tuning.
You'll own features end-to-end: from data model and API, through the systems that run and evaluate AI agents, to the quality controls that make the results trustworthy.
What you'll work on
Backend services & APIs - Build and maintain the services, data models, and APIs that power the platform - designed for correctness, testability, and scale.Simulation & orchestration - Work on the systems that coordinate complex, multi-step interactions between AI agents and external systems, improving their reliability and throughput.Evaluation & scoring - Design systems that grade agent outputs, combining deterministic checks with model-assisted judgment - and make scoring reliable, explainable, and reproducible.Data pipelines - Build pipelines that generate, transform, and quality-check large volumes of structured data and benchmark content.Quality & reliability - Add the tests, instrumentation, and safeguards needed to trust outputs from systems that are inherently non-deterministic.
What we're looking for
Required
4+ years building and shipping production software, with strong proficiency in Python.Deep software engineering fundamentals: system and API design, data modeling, concurrency/async, testing strategy, debugging, and code review.
You can own a non-trivial service end-to-end.Experience designing and operating distributed or service-oriented systems (queues, workers, APIs) -not just calling them.Comfort designing schemas and working with relational databases, plus the migrations and performance concerns that come with them.Working knowledge of LLM APIs - orchestration, structured outputs, and handling non-determinism.
We expect you to *use* LLMs effectively, but this is not a prompt-engineering role.Ability to reason about correctness of *probabilistic* systems: how to test, measure, and trust outputs that aren't byte-for-byte deterministic.High quality bar: you write tests, types, and docs by default, and you keep changes small and reviewable.Must be based in Greece.
Nice to have
Experience building *agentic* or multi-agent systems, tool-use, or orchestration frameworks.Background in *evaluation / benchmarking* of ML or LLM systems (rubrics, golden datasets, model-as-judge, inter-rater reliability).Experience with distributed task queues and async workloads.Modern Python tooling and typed codebases (e.g.
type checkers, linters, Pydantic, FastAPI).Retrieval / search experience and working with data ingest pipelines.Some comfort with the infra side (Docker, CI/CD) so you can ship what you build.
What success looks like
30 days: You're productive in the codebase, have shipped a real feature or fix, and understand how work flows through the platform end-to-end.90 days: You own a meaningful slice of the system and have measurably improved quality, throughput, or reliability.6 months: You're shaping how we build and evolve the platform, and our results are more trustworthy because of your work.
What you'll get
Competitive salary, plus remote working and potential for stock options.A senior team that has done this before.
The people behind Conscium include founders who have built and sold AI companies, and operators who have scaled SaaS businesses with a proven track record.
You work with them directly and draw on that experience as you build.Real ownership.
You own systems from the data model through to production.
The architecture and quality decisions are yours to make, and you see them running in front of customers within weeks, not quarters.A problem with no playbook.
Agent verification is a new field.
There's no established way to build most of what we're building, so you help define how it gets done.Engineering experience that's hard to get elsewhere.
You'll ship production systems that run and grade AI agents at scale.
Engineers who have actually built and tested agent systems are hard to find right now, so the work sets you up for whatever comes next, and you get to be part of a fast growing story from early on.