Description
Armeta is an applied-AI company building engineering-intelligence products for construction and oil & gas.
This is a forward-deployed role.
You won't be shipping against a backlog written by someone else.
You'll sit next to the engineers who use the product, watch how they actually work, and close the gap yourself - frequently in the same session where the problem surfaced.
What you'll buildOwn the product's AI/ML layer end to end — from prototype to production MVP.Build agentic LLM pipelines: orchestration, tool/function calling, structured output, long-context handling, model routing, and cost/latency control.Turn free-form engineer input into structured, validated machine representations — and back again.Design retrieval over a large domain knowledge base.Ship AI features that assist the engineer through the workflow: setup, interpretation of results, and diagnosis and explanation of errors in the user's work.Integrate the AI layer with the computation engine and the product frontend via REST/WebSocket APIs.Build eval pipelines and quality metrics: golden datasets, regression runs, offline/online evaluation — so keep/kill decisions on each hypothesis are objective.
What you'll do in the fieldRun pilots hands-on: deploy into the customer's environment (including on-prem and restricted-network setups), instrument it, and drive it to first real value.Sit in on live sessions with domain engineers — capture how they phrase a task, where the system misreads them, which failure modes actually block adoption.Debug in front of the customer.
You diagnose live, patch, and redeploy the same day where possible.Convert field observations into product: prompt and pipeline changes, new evals, new golden cases, and a clear signal to the squad about what's worth building next.Own the customer's technical relationship alongside the business analyst - demos, integration questions, scoping, and honest answers about what the system can't do yet.
RequirementsPython — 2+ years of commercial experience, with deep command of async (asyncio), typing, Pydantic, and clean architecture.FastAPI and production-API design: REST, WebSocket, background processing, task queues (Celery / RQ / arq).Strong proficiency with Claude Code and its full toolset: subagents, MCP servers, hooks, custom slash commands, CLAUDE.md and context/memory management, plan mode, skills, headless mode.Hands-on LLM experience in production: agent frameworks (LangGraph and similar), function/tool calling, structured outputs, RAG, embeddings and vector DBs (pgvector / Qdrant / Weaviate), fine-tuning, and building and analyzing metrics and evals.PostgreSQL and data handling.Experience taking a product to MVP in a small team: backend, APIs, integrations.Customer-facing ability: you can hold a technical conversation with a practicing engineer who is not an AI person, take criticism without getting defensive, and stay calm when something breaks during a demo.Ship-under-pressure temperament: comfortable debugging unfamiliar environments, working with incomplete information, and making a call without waiting for consensus.Working proficiency in Russian and English (Kazakh is a plus).Willingness to travel to customer sites for pilots and deployments.
Nice to haveTechnical background in engineering or applied math; comfort with quantitative domains.Numerical methods, computational geometry, scientific Python (numpy / scipy).Experience with complex domain-specific data formats and engineering software.Experience as a forward-deployed / solutions engineer, or launching products from scratch with direct customer contact.Experience deploying into on-prem or restricted-network enterprise environments.
How we'll measure youNot by tickets closed.
By whether real users adopt what you built, whether your evals catch regressions before customers do, and how fast the loop runs from a problem observed in the field to a fix in production.