Machine Learning Engineer

Armeta Ai β€” Kazakhstan Β· Posted ~2 hours ago

Mid Full-time

Skills

Machine learning Deep learning NLP Computer vision Python FastAPI Docker Computer Vision LLMs

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An AI engineering role focused on developing intelligent applications and pipelines using modern machine learning technologies. The position involves model development, production services, and full-cycle ownership of AI features.

Highlights

Build production AI applications end-to-end, combining machine learning expertise with software engineering ownership.

Description

TL;DRWe're looking for an ML Engineer who builds applications and pipelines using a modern AI/ML stack and owns features from start to finish. You need solid understanding of ML/DL fundamentals β€” from dataset creation to model training and evaluation. Both NLP and Computer Vision matter: document parsing, structured data extraction (PDFs, Excel, images), object detection, and LLM prompt engineering. You build production-ready FastAPI services, package them in Docker, and own the full stack: AI/ML + engineering. About ArmetaArmeta builds an engineering intelligence layer for the construction and industrial sectors, turning complex technical archives into accessible, actionable information. In Kazakhstan, Armeta analyzes initial permitting documents (Π˜Π Π”) and design-and-estimate documentation (ΠŸΠ‘Π”) to check them for compliance with building codes and regulations, automating review workflows that have historically been fully manual. The platform unifies, contextualizes, and makes queryable critical engineering and construction documents, and is deployed in customers' cloud environments or on-premise for flexibility and control. Armeta is a team of experts with deep expertise in software development and ML/DL, with offices in Astana, San Francisco, and Doha. The technology is already in production, supporting real-world facilities and enabling more efficient engineering and compliance workflows. Role DescriptionThe Machine Learning Engineer will design, develop, and deploy the backend services and ML systems that power Armeta's document understanding and compliance products. You own your solutions end-to-end, covering both the AI/ML and engineering sides. The role sits at the intersection of NLP and Computer Vision: the documents we process are not plain text β€” they are multi-hundred-page packages of drawings, schedules, stamps, tables and prose, where the meaning of a clause often lives in a drawing sheet rather than a paragraph. On the language side, you'll:Build backend ML services with FastAPIIntegrate LLMs and NLP models into applications and internal toolsWork extensively with retrieval-augmented generation (RAG) β€” loading, processing, indexing, and searching data across vector and full-text storesApply LLM prompt engineering for matching, search, reasoning, and classification tasks On the vision side, you'll:Build models and pipelines that read engineering drawings and return structured, checkable resultsDetect and classify graphical elements and symbolsRead title blocks and annotations, extract tables and schedules embedded in sheetsAssociate text with the geometry it refers toReconcile what a drawing shows against what the rest of the documentation statesThe output is not a description β€” it is machine-readable results a reviewer can act on Day-to-day you'll:Build complete end-to-end pipelines: ingestion β†’ inference β†’ production servingTrain and fine-tune models for NLP and CV tasksDevelop document parsing and structured data extraction from different formats (PDFs, Excel, images, drawings)Test, debug, and optimize AI/ML features to keep them accurate, robust, and scalableOwn the quality side: labeled datasets, metrics, regression runs, honest error analysisPackage solutions in Docker with appropriate dependencies and configurationThis is a full-time, on-site position located in Astana, Kazakhstan. QualificationsRequired: 1–3 years of experience in ML with hands-on work in computer vision β€” detection, segmentation, classification, or document/layout analysis β€” alongside NLP or classic deep learningPractical experience with document AI on visually complex inputs: OCR, layout analysis, table extraction, or object detection on scanned or vector documents (drawings, forms, schematics, plans)Familiarity with modern detection/segmentation architectures and the tooling (YOLO-family, DETR-family, Detectron2, MMDetection, or similar), including training on custom datasetsHands-on experience designing and developing microservices with FastAPIProduction experience with vector databases and/or full-text search engines (e.g., Qdrant, Milvus, Elasticsearch)Experience with one or more LLM frameworks (LangGraph, Haystack, LlamaIndex)Experience working with multimodal LLMs (Gemini, Claude, GPT) as well as open-source models (Qwen, Gemma, and similar) β€” including using VLMs on image inputs, not just textProficiency in Python and familiarity with common ML frameworks (PyTorch, TensorFlow, scikit-learn)Comfort with dataset construction, annotation guidelines, class imbalance, and evaluating models whose errors have real consequencesAbility to work on-site in Astana, collaborate with cross-functional teams, and communicate technical concepts clearly Nice to haveFamiliarity with construction, engineering, or industrial domains β€” reading drawings, understanding ΠŸΠ‘Π”/Π˜Π Π” structure, prior work with normative documentationExperience parsing engineering formats (PDF vector layers, DWG/DXF) or with CAD/BIM toolingExperience with GPU training and inference optimizationExperience deploying models in on-premise or restricted-network enterprise environmentsExperience building human-in-the-loop review interfaces or annotation workflows How we'll measure youNot by tickets closed. By whether your features work reliably in production, whether your models stay accurate as data shifts, and how quickly you can diagnose and fix issues when they surface in real usage.