Senior AI Engineer

Reppls — Poland · Posted ~2 hours ago

Senior

Skills

machine learning NLP information extraction model training deployment

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a software development company focused on turning messy documents into clean data. You will build and ship ML/NLP models, collaborate with international teammates, and tackle production‑grade information extraction projects.

Highlights

Design, train, and deploy machine learning models that turn semi‑structured business data into reliable structured information, working on real‑world information‑extraction challenges with a global team.

Description

This role targets candidates searching for Senior AI Engineer, Machine Learning Engineer, NLP Engineer, or Applied Scientist roles with deep, hands-on experience shipping production information-extraction systems. About Us Our customer is a software development company delivering effective digital solutions to clients around the world, with a strong focus on the U.S. market. Headquartered in sunny San Diego, California, we've built a team of over 300 talented professionals across multiple countries. Fueled by courage and curiosity, we go beyond providing services; we dive deep into our clients' businesses to uncover real needs and tackle key challenges, combining human insight with technology. Our success comes from a deep commitment to results and long-term partnerships built on trust. About the Role As a Senior AI Engineer, you will design, train, and ship the models that turn messy, semi-structured business documents into reliable, structured data — for example extracting and decomposing the goods, parties, and terms from trade-finance documents and Letters of Credit. This is a hands-on, data-centric role: you own problems end to end, from pulling and cleaning production data, to training custom on-prem models and standing up LLM-based extractors, to proving they work with honest evaluation and deploying them into a live microservice platform. You will set the technical approach for extraction workstreams, be relentless about data quality and evaluation rigor, and act as a trusted technical partner to the product and domain teams who depend on the results. Key Responsibilities Build production extraction models for complex documents: token classification / sequence labeling, span segmentation and decomposition, and structured attribute extraction from semi-structured text (SWIFT/LC fields, invoices, packing lists, certificates).Fine-tune and train small, on-prem transformer models (encoder token-classifiers, small language models) on GPU, and stand up LLM-based extraction (Azure OpenAI / GPT) as baselines and PoCs — with anti-hallucination grounding and structured output.Own the data: pull from production stores (MongoDB), establish label provenance, deduplicate, and build leakage-free splits (e.g., by transaction/tenant) so metrics mean something.Build and maintain evaluation harnesses: report honest precision / recall / F1, make results reproducible, and actively hunt for contamination, overfitting, and mislabeled data before quoting a number.Make and defend trade-offs across accuracy, precision-vs-recall, cost, and latency, matched to how the model is actually used in production.Package models as containerized GPU/CPU microservices and ship them through GitOps (ArgoCD / Helm / Kapitan) across dev, QA, and production, with the right config and secrets handling.Define standards for evaluation, monitoring, and model lifecycle (MLOps / LLMOps), and build review tooling that lets domain experts validate and correct model output.Partner with product and domain experts to translate business rules and human-labeled data into robust, precision-first extraction; mentor engineers and set best practices.Required Experience & Skills 5+ years of professional AI engineering, including production ML/NLP systems.Strong Python and applied deep learning with PyTorch and Hugging Face Transformers — real experience with token classification, sequence labeling, and fine-tuning encoder models.Hands-on model training on GPU (CUDA, mixed precision, long-context), and comfort operating remote training boxes (SSH / Teleport, Azure or GCP).Practical experience with LLM extraction pipelines (Azure OpenAI / OpenAI): structured output, grounding to source text, and prompt design used as a baseline, not the finish line.Data-centric ML: dataset construction from raw stores (MongoDB), deduplication, leakage-aware train/test splitting, and label-quality analysis.Evaluation rigor: metric design, honest reporting, and an instinct for when a number is too good to be true (leakage, memorization, sparse or synthetic labels).Deployment fundamentals: Docker, Kubernetes, and GitOps (ArgoCD / Helm / Kapitan a strong plus).Clear communication of results, trade-offs, and uncertainty to both engineers and non-technical stakeholders.Nice to Have Domain exposure to trade finance, banking documents, or Letter-of-Credit processing.Document AI / OCR experience (hOCR, Azure Document Intelligence), table extraction, layout analysis, and document classification/routing.Serving/orchestration of ML services (durable queues, model facades) and cost/latency optimization for GPU inference.RAG / vector search where it genuinely fits — treated as one tool among many, not the default answer.Leadership Expectations Set technical direction for extraction and model-training workstreams.Review and approve modeling approaches, experiment designs, and evaluation protocols.Act as the escalation point for hard ML and data-quality problems.Uphold a culture of honest, reproducible evaluation — no headline number ships without provenance.Why Choose Us We’re a company where curiosity fuels everything - from bold ideas to personal growth.We believe common sense drives smart decisions. It keeps us focused and helps move fast, without the noise.We trust each other to deliver. Real ownership, no micromanagement.You’ll be heard here, even when you challenge the status quo. Because courage matters. Who dares — wins. Join us!