LLM Application Engineer

Bolder Apps — Uzbekistan · Posted ~3 hours ago

Senior Contract Remote

Skills

LLM application development structured extraction classification multimodal AI prompt engineering evaluation quality control cost optimization Google Gemini GCP Firebase LLMs

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A mid-to-senior LLM application engineering role focused on turning unstructured emails, documents, web content, and images into reliable structured data. You will build production AI pipelines, establish quality gates and evaluations, and optimize model cost and latency across modern cloud services.

Highlights

Remote monthly-retainer role building production-grade AI pipelines for real-world applications. Offers hands-on work with multimodal LLMs, structured outputs, measurable evaluation, quality targets, and cost and latency optimization.

Description

We're hiring a mid-senior LLM Application Engineer on a remote monthly retainer to design, ship, and harden production AI pipelines for client products at Bolder Apps. You'll own structured extraction and classification systems that turn messy real-world inputs (email, HTML, PDFs, images) into reliable product data, with measurable quality gates, evals, and cost control. You'll work on Firebase / GCP-style backends with product, mobile, and QA. We want someone who has shipped LLM apps for real users, not demos. You should be strong across modern LLMs and especially fluent with Google Gemini (multimodal prompts, structured outputs, failure modes, and cost/latency tradeoffs), with solid experience on other major providers too. If you can hit hard quality targets, keep dollars-per-run honest, and leave runbooks another engineer can pick up, we want to talk. About Us Bolder Apps is a product development studio that partners with US-based startups and established companies to build and scale innovative digital products. We specialize in AI-powered development, full-cycle product creation, and engineering team augmentation. Our mission is simple: build bolder, faster, and smarter. Our Culture & Values (read before applying) We move fast. We take ownership. We work with AI, not against it. And we expect everyone to bring ideas, not wait for instructions. There are no daily checklists, no micromanagement, and no corporate politics. Instead, you'll have autonomy, trust, and a team that's always ready to help you grow. At Bolder Apps, impact matters more than titles, and curiosity matters more than seniority. If you want a place where you can level up fast and actually see your work making a difference - welcome aboard. Requirements Responsibilities Own production LLM pipelines end to end: ingestion, multimodal model calls, structured records, and storage, including confidence flags, retries, and idempotent rescansDesign prompt and schema strategies (including schema-aligned or constrained outputs) so results are consistent and product-readyBuild classification and filtering layers on top of extraction (taxonomy mapping, demographic or audience filters, deduplication, and related cleanup logic)Define and run evaluation harnesses (golden sets, regression suites, online metrics) so quality does not regress when prompts, models, or parsers changeHit and report against hard quality targets (precision-style gates for completeness, duplicates, incorrect inclusions, image presence, and similar product SLAs)Optimize token usage, model tiering, caching, and batching to keep dollars-per-run and latency under controlHarden reliability for long-running async jobs (timeouts, partial recovery, memory limits, safe production deploys)Partner with Flutter / mobile and QA on field contracts, review queues, and incident debuggingDocument architecture and runbooks so ownership is shared, not a single point of failureStay current on Gemini and peer LLM APIs; recommend when to swap models, add fallbacks (e.g. document AI), or tighten schemasShipped LLM applications in production (not demos only): prompts, structured outputs, retries, observability, and real failure handlingStrong hands-on experience with Google Gemini, including multimodal (text + image / document-style) workflows and structured extractionPractical experience with at least one other major LLM stack (OpenAI, Anthropic, or similar) and good judgment on when to use whichStructured extraction from messy inputs: HTML, PDFs, images, and mixed email-like contentClassification / taxonomy systems on top of LLM outputsEvaluation discipline: offline evals, regression suites, and production quality metrics tied to clear acceptance criteriaCost and latency awareness: token budgeting, cheaper tiers, caching, batching; can explain dollars-per-run tradeoffs to a PMPython backend experience on serverless cloud (Cloud Functions or equivalent) and document stores (e.g. Firestore) or similar GCP patternsEnglish at C1 or above for client-adjacent debugging with a PMUS hours overlap through roughly 5 PM EST when live coordination is neededOwnership habits: honest estimates, early blockers, finished releases Nice to have Schema-aligned LLM frameworks (BAML, Instructor, Outlines, or similar)Google Document AI or other OCR / document intelligence as a fallback pathGmail API / OAuth products and restricted-scope compliance familiarityComputer vision for product-image quality checksFirebase + GCP ops (secrets, regions, schedulers, cost monitoring)Building eval corpora from real production data and iterating until contractual SLAs passAgency or multi-client studio experience Benefits Fully remote and async-friendly, with required overlap through :5 PM EST when client or release coordination needs itMonthly retainer structure with recurring AI pipeline work for engineers who keep production quality and cost honestReal autonomy over how you structure prompts, schemas, evals, and deploys. We do not hand you a rigid playbookDirect line to PMs, mobile engineers, and decision-makersTooling budget for the LLM and cloud tools you need to move fastA peer network of product-minded builders across overlapping client projects