Summary
✨ AI‑Generated
Build and improve production NLP and LLM systems covering retrieval-augmented generation, search, AI agents, document intelligence, and multilingual NLP. You will design retrieval pipelines, multi-step agent workflows, evaluation systems, and production APIs, while reducing hallucinations and improving reliability, latency, token efficiency, and inference cost. The stack includes Python, FastAPI, Docker, Git, testing, and CI/CD.
Highlights
Build production-grade NLP and LLM systems spanning retrieval, agents, document intelligence, multilingual processing, and evaluation. The role combines applied machine learning with production engineering and emphasizes quality, reliability, latency, and cost optimization.
Description
Responsibilities
Design, build, and improve production NLP/LLM systems, including RAG, search, AI agents, document intelligence, and multilingual NLP.
Build retrieval pipelines involving chunking, embeddings, hybrid search, filtering, reranking, and context construction.
Develop multi-step LLM workflows and agents with routing, tool use, structured outputs, retries, and failure recovery.
Build evaluation pipelines for retrieval and generation quality.
Analyze and reduce hallucinations and improve grounded generation and answer quality.
Create and clean datasets, perform error analysis, generate synthetic data, and apply model adaptation techniques such as SFT, LoRA, QLoRA, and PEFT when appropriate.
Develop production APIs and services using Python and tools such as FastAPI, Docker, Git, automated testing, and CI/CD.
Improve reliability, latency, token usage, inference cost, and overall system performance.
Run controlled experiments, regression tests, A/B tests, and quantitative evaluations.
Requirements
3+ years of professional experience in NLP, Machine Learning, Information Retrieval, Applied AI, or a related field.
Strong Python engineering skills and experience maintaining production software.
Hands-on experience shipping at least one LLM, RAG, NLP, search, or agent-based system into real production use.
Strong understanding of transformers, tokenization, embeddings, context windows, prompting, structured generation, and common LLM failure modes.
Practical understanding of retrieval beyond simply connecting an embedding model to a vector database.
Experience with semantic search, keyword search, hybrid retrieval, filtering, reranking, chunking, and indexing.
Experience with at least one modern LLM ecosystem such as OpenAI, Gemini, Anthropic, Hugging Face, or open-source models.
Experience with PyTorch or another modern deep-learning framework.
Ability to evaluate AI systems quantitatively rather than relying only on manual inspection.
Experience with API development, FastAPI, Docker, Git, automated testing, and CI/CD.
Professional working English for technical documentation and engineering communication.
Experience with BM25, RRF, Elasticsearch, Qdrant, Milvus, Pinecone, Weaviate, LangGraph, LangChain, LlamaIndex, Recall@K, MRR, nDCG, RAGAS, DeepEval, Langfuse, LoRA/QLoRA/DPO, vLLM, TGI, Triton, Graph RAG, OCR, or multilingual NLP is a strong advantage.
This role is probably not a fit if your experience is mainly limited to prompt engineering, ChatGPT usage, online courses/tutorials, simple “chat with PDF” projects, API/LangChain integrations without understanding the underlying architecture, or demo/hackathon projects that were never deployed to production.
Conditions
Full-time employment.
Location: Incheon, South Korea.
Work style: On-site / hybrid, depending on experience.
Work on real production AI systems involving RAG, agents, evaluation, model adaptation, and deployment.
Opportunity to influence architecture and technical direction.
Direct collaboration with founders and product leadership.
Access to high-performance compute, AI tools, and infrastructure.
Market-competitive compensation based on experience and demonstrated technical ability.
Significant opportunities for technical growth and increased responsibility.
When Applying
Please submit your CV/Resume and, if available, your GitHub, portfolio, publications, or technically relevant projects.
Please also briefly describe one NLP / LLM / Search system that you personally built or significantly contributed to, including:
What problem were you solving? What part did you personally own? What was the architecture? How did you evaluate whether it worked? What was the most important failure you discovered? What did you change because of that failure? Was the system used in production? If yes, at approximately what scale?
We are looking for engineers who can explain not only what they built, but also why the system works, why it fails, how they measured it, and what they changed to improve it.