Senior Software Engineer - LLM Inference Engine

Furiosaai — South Korea · Posted ~3 days ago

Senior

Skills

inference engine development LLM inference optimization multimodal LLMs performance optimization distributed systems NPU computing LLMs inference engines NPUs tensor parallelism KV cache

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join an advanced AI systems team developing a high-performance inference engine for large and multimodal language models running on specialized accelerators. You will research and implement state-of-the-art techniques for throughput, latency, memory efficiency, parallelism, scheduling, and distributed execution while collaborating closely with compiler and hardware engineers.

Highlights

Work on next-generation inference infrastructure for large and multimodal language models, with deep exposure to cutting-edge optimization techniques. The role offers close collaboration with compiler and hardware specialists and focuses on throughput, latency, memory efficiency, and scalable production systems.

Description

About The Job Software Engineer (Inference Engine) is responsible for developing and optimizing a high-performance inference engine for Large Language Models (LLMs) and multimodal LLMs running on FuriosaAI NPUs. In this role, you will proactively research and apply the state-of-the-art inference optimization techniques to our inference engine. You will work in close collaboration with the compiler and hardware teams to enhance the engine's performance to its full potential. Responsibilities Design and implement FuriosaAI’s next-generation inference engine for large and multimodal language models—comparable in capability to frameworks such as vLLM and SGLang—optimized for throughput, latency, and memory efficiency.Design and implement advanced inference optimizations—such as speculative decoding, KV-cache management, tensor/model parallelism, memory-efficient execution, and scheduling—in our production inference engine.Design and develop capabilities for distributed and scalable inference, including prefill–decode (PD) and encode–prefill–decode (EPD) disaggregation, disaggregated speculative decoding, and hierarchical and external KV-cache storage such as HiCache and Mooncake.Collaborate closely with the Compiler team to co-design and optimize execution for FuriosaAI NPUs, improving system-level throughput, latency, and memory utilization.Proactively research, evaluate, and integrate state-of-the-art inference optimization techniques and key features of LLM serving frameworks into our production inference engine. Minimum Qualifications BS degree in Computer Science, Engineering, or a related field, with at least 3 years of relevant industry experience, or equivalent practical experienceProficiency in Rust or C++ programming skillKnowledge and passion of deep learning, LLM, and/or generative AI modelsExcellent problem-solving and data analysis skills.Strong communication and collaboration skills. Preferred Qualifications Experience in building inference serving systems for large models, encompassing batching, scheduling, caching, and load balancing.A deep understanding of performance optimization systems.Proficiency in C++/CUDA or Triton kernel developmentContributions to open-source inference frameworks such as vLLM, SGLang, or TensorRT-LLM. Contact recruit@furiosa.ai