AI Kernel Software Engineer

Furiosaai — South Korea · Posted ~1 day ago

Mid Full-time

Skills

Kernel programming Low-level systems programming AI accelerators GPU/NPU architectures Performance optimization Profiling C/C++ C++ GPU NPU

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary

A software engineering role focused on developing optimized kernels for AI hardware platforms, improving model performance, and creating tools for efficient AI workload execution.

Highlights

Work on advanced AI infrastructure, optimize high-performance computing workloads, and collaborate across research and engineering teams.

Description

About The Job Lead the integration of diverse AI models including VLA, Vision, and Multimodal architectures by utilizing our kernel programming language to ensure both accuracy and performance while keeping the stack ready for developers to use. Responsibilities Design and implement efficient kernels on FuriosaAI’s kernel programming stack (including vISA, TCL), targeting Tensor Contract Processor (TCP) architectures.Diagnose and optimize kernel performance with profiling tools and roofline analysis for each RNGD-accelerated AI model.Develop and apply automated kernel generation and optimization for AI workloads.Build diagnostic tools or testbeds for robust and reliable kernel validation.Drive end-to-end programming enablement on RNGDs, creating reproducible guides and reference implementations. Minimum Qualifications BS in Computer Science, Artificial Intelligence, Electrical Engineering, or a related field.Experience in low-level systems programming targeting XPU (e.g., NPU, GPU) architectures.Experience collaborating across engineering, research, and product teams to align software development with product requirements. Preferred Qualifications MS or PhD in Computer Science, Artificial Intelligence, Electrical Engineering, or a related field.Experience in optimizing high-performance kernels on AI accelerators (e.g., GPU, TPU) for AI products.Understanding of XPU architecture (computation patterns, data movement) and software-hardware co-optimization strategies.Experience in open-source or research projects on AI model architectures such as Diffusion, Mamba, and VLA.Experience in designing efficient deep learning architectures and developing algorithms for AI applications. Contact recruit@furiosa.ai