Senior Triton Compiler and Kernel Software Engineer

Amd — United States · Posted ~6 hours ago

Senior Full-time

Skills

Triton GPU programming compiler development Python high-performance computing AI software stack GPU Compiler AI frameworks

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior engineering role focused on improving GPU computing capabilities through compiler development and high-performance kernel optimization. The position involves working with open-source technologies, AI workloads, and advanced software-hardware integration.

Highlights

Opportunity to work on advanced AI infrastructure, GPU performance optimization, and cutting-edge compiler technologies with significant technical impact.

Description

ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career. The Role We are seeking a Senior Triton Compiler and Kernel Engineer to advance Triton performance and capabilities on AMD GPUs. Triton is an open-source language and compiler for developing high-performance GPU kernels in Python. It is a critical layer in the AI software stack, connecting frameworks and workloads to GPU hardware. Triton is strategic to AMD’s AI roadmap, and AMD is investing fully in making it a first-class platform for current and future AMD GPUs. You will work across GPU architecture, compilers, kernels, multi-GPU communication, and AI frameworks while contributing to upstream Triton and AMD’s ROCm software stack. The Person The ideal candidate has strong experience in several of the following areas: GPU architecture and programmingCompiler development and optimizationHigh-performance GPU kernelsMulti-GPU communication and collective operationsDistributed AI training and inferenceAI workload and framework performanceLow-level performance analysis You can reason across the stack—from distributed AI algorithms and Triton programs to compiler transformations, communication libraries, generated instructions, and GPU hardware. Key Responsibilities Develop and optimize Triton compiler support for AMD GPUs.Improve compiler lowering, optimization, scheduling, and code generation.Create high-performance kernels for attention, GEMM, MoE, and other AI workloads.Develop and optimize multi-GPU kernels and communication primitives.Enable new AMD GPU and interconnect capabilities through effective Triton abstractions.Analyze compute, memory, communication, occupancy, register usage, and generated code.Resolve complex correctness and performance issues across single- and multi-GPU workloads.Collaborate with GPU architecture, ROCm, PyTorch, and AI framework teams.Contribute designs and implementations to upstream Triton and LLVM/MLIR.Provide technical leadership, code reviews, and mentoring. Preferred Experience Experience with AMD GPU architecture and ROCm is highly desirable.Experience optimizing kernels with Triton, HIP, CUDA, or GPU assembly.Experience with Triton, LLVM, MLIR, or another optimizing compiler.Knowledge of GPU execution models, memory hierarchies, synchronization, and instruction pipelines.Experience with collective communication, distributed programming, and libraries such as RCCL or NCCL.Understanding of communication topologies, interconnects, synchronization, and communication–computation overlap.Familiarity with distributed training, inference, tensor parallelism, expert parallelism, or pipeline parallelism.Experience using AI coding tools and autonomous agents to accelerate software development, debugging, benchmarking, and kernel performance tuning.Familiarity with AI primitives, reduced-precision formats, and performance profiling.Contributions to complex or open-source software projects.Strong analytical, debugging, communication, and collaboration skills. Academic Credentials Bachelor’s or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent This role is not eligible for visa sponsorship. Benefits offered are described: AMD benefits at a glance. AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here. This posting is for an existing vacancy.