Senior Systems Design Engineer – Data Center GPU

Amd — Canada · Posted ~11 hours ago

Senior Full-time

Skills

Systems design GPU architecture Data center systems Hardware engineering System validation GPU High-performance computing AI accelerators Cloud computing

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a senior systems engineering team focused on next-generation data-center accelerators. You will help improve system design and delivery for high-performance computing and AI workloads, working across complex hardware and system-level engineering challenges.

Highlights

Contribute to the design and delivery of next-generation data-center GPU systems in a technically advanced engineering environment. The role offers exposure to high-performance computing, AI, cloud infrastructure, and complex system-level engineering.

Description

ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we’re looking for talent who feel the same: people who want to leave the planet better than they found it, those who don’t shy away from humanity’s challenges but are determined to help solve them. AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger — technology that moves the world forward. THE ROLE:  We are looking for a dynamic, energetic Sr. Systems Design Engineer to join our growing Data Center GPU team. As a key contributor to the success of AMD’s product, you will be part of a leading team to drive and improve AMD’s abilities to deliver the highest quality, industry-leading technologies to market. The Systems Design Engineering team fosters and encourages continuous technical innovation to showcase successes as well as facilitate continuous career development.    THE PERSON:  In this role, you will drive balanced, scalable, and automated solutions. In this high visibility position, your software systems engineering expertise will be necessary towards Product development, definition, and root cause resolution.    Key Responsibilities p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Executing and contributing to post-silicon validation efforts for SOC-level IP blocks, including test plan development, test execution, coverage tracking, and issue reporting across program milestonesp]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Debugging hardware and system-level issues found during bring-up, validation, and production phases of SOC programs, with a focus on system IP blocks such as DMA engines, interrupt controllers, and data path logicp]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Performing post-silicon debug analysis using scan dump tools and debug reports to investigate hangs, stalls, and error conditions at the IP and system levelp]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Triaging failures by collecting and correlating debug data across multiple IP domains, with guidance from senior engineers on complex cross-chiplet issuesp]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Working with multiple teams and tracking test execution to make sure all features are validated and optimized on timep]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Working closely with design, firmware, driver, software runtime, and kernel teams to understand IP behavior, error propagation, software-hardware interactions, and power management dependenciesp]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Developing and maintaining validation test content targeting error handling, fault injection, and power management scenariosp]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Engaging in hardware/software modeling and debug frameworks to reproduce and root-cause silicon failuresp]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Participating in collaborative triage and debug efforts across multiple teams and IP domains, progressively taking ownership of specific IP areas Preferred Experience Experience in post-silicon validation, hardware debug, or a related semiconductor engineering roleProgramming/scripting skills (e.g., C/C++, Python, Perl)p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Debug techniques and methodologies for post-silicon validationp]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Experience with board/platform-level debugging, including bring-up, sequencing, analysis, and optimizationp]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Knowledge of SoC system architecture, including multi-die or chiplet-based designsp]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Understanding of cache hierarchies, buffer allocation and management, ring buffers, FIFO structures, credit-based flow control, and tag tracking mechanismsp]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Familiarity with DMA, interrupt handling, memory subsystem, or I/O subsystem architecturesp]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Knowledge of memory management concepts including address translation, virtual memory, TLB operation, and page fault handlingp]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Familiarity with software runtime environments, kernel-level drivers, or OS-level interfaces that interact with hardware IPsp]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Exposure to RAS concepts — error detection, poison propagation, machine check logging, and watchdog mechanismsp]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Experience with scan dump analysis or JTAG-based post-silicon debug toolsp]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Strong analytical/problem-solving skills and pronounced attention to detailp]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item">Must be a self-starter, able to independently drive tasks to completion and willing to ramp up on new IP domains through documentation and hands-on debug ACADEMIC CREDENTIALS:  Bachelors or Masters degree in electrical or computer engineering LOCATION: Markham, ON Benefits offered are described: AMD benefits at a glance. AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here. This posting is for an existing vacancy.