GPU Systems Engineer

Tossglobal โ€” South Korea ยท Posted ~2 weeks ago

Mid Full-time

Skills

Linux system engineering large-scale server infrastructure GPU infrastructure HPC or AI clusters infrastructure automation OS/kernel/driver troubleshooting high-performance networking storage design infrastructure monitoring GPU CUDA NCCL NVLink RDMA InfiniBand RoCE HPC

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary โœจ AIโ€‘Generated

Join a systems engineering team responsible for designing and operating large-scale GPU infrastructure for AI and ML workloads. You will build GPU clusters from bare metal, evaluate new architectures, standardize operating systems and software stacks, optimize high-speed networking and storage, automate operations, and plan capacity expansion.

Highlights

Work on large-scale GPU infrastructure for AI and ML workloads, shape next-generation infrastructure roadmaps, evaluate cutting-edge hardware through PoCs and benchmarks, optimize high-performance clusters, and automate operations.

Description

ํ•จ๊ป˜ํ•˜๊ฒŒ ๋  ํŒ€์— ๋Œ€ํ•ด ์•Œ๋ ค๋“œ๋ ค์š” ํ† ์Šค์˜ System Platform ํŒ€์€ Network Engineer, InfraOps Engineer์™€ ํ•จ๊ป˜ Infra Engineering Tribe์— ์†ํ•ด ์žˆ์–ด์š”.System Platform ํŒ€์€ IDC ๋ฒ ์–ด๋ฉ”ํƒˆ๋ถ€ํ„ฐ ํผ๋ธ”๋ฆญ ํด๋ผ์šฐ๋“œ๊นŒ์ง€, ์„œ๋น„์Šค๊ฐ€ ์˜ฌ๋ผ๊ฐ€๋Š” ์‹œ์Šคํ…œ ๊ธฐ๋ฐ˜์„ ์ง์ ‘ ์„ค๊ณ„ํ•˜๊ณ  ์šด์˜ํ•ด์š”.์ด ํฌ์ง€์…˜์€ AI/ML ์›Œํฌ๋กœ๋“œ๊ฐ€ ์•ˆ์ •์ ์œผ๋กœ ๋™์ž‘ํ•  ์ˆ˜ ์žˆ๋Š” GPU ์ธํ”„๋ผ๋ฅผ ๋‹ด๋‹นํ•ด์š”.์–ด๋–ค GPU๋ฅผ ์–ด๋–ค ๊ตฌ์กฐ๋กœ ๋„์ž…ํ•˜๊ณ  ์–ด๋–ป๊ฒŒ ํ™•์žฅํ• ์ง€ ํŒ๋‹จํ•˜๋ฉฐ ์ฐจ์„ธ๋Œ€ AI ์ธํ”„๋ผ ๋กœ๋“œ๋งต์„ ํ•จ๊ป˜ ๋งŒ๋“ค์–ด๊ฐ€์š”.ํ•ฉ๋ฅ˜ํ•˜๋ฉด ํ•จ๊ป˜ ํ•  ์—…๋ฌด์˜ˆ์š” GPU ์ปดํ“จํŒ… ํด๋Ÿฌ์Šคํ„ฐ ์ธํ”„๋ผ๋ฅผ ์„ค๊ณ„ํ•˜๊ณ , ๋ฒ ์–ด๋ฉ”ํƒˆ๋ถ€ํ„ฐ ๊ตฌ์ถ•ํ•ด ์šด์˜ํ•ด์š”.์ตœ์‹  GPU ๋ผ์ธ์—…๊ณผ ์•„ํ‚คํ…์ฒ˜ ๋ณ€ํ™”๋ฅผ ์ถ”์ ํ•˜๊ณ , PoC์™€ ๋ฒค์น˜๋งˆํฌ๋กœ ๋„์ž…์„ ๊ฒ€์ฆํ•˜๋ฉฐ ์ธํ”„๋ผ ๋กœ๋“œ๋งต์„ ์„ธ์›Œ์š”.GPU ์„œ๋ฒ„์˜ ํ‘œ์ค€ OS์™€ ์†Œํ”„ํŠธ์›จ์–ด ์Šคํƒ์„ ๊ตฌ์„ฑํ•˜๊ณ , ๋Œ€๊ทœ๋ชจ ๋…ธ๋“œ๋ฅผ ์žฌํ˜„ ๊ฐ€๋Šฅํ•˜๊ฒŒ ํ”„๋กœ๋น„์ €๋‹ํ•ด์š”.๊ณ ์„ฑ๋Šฅ ๋„คํŠธ์›Œํฌ์™€ ์Šคํ† ๋ฆฌ์ง€๋ฅผ ์„ค๊ณ„ํ•ด ํด๋Ÿฌ์Šคํ„ฐ ์„ฑ๋Šฅ์„ ์ตœ์ ํ™”ํ•ด์š”.GPU ์ž์›์„ ๋ชจ๋‹ˆํ„ฐ๋งํ•˜๊ณ , ์žฅ์• ๋ฅผ ์ง„๋‹จํ•˜๊ณ  ๊ทผ๋ณธ ์›์ธ์„ ์ฐพ์•„ ๋Œ€์‘ํ•˜๋ฉฐ ์šด์˜์„ ์ž๋™ํ™”ํ•ด์š”.์‚ฌ์šฉ๋ฅ ๊ณผ ์šฉ๋Ÿ‰ ์ง€ํ‘œ๋ฅผ ๊ทผ๊ฑฐ๋กœ ์ฆ์„ค ์‹œ์ ์„ ํŒ๋‹จํ•ด ์ „๋ ฅ, ๋ƒ‰๊ฐ, ๋ž™ ๋“ฑ ์ธํ”„๋ผ ํ™˜๊ฒฝ์„ ๊ณ ๋ คํ•˜์—ฌ ์•ˆ์ •์ ์œผ๋กœ ์ฆ์„คํ•ด์š”.์ด๋Ÿฐ ๋ถ„๊ณผ ํ•จ๊ป˜ํ•˜๊ณ  ์‹ถ์–ด์š”. ์‹œ์Šคํ…œ ์—”์ง€๋‹ˆ์–ด๋กœ 3๋…„ ์ด์ƒ, ๋Œ€๊ทœ๋ชจ Linux ์„œ๋ฒ„ ์ธํ”„๋ผ๋ฅผ ์šด์˜ํ•ด ๋ณด์‹  ๋ถ„์ด๋ฉด ์ข‹์•„์š”.OS, ์ปค๋„, ๋“œ๋ผ์ด๋ฒ„ ์ค‘ ํ•œ ์˜์—ญ ์ด์ƒ์„ ๊นŠ์ด ๋‹ค๋ค„ ๋ฌธ์ œ๋ฅผ ํ•ด๊ฒฐํ•ด ๋ณธ ๋ถ„์ด๋ฉด ์ข‹์•„์š”.GPU ์„œ๋ฒ„๋ฅผ ํ™œ์šฉํ•œ HPC ๋˜๋Š” AI ํด๋Ÿฌ์Šคํ„ฐ๋ฅผ ๊ตฌ์ถ•ํ•˜๊ณ  ์šด์˜ํ•ด ๋ณธ ๊ฒฝํ—˜์ด ์žˆ์œผ์‹  ๋ถ„์ด๋ฉด ์ข‹์•„์š”.๋ฐ˜๋ณต๋˜๋Š” ์šด์˜์„ ์ฝ”๋“œ๋กœ ์ž๋™ํ™”ํ•˜๊ณ , ํ‘œ์ค€๊ณผ ๋ฌธ์„œ๋กœ ๋‚จ๊ฒจ ๋ณธ ๋ถ„์ด๋ฉด ์ข‹์•„์š”.์—ฌ๋Ÿฌ ํŒ€๊ณผ ๊ธฐ์ˆ ์ ์ธ ๋…ผ๋ฆฌ์— ๊ธฐ๋ฐ˜ํ•ด ํ˜‘์—…ํ•  ์ˆ˜ ์žˆ๋Š” ๋ถ„์ด๋ฉด ์ข‹์•„์š”.์ด๋Ÿฐ ๊ฒฝํ—˜์ด ์žˆ๋‹ค๋ฉด ๋” ์ข‹์•„์š” CUDA, NCCL, NVLink, RDMA์™€ ๊ฐ™์€ GPU ์ปดํ“จํŒ… ๊ธฐ์ˆ ์— ๋Œ€ํ•œ ์ดํ•ด๊ฐ€ ์žˆ์œผ๋ฉด ๋” ์ข‹์•„์š”.InfiniBand ๋˜๋Š” RoCE ๊ธฐ๋ฐ˜ ๊ณ ์† ๋„คํŠธ์›Œํฌ๋ฅผ ์„ค๊ณ„ํ•˜๊ณ  ์šด์˜ํ•ด ๋ณธ ๊ฒฝํ—˜์ด ์žˆ์œผ๋ฉด ๋” ์ข‹์•„์š”.ํŽŒ์›จ์–ด, BIOS, BMC ๋“ฑ ํ•˜๋“œ์›จ์–ด ๋ ˆ๋ฒจ์„ ์ง์ ‘ ๋‹ค๋ค„๋ณธ ๊ฒฝํ—˜์ด ์žˆ์œผ๋ฉด ๋” ์ข‹์•„์š”.DCGM, Prometheus, Grafana ๋“ฑ์œผ๋กœ ๋ชจ๋‹ˆํ„ฐ๋ง ์ฒด๊ณ„๋ฅผ ๊ตฌ์ถ•ํ•ด ๋ณธ ๊ฒฝํ—˜์ด ์žˆ์œผ๋ฉด ๋” ์ข‹์•„์š”.AI/ML ์กฐ์ง๊ณผ ํ˜‘์—…ํ•˜๊ฑฐ๋‚˜ MLOps ํŒŒ์ดํ”„๋ผ์ธ์„ ์ดํ•ดํ•˜๊ณ  ์žˆ์œผ๋ฉด ๋” ์ข‹์•„์š”.IDC ์ „๋ ฅ, ๋ƒ‰๊ฐ, ๋ž™ ๋“ฑ ๋ฌผ๋ฆฌ ์ธํ”„๋ผ๋ฅผ ๊ณ ๋ คํ•ด ๋Œ€๊ทœ๋ชจ ์ปดํ“จํŒ… ํ™˜๊ฒฝ์„ ์„ค๊ณ„ํ•ด ๋ณธ ๊ฒฝํ—˜์ด ์žˆ์œผ๋ฉด ๋” ์ข‹์•„์š”.