Site Reliability Engineer

Hanwha Vision โ€” South Korea ยท Posted ~4 hours ago

Mid Full-time

Skills

Observability Monitoring Incident response Root cause analysis Kubernetes Terraform Helm DevOps Datadog PagerDuty Prometheus Grafana

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary โœจ AIโ€‘Generated

Join a team focused on enhancing the reliability and availability of global cloud services. Define and monitor service indicators, build robust observability stacks, and lead incident response and rootโ€‘cause analysis. Operate and improve Kubernetesโ€‘based platforms while automating infrastructure with IaC tools. Collaborate closely with development and operations groups to ensure stable, highโ€‘performing services.

Highlights

Enhance reliability and availability of global cloud services, define and monitor SLIs/SLOs, build observability stacks, lead incident response and rootโ€‘cause analysis, operate Kubernetes platforms, and automate infrastructure with IaC tools while collaborating with development and operations teams.

Description

๐ŸŽตOur Job (์•ž์œผ๋กœ ์ €ํฌ์™€ ํ•จ๊ป˜ํ•˜์‹ค ์ผ๋“ค์ด์—์š”)๐ŸŽต ํ•œํ™”๋น„์ „์˜ ๊ธ€๋กœ๋ฒŒ ํด๋ผ์šฐ๋“œ ์„œ๋น„์Šค ๋ฐ ํ”Œ๋žซํผ์˜ ์•ˆ์ •์„ฑ๊ณผ ๊ฐ€์šฉ์„ฑ์„ ํ–ฅ์ƒ์‹œํ‚ค๊ณ  ์šด์˜ํ•ฉ๋‹ˆ๋‹ค.SLI/SLO๋ฅผ ์ •์˜ํ•˜๊ณ  ๋ชจ๋‹ˆํ„ฐ๋ง ์ฒด๊ณ„๋ฅผ ๊ตฌ์ถ•ํ•˜์—ฌ ์„œ๋น„์Šค ์‹ ๋ขฐ์„ฑ์„ ์ง€์†์ ์œผ๋กœ ๊ฐœ์„ ํ•ฉ๋‹ˆ๋‹ค.Datadog, PagerDuty, Prometheus, Grafana ๋“ฑ์„ ํ™œ์šฉํ•ด Observability ํ™˜๊ฒฝ์„ ๊ณ ๋„ํ™”ํ•ฉ๋‹ˆ๋‹ค.์žฅ์•  ๋Œ€์‘, ์›์ธ ๋ถ„์„ ๋ฐ ์žฌ๋ฐœ ๋ฐฉ์ง€ ํ™œ๋™์„ ์ˆ˜ํ–‰ํ•ฉ๋‹ˆ๋‹ค.์„œ๋น„์Šค ๋ฐ ์ธํ”„๋ผ์˜ ์„ฑ๋Šฅ์„ ๋ถ„์„ํ•˜๊ณ  ์ตœ์ ํ™”ํ•ฉ๋‹ˆ๋‹ค.Amazon EKS ๊ธฐ๋ฐ˜ Kubernetes ํ”Œ๋žซํผ์„ ์šด์˜ยท๊ฐœ์„ ํ•ฉ๋‹ˆ๋‹ค.Terraform, Helm ๊ธฐ๋ฐ˜ IaC ํ™˜๊ฒฝ์„ ๊ตฌ์ถ•ํ•˜๊ณ  ์šด์˜ ์ž๋™ํ™”๋ฅผ ์ถ”์ง„ํ•ฉ๋‹ˆ๋‹ค.DevOps ๋ฐ ๊ฐœ๋ฐœ ์กฐ์ง๊ณผ ํ˜‘์—…ํ•˜์—ฌ ์•ˆ์ •์ ์ธ ์„œ๋น„์Šค ์šด์˜์„ ์ง€์›ํ•ฉ๋‹ˆ๋‹ค. Improve and operate the reliability, availability, and scalability of Hanwha Vision's global cloud services and platforms.Define SLIs/SLOs and continuously improve service reliability through monitoring and observability.Build and enhance observability platforms using Datadog, PagerDuty, Prometheus, and Grafana.Lead incident response, root cause analysis, and preventive improvement initiatives.Analyze and optimize service and infrastructure performance.Operate and enhance Amazon EKS-based Kubernetes environments.Build and manage Infrastructure as Code using Terraform and Helm.Collaborate with development and DevOps teams to support reliable service operations. ๐ŸฅบOur Expectation (์ด๋Ÿฐ ๋ถ„๋“ค์„ ์ฐพ๊ณ  ์žˆ์–ด์š”)๐Ÿฅบ SRE, DevOps ๋˜๋Š” ์ธํ”„๋ผ ์—”์ง€๋‹ˆ์–ด๋ง ๋ถ„์•ผ์—์„œ 8๋…„ ์ด์ƒ์˜ ๊ฒฝ๋ ฅ์„ ๋ณด์œ ํ•˜์‹  ๋ถ„ํด๋ผ์šฐ๋“œ ํ™˜๊ฒฝ์—์„œ ๋Œ€๊ทœ๋ชจ ์„œ๋น„์Šค ๋˜๋Š” ์ธํ”„๋ผ ์šด์˜ ๊ฒฝํ—˜์ด ์žˆ์œผ์‹  ๋ถ„SLI/SLO ๊ธฐ๋ฐ˜ ์„œ๋น„์Šค ์šด์˜ ๋ฐ ์‹ ๋ขฐ์„ฑ ๊ฐœ์„  ๊ฒฝํ—˜์ด ์žˆ์œผ์‹  ๋ถ„Datadog, PagerDuty, Prometheus, Grafana ๋“ฑ ๋ชจ๋‹ˆํ„ฐ๋ง ๋„๊ตฌ ํ™œ์šฉ ๊ฒฝํ—˜์ด ์žˆ์œผ์‹  ๋ถ„์‹œ์Šคํ…œ ์„ฑ๋Šฅ ๋ถ„์„, ์žฅ์•  ๋Œ€์‘ ๋ฐ ํŠธ๋Ÿฌ๋ธ”์ŠˆํŒ… ๊ฒฝํ—˜์ด ์žˆ์œผ์‹  ๋ถ„Kubernetes(EKS) ์šด์˜ ๊ฒฝํ—˜์ด ์žˆ์œผ์‹  ๋ถ„Terraform ๊ธฐ๋ฐ˜ Infrastructure as Code(IaC) ์šด์˜ ๊ฒฝํ—˜์ด ์žˆ์œผ์‹  ๋ถ„Go ๋˜๋Š” Python์„ ํ™œ์šฉํ•œ ์ž๋™ํ™” ๊ฒฝํ—˜์ด ์žˆ์œผ์‹  ๋ถ„๊ธ€๋กœ๋ฒŒ ์กฐ์ง๊ณผ์˜ ํ˜‘์—…์ด ๊ฐ€๋Šฅํ•œ ์˜์–ด ์ปค๋ฎค๋‹ˆ์ผ€์ด์…˜ ์—ญ๋Ÿ‰์„ ๋ณด์œ ํ•˜์‹  ๋ถ„(์šฐ๋Œ€) Helm, Karpenter ํ™œ์šฉ ๊ฒฝํ—˜(์šฐ๋Œ€) AWS ๋˜๋Š” Azure ๊ธฐ๋ฐ˜ ๊ธ€๋กœ๋ฒŒ ์„œ๋น„์Šค ๊ตฌ์ถ• ๋ฐ ์šด์˜ ๊ฒฝํ—˜(์šฐ๋Œ€) Prometheus, Grafana, Loki ๊ธฐ๋ฐ˜ Observability ํ™˜๊ฒฝ ๊ตฌ์ถ• ๊ฒฝํ—˜(์šฐ๋Œ€) DynamoDB, Redis, Kafka/MSK ๋“ฑ ๋ฐ์ดํ„ฐ ํ”Œ๋žซํผ ์šด์˜ ๊ฒฝํ—˜ 8+ years of experience in SRE, DevOps, or Infrastructure EngineeringExperience operating large-scale cloud services or infrastructureExperience defining and managing SLIs/SLOs to improve service reliabilityHands-on experience with Datadog, PagerDuty, Prometheus, and GrafanaStrong troubleshooting, incident response, and performance optimization skillsExperience operating Kubernetes environments (EKS preferred)Experience managing Infrastructure as Code using TerraformProgramming experience with Go or Python for automationStrong English communication skills for collaboration with global teams(Preferred) Experience with Helm and Karpenter(Preferred) Experience building and operating global AWS or Azure environments(Preferred) Experience implementing observability solutions using Prometheus, Grafana, and Loki(Preferred) Experience with DynamoDB, Redis, Kafka/MSK, or similar data platforms(Preferred) Experience supporting security and compliance initiatives such as SOC2 * ๊ด€๋ จ ํฌํŠธํด๋ฆฌ์˜ค๊ฐ€ ์žˆ์œผ์‹œ๋‹ค๋ฉด ์ œ์ถœ ์š”์ฒญ๋“œ๋ฆฝ๋‹ˆ๋‹ค * K8s helm, Terraform, Bicep, CDK, AWS | Azure Architecture DiagramCloud ๊ด€๋ จ ๋ณธ์ธ ๊ธฐ์ˆ  ๋ธ”๋กœ๊ทธ | ๋…ธ์…˜ | ํ™ˆํŽ˜์ด์ง€ | YoutubeSite Reliability Engineer ์— ๊ด€๋ จํ•œ ์ž์‹ ์„ ๋‚˜ํƒ€๋‚ผ ์ˆ˜ ์žˆ๋Š” ๋ชจ๋“  ๊ธฐ์ˆ  ์ž๋ฃŒ ๊ฐ€๋Šฅ์ด ์™ธ ์ž์‹ ์„ ๋‚˜ํƒ€๋‚ผ ์ˆ˜ ์žˆ๋Š” ๋ชจ๋“  ๊ธฐ์ˆ  ์ž๋ฃŒ ์ œ์ถœ ๊ฐ€๋Šฅ