Summary
โจ AIโGenerated
Join a team focused on enhancing the reliability and availability of global cloud services. Define and monitor service indicators, build robust observability stacks, and lead incident response and rootโcause analysis. Operate and improve Kubernetesโbased platforms while automating infrastructure with IaC tools. Collaborate closely with development and operations groups to ensure stable, highโperforming services.
Highlights
Enhance reliability and availability of global cloud services, define and monitor SLIs/SLOs, build observability stacks, lead incident response and rootโcause analysis, operate Kubernetes platforms, and automate infrastructure with IaC tools while collaborating with development and operations teams.
Description
๐ตOur Job (์์ผ๋ก ์ ํฌ์ ํจ๊ปํ์ค ์ผ๋ค์ด์์)๐ต
ํํ๋น์ ์ ๊ธ๋ก๋ฒ ํด๋ผ์ฐ๋ ์๋น์ค ๋ฐ ํ๋ซํผ์ ์์ ์ฑ๊ณผ ๊ฐ์ฉ์ฑ์ ํฅ์์ํค๊ณ ์ด์ํฉ๋๋ค.SLI/SLO๋ฅผ ์ ์ํ๊ณ ๋ชจ๋ํฐ๋ง ์ฒด๊ณ๋ฅผ ๊ตฌ์ถํ์ฌ ์๋น์ค ์ ๋ขฐ์ฑ์ ์ง์์ ์ผ๋ก ๊ฐ์ ํฉ๋๋ค.Datadog, PagerDuty, Prometheus, Grafana ๋ฑ์ ํ์ฉํด Observability ํ๊ฒฝ์ ๊ณ ๋ํํฉ๋๋ค.์ฅ์ ๋์, ์์ธ ๋ถ์ ๋ฐ ์ฌ๋ฐ ๋ฐฉ์ง ํ๋์ ์ํํฉ๋๋ค.์๋น์ค ๋ฐ ์ธํ๋ผ์ ์ฑ๋ฅ์ ๋ถ์ํ๊ณ ์ต์ ํํฉ๋๋ค.Amazon EKS ๊ธฐ๋ฐ Kubernetes ํ๋ซํผ์ ์ด์ยท๊ฐ์ ํฉ๋๋ค.Terraform, Helm ๊ธฐ๋ฐ IaC ํ๊ฒฝ์ ๊ตฌ์ถํ๊ณ ์ด์ ์๋ํ๋ฅผ ์ถ์งํฉ๋๋ค.DevOps ๋ฐ ๊ฐ๋ฐ ์กฐ์ง๊ณผ ํ์
ํ์ฌ ์์ ์ ์ธ ์๋น์ค ์ด์์ ์ง์ํฉ๋๋ค.
Improve and operate the reliability, availability, and scalability of Hanwha Vision's global cloud services and platforms.Define SLIs/SLOs and continuously improve service reliability through monitoring and observability.Build and enhance observability platforms using Datadog, PagerDuty, Prometheus, and Grafana.Lead incident response, root cause analysis, and preventive improvement initiatives.Analyze and optimize service and infrastructure performance.Operate and enhance Amazon EKS-based Kubernetes environments.Build and manage Infrastructure as Code using Terraform and Helm.Collaborate with development and DevOps teams to support reliable service operations.
๐ฅบOur Expectation (์ด๋ฐ ๋ถ๋ค์ ์ฐพ๊ณ ์์ด์)๐ฅบ
SRE, DevOps ๋๋ ์ธํ๋ผ ์์ง๋์ด๋ง ๋ถ์ผ์์ 8๋
์ด์์ ๊ฒฝ๋ ฅ์ ๋ณด์ ํ์ ๋ถํด๋ผ์ฐ๋ ํ๊ฒฝ์์ ๋๊ท๋ชจ ์๋น์ค ๋๋ ์ธํ๋ผ ์ด์ ๊ฒฝํ์ด ์์ผ์ ๋ถSLI/SLO ๊ธฐ๋ฐ ์๋น์ค ์ด์ ๋ฐ ์ ๋ขฐ์ฑ ๊ฐ์ ๊ฒฝํ์ด ์์ผ์ ๋ถDatadog, PagerDuty, Prometheus, Grafana ๋ฑ ๋ชจ๋ํฐ๋ง ๋๊ตฌ ํ์ฉ ๊ฒฝํ์ด ์์ผ์ ๋ถ์์คํ
์ฑ๋ฅ ๋ถ์, ์ฅ์ ๋์ ๋ฐ ํธ๋ฌ๋ธ์ํ
๊ฒฝํ์ด ์์ผ์ ๋ถKubernetes(EKS) ์ด์ ๊ฒฝํ์ด ์์ผ์ ๋ถTerraform ๊ธฐ๋ฐ Infrastructure as Code(IaC) ์ด์ ๊ฒฝํ์ด ์์ผ์ ๋ถGo ๋๋ Python์ ํ์ฉํ ์๋ํ ๊ฒฝํ์ด ์์ผ์ ๋ถ๊ธ๋ก๋ฒ ์กฐ์ง๊ณผ์ ํ์
์ด ๊ฐ๋ฅํ ์์ด ์ปค๋ฎค๋์ผ์ด์
์ญ๋์ ๋ณด์ ํ์ ๋ถ(์ฐ๋) Helm, Karpenter ํ์ฉ ๊ฒฝํ(์ฐ๋) AWS ๋๋ Azure ๊ธฐ๋ฐ ๊ธ๋ก๋ฒ ์๋น์ค ๊ตฌ์ถ ๋ฐ ์ด์ ๊ฒฝํ(์ฐ๋) Prometheus, Grafana, Loki ๊ธฐ๋ฐ Observability ํ๊ฒฝ ๊ตฌ์ถ ๊ฒฝํ(์ฐ๋) DynamoDB, Redis, Kafka/MSK ๋ฑ ๋ฐ์ดํฐ ํ๋ซํผ ์ด์ ๊ฒฝํ
8+ years of experience in SRE, DevOps, or Infrastructure EngineeringExperience operating large-scale cloud services or infrastructureExperience defining and managing SLIs/SLOs to improve service reliabilityHands-on experience with Datadog, PagerDuty, Prometheus, and GrafanaStrong troubleshooting, incident response, and performance optimization skillsExperience operating Kubernetes environments (EKS preferred)Experience managing Infrastructure as Code using TerraformProgramming experience with Go or Python for automationStrong English communication skills for collaboration with global teams(Preferred) Experience with Helm and Karpenter(Preferred) Experience building and operating global AWS or Azure environments(Preferred) Experience implementing observability solutions using Prometheus, Grafana, and Loki(Preferred) Experience with DynamoDB, Redis, Kafka/MSK, or similar data platforms(Preferred) Experience supporting security and compliance initiatives such as SOC2
* ๊ด๋ จ ํฌํธํด๋ฆฌ์ค๊ฐ ์์ผ์๋ค๋ฉด ์ ์ถ ์์ฒญ๋๋ฆฝ๋๋ค *
K8s helm, Terraform, Bicep, CDK, AWS | Azure Architecture DiagramCloud ๊ด๋ จ ๋ณธ์ธ ๊ธฐ์ ๋ธ๋ก๊ทธ | ๋
ธ์
| ํํ์ด์ง | YoutubeSite Reliability Engineer ์ ๊ด๋ จํ ์์ ์ ๋ํ๋ผ ์ ์๋ ๋ชจ๋ ๊ธฐ์ ์๋ฃ ๊ฐ๋ฅ์ด ์ธ ์์ ์ ๋ํ๋ผ ์ ์๋ ๋ชจ๋ ๊ธฐ์ ์๋ฃ ์ ์ถ ๊ฐ๋ฅ