Senior Site Reliability Engineer

Lunit โ€” South Korea ยท Posted ~2 hours ago

Senior

Skills

site reliability engineering backend engineering application development operations system reliability Software as a Medical Device medical device compliance AI/ML systems backend systems

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary โœจ AIโ€‘Generated

Join a senior reliability engineering team responsible for keeping AI-powered medical software reliable across diverse deployment environments. You will bridge development and operations, improve application reliability and performance, and help ensure software meets stringent medical-device requirements.

Highlights

The role sits at the intersection of software development and operations, with a strong focus on reliability, performance, and production readiness. It offers the opportunity to work on regulated medical software and highly dependable AI-driven systems deployed across diverse environments.

Description

"Conquering cancer through AI" Lunit, a portmanteau of โ€˜Learning unit,' is a medical AI software company devoted to providing AI-powered total cancer care. Our AI solutions help discover cancer and predict cancer treatment outcomes, achieving timely and individually-tailored cancer treatment. ๐Ÿ—จ๏ธ About The Team Lunit's Software Engineering (hereafter, "SE") department develops the Lunit INSIGHT product. Within the SE department, the Core Product Engineering team performs the Backend/Application development needed to productize INSIGHT AI models. We translate product requirements into development requirements and build Software as a Medical Device (SaMD) that complies with medical device guidelinesOur team's core goal is to ensure the INSIGHT product has optimized performance and reliability so it can operate across diverse countries and deployment environments ๐Ÿ—จ๏ธ About The Position The Site Reliability Engineer role sits between development and operations, improving reliability, observability, automation, and operational systems so the INSIGHT product can run stablyRather than simply performing operational tasks, you will discover recurring problems and operational inefficiencies in production, analyze their root causes, and translate the findings into automation, improved observability, and improved operational processesYou will collaborate with Lunit International's SRE team, understand differing operational environments and processes, and connect and align the global operating system with the domestic product development environment ๐Ÿšฉ Roles & Responsibilities You will design and improve the reliability, observability, deployment automation, and cloud infrastructure operations of the Lunit INSIGHT product to ensure it runs stably. This role is not limited to performing predefined operational tasks. You will identify gaps in the team's current SRE capabilities and operational systems, then define and drive the improvements needed. Improve the stability, availability, performance, and operational quality of the Lunit INSIGHT product and related servicesDesign and enhance cloud infrastructure, deployment, monitoring, and operational automation systemsAnalyze cloud resource usage and cost, and continuously drive cost optimization while considering stability and performanceDiagnose, recover from, and perform root-cause analysis on incidents, and prevent recurrence through monitoring, alerting, runbooks, and automationReliably manage and improve operational configurations and security elements such as per-environment/per-customer settings, certificates, secrets, and access permissionsUnderstand the product domain and collaborate with development teams to build reliability in from the design and development stagesCollaborate in English with Lunit International's SRE team, understanding and aligning differing operational environments and processesBuild an on-call and incident-response system for 24/7 service operations, and participate in the actual on-call rotation to handle production incidents. As the team grows, evolve and lead a sustainable on-call operating model Requirements ๐ŸŽฏ Requirements 5+ years of experience in SRE, DevOps, Platform Engineering, or production infrastructure operationsExperience directly designing and operating Azure-based production environmentsHands-on experience with Linux, networking, and containersExperience designing or improving CI/CD, Infrastructure as Code, or operational automation systemsExperience analyzing production incidents based on monitoring, and performing recovery and recurrence-prevention activitiesExperience collaborating with product and development teams to balance stability and development velocityAbility to discuss and coordinate technical decisions fluently in English with overseas engineers and colleagues from diverse roles ๐Ÿ… Preferred Experiences Experience establishing and improving the technical direction or operating systems of the SRE or Platform domainExperience identifying the causes of recurring incidents or operational inefficiencies and leading structural improvements such as recurrence prevention or automationExperience building or operating SRE practices such as on-call, incident management, postmortems, and SLO/SLAExperience with Infrastructure as Code (Terraform, Bicep, ARM templates) and operational automation using Python, Bash, etcExperience operating container and deployment environments such as Kubernetes, GitHub Actions, and Azure DevOpsExperience using observability tools such as Azure Monitor, Application Insights, and Log Analytics, and incident-response/operations tools such as PagerDuty and ServiceNow ๐Ÿ“Œ Tech Stack Scripting/Language: Python, BashCloud/Infrastructure: Azure, Linux, Docker, Kubernetes, Terraform, BicepCI/CD: GitHub Actions, Azure DevOps PipelinesObservability/Operations: Azure Monitor, Application Insights, Log Analytics, PagerDuty, ServiceNow, Custom DashboardsProduct Environment: Python, Go, FastAPI, PostgreSQL, Redis, RabbitMQCollaboration: Jira, Confluence, Slack, GitHub ๐Ÿ“How To Apply CV/Resume in English, free formatPlease note that an English resume is required for this position๋ณธ ํฌ์ง€์…˜์€ ์˜๋ฌธ ์ด๋ ฅ์„œ ์ œ์ถœ์ด ํ•„์ˆ˜์ž…๋‹ˆ๋‹ค ๐Ÿƒโ€โ™€๏ธ Hiring Process Document Screening โ†’ Assignment โ†’ 1st Competency-based Interview (On-site) โ†’ 2nd Competency-based Interview (Online) โ†’ Culture-fit Interview (On-site) โ†’ OnboardingInterviews will be conducted in English, and detailed instructions for the Assignment will be provided separately. After the final interview, we may proceed with reference checks if needed ๐Ÿค Work Conditions and Environment Work type: Full-time (3-month probation)Work location : Lunit HQ(5F, 374, Gangnam-daero, Gangnam-gu, Seoul)Salary: After negotiation ๐ŸŽธ ETC If you misrepresent your experience or education or provide false or fraudulent information in or with your application, it may be grounds for cancellation of the employmentLunit is committed in providing the preferential processing to those eligible for employment protection (national merits and people with disabilities) relevant to related laws and regulations Benefits ๐ŸŒป Benefits & Perks The office is at a very convenient location, just a minute away from Gangnam Station Exit 3Meal Allowance is provided (up to 12,000 KRW per meal) when working at the officeLatest computer models, such as Macs and 4K monitors are provided and can be renewed every three yearsSeminar registration fees and book purchases are coveredRegular in-house AI and medical seminars are heldIn-house English lessons (aka Luniversal) are provided for English developmentAccess to high-quality AI learning resources & deep learning DevOps systemUp to 1.2 million KRW worth of benefits points can be claimed annuallyHoliday Allowances are provided in the form of gifts or vouchers for Korean National holidays, Seollal and ChuseokCongratulatory and Condolence allowances, along with paid time off are providedAnnual medical checkups and employee accident insurance are provided