Lead Site Reliability Engineer
Tgjapan — Japan · Posted ~3 days ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
Compensation: ¥9M – ¥12M JPY per year
Employment Type: Full-time
Experience Level: Senior
About the Role
We are seeking an experienced Lead Site Reliability Engineer (SRE) to join a global engineering team responsible for ensuring the reliability, scalability, and performance of large-scale distributed systems serving millions of users worldwide.
In this role, you will provide technical leadership for mission-critical services, drive SRE best practices, and lead initiatives focused on automation, observability, incident management, and system optimization.
You will work closely with cross-functional engineering teams to improve operational excellence and maintain high service availability across a global cloud platform.
Responsibilities
Define and drive Service Level Objectives (SLOs) and Service Level Agreements (SLAs).Establish and manage error budgets to balance reliability and feature delivery.Lead performance optimization initiatives, including latency reduction and scalability improvements.Own incident management processes and coordinate major incident response.Conduct root cause analysis (RCA) and implement long-term preventive solutions.Design and improve monitoring, alerting, logging, and observability platforms.Drive automation initiatives to reduce operational overhead and improve engineering efficiency.Develop operational standards, runbooks, and reliability best practices.Mentor engineers and provide technical leadership across SRE initiatives.Collaborate with software engineering, infrastructure, and platform teams to improve service reliability.Participate in on-call rotations and continuously improve operational processes.
Required Qualifications
- Bachelor's degree in Computer Science or related field, or equivalent practical experience.
- More than 7 years of hands-on experience in SRE, infrastructure engineering, or a related field, with demonstrated technical leadership experience.
- Experience building and operating production systems in GCP (Google Cloud Platform).
- Extensive experience designing, building, operating, and scaling Kubernetes environments.
- Deep knowledge and hands-on experience building and operating modern monitoring, alerting, and logging tools (e.g., Prometheus, Grafana, ELK Stack, Datadog).
- In-depth knowledge of UNIX-like operating system internals and/or networking.
- Deep knowledge of IP network systems and protocols (TCP/IP, HTTP, etc.) and hands-on troubleshooting experience.
- Experience building automated workflows using CI/CD tools (e.g., Jenkins, CircleCI, GitLab, CI/CD).
- Experience developing operational automation tools and scripts using scripting languages such as Shell, Python, etc.
- Proven track record of leading production incident handling end-to-end (detection, triage, short-term / long-term fix, root cause analysis).
- Experience in system performance tuning and capacity planning.
- Proficiency with Git and GitHub for version control and collaboration.
- Strong communication, negotiation, and collaboration skills to articulate complex technical issues and align with internal and external stakeholders.
Preferred Qualifications
- Experience developing or maintaining GCP environments (e.g., GKE, Cloud Run, BigQuery, Cloud Monitoring, IAM).
- Experience in web application development.
- Deep knowledge and practical experience in observability, and a strong drive to improve services leveraging SLIs/SLOs.
- Experience implementing and operating error budgets, or a proven track record in toil reduction initiatives.
- Experience driving cross-team or org-wide reliability improvements (e.g., defining standards, leading postmortem culture).
- Experience working with cross-cultural global teams in different locations.
Language Requirements
English: FluentJapanese: Optional
Work Environment
Flexible working hours with core collaboration hoursInternational engineering team
Hybrid Position
Apply now or contact us for further information:
TG Japan Inc.
03-5775-6618
RSU1_Agt@tg-hr.com
We have 86,246 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 86,246 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume