Software Engineer, Site Reliability

Huxley — Japan · Posted ~2 hours ago

Senior Full-time

Skills

Cloud platform reliability SLIs/SLOs Production readiness Infrastructure scaling Database scaling Observability Monitoring Distributed tracing Logging Incident response Root cause analysis On-call operations CI/CD Deployment workflows Developer tooling Automation Cloud platforms Databases AI LLM

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a high-impact reliability engineering team responsible for a large-scale, multi-region cloud platform. You will define reliability standards, improve observability and incident response, scale infrastructure and databases, automate operational work, and enhance CI/CD and developer tooling. The role is ideal for an engineer who enjoys distributed systems, production engineering, automation, and applying modern AI technologies to operational challenges.

Highlights

Own reliability for a large-scale, multi-region cloud platform, with strong ownership across observability, incident management, automation, infrastructure scaling, and developer experience. The role offers significant technical impact and opportunities to improve availability and engineering efficiency.

Description

Key Responsibilities Platform Reliability Own and improve the reliability of a large-scale cloud platform operating across multiple regions.Define and manage SLIs/SLOs, production readiness reviews, and reliability standards.Scale infrastructure and databases to support rapid user growth and global expansion.Drive improvements in service availability, incident response, and operational excellence.Observability & Incident Management Build and maintain observability frameworks, including monitoring, distributed tracing, and logging.Lead incident response, root cause analysis, and post-incident reviews.Participate in on-call rotations and reduce operational burden through automation and AI-powered solutions.Developer Experience & Platform Enablement Enhance CI/CD pipelines, deployment workflows, and developer tooling.Build self-service platforms and automation that improve engineering velocity.Leverage AI and LLM-based technologies to optimize development and operational workflows.Improve onboarding, development environments, and overall engineering productivity.Cross-Functional Leadership Partner with engineering, product, and operations teams across multiple regions and time zones.Align reliability initiatives with business objectives and customer experience goals.Influence engineering best practices and help establish a strong reliability culture across the organization. Requirements Strong experience operating cloud-native production environments, ideally with GCP or AWS and Kubernetes.Proven SRE, DevOps, or Platform Engineering experience supporting large-scale systems.Experience with observability platforms, incident management, and reliability metrics (SLIs/SLOs).Solid understanding of databases, distributed systems, and modern web architectures.Experience improving developer productivity through tooling, automation, or platform engineering initiatives.Hands-on exposure to AI/LLM technologies for engineering workflows is highly desirable.Strong communication skills and ability to work in cross-functional, international environments.