Summary
✨ AI‑Generated
Join a high-impact reliability engineering team responsible for a large-scale, multi-region cloud platform. You will define reliability standards, improve observability and incident response, scale infrastructure and databases, automate operational work, and enhance CI/CD and developer tooling. The role is ideal for an engineer who enjoys distributed systems, production engineering, automation, and applying modern AI technologies to operational challenges.
Highlights
Own reliability for a large-scale, multi-region cloud platform, with strong ownership across observability, incident management, automation, infrastructure scaling, and developer experience. The role offers significant technical impact and opportunities to improve availability and engineering efficiency.
Description
Key Responsibilities
Platform Reliability
Own and improve the reliability of a large-scale cloud platform operating across multiple regions.Define and manage SLIs/SLOs, production readiness reviews, and reliability standards.Scale infrastructure and databases to support rapid user growth and global expansion.Drive improvements in service availability, incident response, and operational excellence.Observability & Incident Management
Build and maintain observability frameworks, including monitoring, distributed tracing, and logging.Lead incident response, root cause analysis, and post-incident reviews.Participate in on-call rotations and reduce operational burden through automation and AI-powered solutions.Developer Experience & Platform Enablement
Enhance CI/CD pipelines, deployment workflows, and developer tooling.Build self-service platforms and automation that improve engineering velocity.Leverage AI and LLM-based technologies to optimize development and operational workflows.Improve onboarding, development environments, and overall engineering productivity.Cross-Functional Leadership
Partner with engineering, product, and operations teams across multiple regions and time zones.Align reliability initiatives with business objectives and customer experience goals.Influence engineering best practices and help establish a strong reliability culture across the organization.
Requirements
Strong experience operating cloud-native production environments, ideally with GCP or AWS and Kubernetes.Proven SRE, DevOps, or Platform Engineering experience supporting large-scale systems.Experience with observability platforms, incident management, and reliability metrics (SLIs/SLOs).Solid understanding of databases, distributed systems, and modern web architectures.Experience improving developer productivity through tooling, automation, or platform engineering initiatives.Hands-on exposure to AI/LLM technologies for engineering workflows is highly desirable.Strong communication skills and ability to work in cross-functional, international environments.