Description
Sr.
Director, Site Reliability Engineering
Coupang operates one of the largest and most complex technology platforms in the world.
We are seeking a Senior Director, Site Reliability Engineering (Head of SRE) to define and lead company-wide reliability, resilience, scalability, and operational excellence.
This leader will transform reliability from a collection of team-specific practices into platform mechanisms that services inherit by tier, while advancing incident response toward an intelligent, AI-assisted, and increasingly autonomous operating model.
We are looking for a visionary, industry-recognized technology leader who has previously conceived, built, and scaled a comparable SRE, production engineering, resilience, or autonomous-operations organization at a leading global technology company.
The successful candidate must combine deep technical credibility with the organizational leadership required to align executives, influence architecture across the company, and build a world-class leadership bench.
Key Responsibilities
Set a bold, multi-year vision for company-wide reliability, resilience, and autonomous operations, and translate that vision into an executable roadmap with measurable business outcomes.Define and own the SRE strategy, operating model, engineering standards, and reliability governance across Coupang.Build platform mechanisms that allow services to inherit reliability requirements based on service tier rather than recreate them independently.Lead initiatives that materially improve availability, resilience, scalability, performance, and operational readiness.Partner with engineering, product, infrastructure, security, finance, and business leaders to align reliability investments with customer and business priorities.Own executive reliability metrics, including availability, detection and recovery performance, change risk, incident recurrence, capacity readiness, and operational toil.Build and scale a world-class SRE organization capable of influencing engineering practices across the company.
Reliability Strategy, SLOs & Engineering Governance
Establish and evolve service-tier definitions, SLOs, SLAs, error budgets, reliability scorecards, and objective certification mechanisms such as RBD/RBO.Create clear reliability requirements for Tier 0, Tier 1, and Tier 2 services, including redundancy, load testing, disaster recovery, observability, and incident response.Ensure reliability governance is embedded in architecture, development, release, and production operations rather than applied as a final review.Drive systematic reduction of recurring incidents, reliability risks, operational debt, and unsafe change patterns.Influence company-wide architecture for graceful degradation, fault isolation, load shedding, circuit breaking, and failure containment.
Incident Management & Autonomous Operations
Transform incident management into a fast, disciplined, data-driven, and increasingly autonomous operating model.Enable AI-assisted detection, event correlation, triage, escalation, root-cause drafting, remediation recommendations, and selected guardrailed auto-remediation.Improve incident command, on-call quality, escalation mechanisms, communication, post-incident learning, and corrective-action completion.Reduce noisy alerts, manual on-call work, repeated diagnosis, and time spent coordinating across fragmented systems.Use incident and telemetry data to continuously improve platform standards, testing, capacity models, and engineering roadmaps.
Disaster Recovery, Resilience & Capacity
Own the strategy and execution model for disaster recovery, regional resilience, availability-zone loss, capacity-constrained recovery, and critical business continuity.Build reusable DR and failover mechanisms that services inherit from the platform rather than implement as bespoke projects.Establish objective RPO/RTO targets, automated readiness gates, regular game days, fault injection, and evidence-based recovery certification.Drive proactive and intelligent capacity management using forecasting, reservations, workload prioritization, and automated response to demand and failure scenarios.Partner with compute, traffic, networking, storage, and application leaders to enable safe zone evacuation, regional failover, and surge readiness.
Observability, Testing & Reliability Intelligence
Partner with Observability and TestOps leaders to integrate logs, metrics, traces, continuous profiling, testing, and incident intelligence into one reliability feedback loop.Ensure every critical service has actionable telemetry, meaningful SLOs, release-quality signals, and production-readiness evidence.Use production incidents and operational patterns to drive targeted integration, load, resilience, and regression testing.Establish executive reliability dashboards that provide trusted views of service health, risk, capacity, and operational effectiveness.
Talent Leadership & Organization
Lead multiple layers of SRE leaders, including senior managers, directors, principal engineers, and senior individual contributors.Own organizational design, global hiring strategy, leadership development, succession planning, and the creation of a strong leadership bench.Attract exceptional SRE, distributed systems, resilience, incident-management, and capacity-engineering talent from best-in-class technology organizations.Build an empowered organization with clear accountability, strong technical judgment, high execution velocity, and a company-wide perspective.Act as a force multiplier by mentoring technical and organizational leaders and raising reliability capabilities across engineering.
Technical Leadership & Architecture
Own reliability architecture decisions across large-scale distributed systems and cloud-native infrastructure.Define resilient patterns for redundancy, failover, traffic management, data recovery, workload prioritization, and dependency isolation.Guide architecture reviews and platform standards for safe scaling, fault tolerance, and operational simplicity.Balance availability, customer impact, engineering velocity, cost, and operational complexity in major technical decisions.Maintain sufficient technical depth to challenge assumptions, guide principal engineers, and make high-quality decisions during critical incidents.
Execution & Impact
Deliver measurable improvements in availability, time to detect, time to mitigate, time to recover, incident recurrence, change-failure rate, and on-call burden.Create disciplined operating rhythms, milestones, ownership models, and quarterly targets for strategic reliability programs.Increase adoption of common reliability mechanisms and reduce team-specific implementations and manual operations.Demonstrate business impact through improved customer experience, reduced outage exposure, stronger peak readiness, and more efficient use of infrastructure capacity.Build credibility through predictable delivery, transparent risk management, and objective evidence of reliability improvement.
Essential Qualifications
Leadership experience in a best-in-class SRE, production engineering, infrastructure reliability, or cloud operations organization at hyperscaler, major cloud provider, global marketplace, leading fintech, or similarly scaled technology company.Experience building an SRE practice comparable in maturity to leading industry organizations, rather than operating a traditional support or operations function renamed as SRE.Experience with Kubernetes, service mesh, cloud-native platforms, traffic engineering, and large-scale capacity management.Experience with chaos engineering, fault injection, regional resilience, and automated disaster recovery.Experience building AI-assisted operations, incident intelligence, predictive reliability, or self-healing systems.15+ years of experience in software engineering, infrastructure engineering, distributed systems, or site reliability engineering.8+ years leading large-scale, multi-layer engineering organizations, including senior managers, directors, and senior individual contributors.Demonstrated experience personally defining the vision and leading the architecture, build-out, launch, and scaled adoption of a company-wide SRE, reliability, resilience, or autonomous-operations program.Prior experience building reliability systems and operating practices for high-scale, high-availability, customer-critical distributed systems.Deep expertise in SLOs, error budgets, observability, incident management, disaster recovery, capacity planning, and resilience engineering.Proven ability to lead through major incidents while also creating durable mechanisms that prevent recurrence.Recognized as a visionary technology and organizational leader who can influence executive stakeholders, align multiple engineering organizations, and attract exceptional talent.Proven ability to convert long-term strategy into measurable execution and company-wide adoption.