Senior Site Reliability Engineer

Kodypay — Hong Kong Sar · Posted ~3 hours ago

Senior

Skills

site reliability engineering incident management production operations observability Kubernetes cloud infrastructure databases messaging systems SLO management incident response

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A global technology organization is looking for a Senior Site Reliability Engineer to take end-to-end ownership of reliability, availability, scalability, and operational excellence. You will lead major incident response, improve observability and SLOs, and operate Kubernetes, databases, messaging systems, and cloud infrastructure supporting mission-critical services.

Highlights

Own reliability and operational excellence for mission-critical systems at global scale. The role provides senior-level ownership of observability, incident response, SLOs, cloud infrastructure, and complex production environments.

Description

Job Summary Kody is seeking a Senior Site Reliability Engineer (8+ years of experience) to drive the reliability, availability, scalability, and operational excellence of our global payment platform. Based in Hong Kong or Shenzhen, you will take end-to-end ownership of production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating across Europe, Asia, and North America. Key Responsibilities Incident Management & On-Call: Participate in a follow-the-sun production on-call rotation as a senior incident responder. Lead incident management during SEV1/SEV2 events to optimize MTTR and operational effectivenessProduction Operations: Diagnose, triage, mitigate, and coordinate the resolution of complex production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructureSLO & Reliability Engineering: Define, implement, and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes across distributed servicesContinuous Optimization: Drive systemic reliability improvements through infrastructure automation, observability enhancement, capacity planning, performance tuning, and post-incident root-cause analysis (RCA)Security & Compliance: Partner with global engineering teams to strengthen architectural resilience, security posture, and operational maturity in PCI-DSS-regulated payment environmentsTechnical Leadership: Mentor junior engineers, eliminate operational toil through automation, and influence engineering teams to adopt resilience-by-design practices Requirements Qualifications & Requirements Experience: 8+ years of hands-on experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting high-availability, mission-critical production systemsCore Technical Stack: Strong expertise in AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms (e.g., Datadog, Prometheus, Grafana)Distributed Systems Mastery: Deep understanding of distributed systems architecture, high availability, disaster recovery, capacity planning, and microservices orchestrationDomain Expertise: Proven track record operating in payment, banking, fintech, or other highly regulated environments with strict PCI-DSS, security, and uptime standardsSRE Methodology: Deep knowledge of core SRE principles, including SLO/SLI design, error budget management, alert governance, and toil reductionLocation & Communication: Based in Hong Kong or Shenzhen. Excellent command of English (written and spoken) to lead cross-functional incident responses and collaborate seamlessly with global teams Leadership & Operational Excellence Ownership: Demonstrates strong end-to-end accountability for service reliability and customer impact under high pressureStructured Problem Solving: Applies a systematic and data-driven approach to troubleshooting, telemetry analysis, and incident resolution in complex distributed environmentsCrisis Management: Proven ability to command cross-functional incident response efforts, align stakeholders, and maintain clear communication during critical outagesEngineering Culture: Champions a blameless post-incident culture, operational readiness, continuous learning, and technical mentorship Benefits Competitive Package A dynamic and innovative team Collaborative, inclusive working environment