Principal Site Reliability Engineer

Node Nine Limited — Hong Kong Sar · Posted ~23 hours ago

Lead Full-time Onsite

Skills

SRE cloud infrastructure infrastructure engineering observability reliability engineering secure infrastructure scalable systems cross-functional collaboration Cloud Observability

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A principal-level SRE will design, implement, and maintain reliable, scalable, and secure infrastructure for mission-critical client systems. The role emphasizes observability, operational excellence, direct reliability ownership, and embedding modern SRE practices into cross-functional engineering teams.

Highlights

Principal-level opportunity in a technically mature environment, with ownership of reliability and observability for mission-critical systems and close collaboration with engineering teams.

Description

Company Description Node Nine Limited is an independent consultancy focused on cloud-native observability and site reliability engineering for mission-critical systems. The team partners closely with client engineering groups to embed modern operational practices and improve reliability across complex environments. By implementing robust tooling and clear processes, Node Nine helps organizations gain confidence in running and scaling their production systems. Applicants can expect to work in a technically mature environment that values collaboration, continuous improvement, and practical solutions to real-world reliability challenges. Role Description The Principal SRE Engineer role is a full-time, on-site position based in Hong Kong SAR. In this role, you will design, implement, and maintain reliable, scalable, and secure infrastructure for client systems, with a strong focus on embedding SRE practices into cross-functional teams, having direct ownership of reliability, observability and performance of the systems. You will work hands-on with engineering teams to improve deployment pipelines, define and monitor service-level objectives, automate operations, and drive incident response and blameless postmortems. Day-to-day responsibilities include mentoring engineers on SRE best practices and establishing SRE practices in an environment with explicit executive mandate and teams volunteering for adoption. You will contribute to internal frameworks, tools, and documentation that support consistent, high-quality operational standards across client engagements. Qualifications Principal-level Site Reliability Engineering skills, including reliability design, incident management, and observability practice definition experience.Proficiency in troubleshooting complex distributed systems, performance bottlenecks, and production incidents in a highly collaborative environment.Solid Software Development experience in one or more programming languages (e.g., Python, Go, Java) and familiarity with CI/CD pipelines.Hands-on System Administration background, including Linux-based systems, networking fundamentals, and security-conscious operations.Experience with Infrastructure design and automation using tools such as Kubernetes, Terraform, cloud platforms (e.g., AWS, Azure), gitops (argocd) and configuration management.Proven ability to lead technical initiatives, mentor engineers, and collaborate effectively with cross-functional teams and clients.Strong communication skills with the ability to explain complex technical concepts to diverse stakeholders.Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience; prior consultancy or client-facing experience is an advantage. Hi