Site Reliability Engineer

Longbridge Sg — Singapore · Posted ~1 day ago

Full-time

Skills

Site Reliability Engineering Distributed Systems System Reliability Monitoring Alerting Infrastructure as Code Terraform Ansible Helm Incident Response Root Cause Analysis Automation High Availability Infrastructure Security

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary

A rapidly growing financial technology organization is seeking an SRE to design, operate, and safeguard highly available distributed systems. The role focuses on infrastructure automation, monitoring and alerting, reliability engineering, global engineering collaboration, on-call operations, and continuous root-cause analysis.

Highlights

High-impact opportunity focused on highly available, secure distributed systems within a rapidly expanding financial technology environment. The role offers global collaboration, ownership of reliability, large-scale automation, and leadership of incident-response practices.

Description

Job Description About Us Longbridge is a fast-growing online brokerage platform on a mission to make investing smarter, simpler, and more accessible for everyone.As part of our global expansion, we’re looking for a hands-on Site Reliability Engineer (SRE) to design, scale, and safeguard the reliability of our next-generation financial platforms. This is a high-impact role where you’ll partner closely with product and engineering teams across the globe. What You’ll Do Own system reliability: Design, implement, and operate highly available, secure distributed systems to meet strict uptime and performance targets.Build automation at scale: Develop and enforce best practices in monitoring, alerting, and infrastructure-as-code (e.g., Terraform, Ansible, Helm).Partner globally: Work with development teams from design through deployment, ensuring reliability and resiliency are built in from day one.Lead incident response: Drive on-call processes, conduct root-cause analysis, and continuously reduce MTTR and failure recurrence.Future-proof our stack: Evaluate and adopt modern cloud-native technologies (e.g., Kubernetes, Prometheus, AWS) to keep systems secure and scalable.Stress-test and safeguard: Lead disaster recovery, chaos testing, and capacity planning for critical wealth management services What We’re Looking For 5+ years of experience in SRE, DevOps, or production engineering roles.Strong background in AWS (or GCP/Azure) and container orchestration (Docker, Kubernetes).Proficiency in at least one programming language (Python, Go, or similar) for automation and tooling.Solid Linux administration skills and experience with CI/CD pipelines.Proven ability in incident management and troubleshooting distributed systems.Strong collaboration and communication skills across global teams.Bilingual proficiency in English and Chinese (Mandarin preferred) is highly valued, given the need to collaborate effectively with stakeholders across Southeast Asia and the Greater China region.Experience supporting regulated financial systems is a strong plus.