Site Reliability Engineer (Cloud Operations)

Liquidgroupsg — Singapore · Posted ~3 days ago

Senior Full-time

Skills

Site Reliability Engineering AWS AWS ECS AWS EKS Kubernetes AWS MSK Kafka Terraform AWS RDS Oracle PostgreSQL Observability APM Production operations 24/7 on-call New Relic

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary

Join an Infrastructure & Reliability team responsible for highly available, high-performance payment platforms. You will bridge software development and infrastructure operations, working with AWS container platforms, Kafka-based streaming, Terraform, database services, and advanced observability.

Highlights

Own reliability, observability, automation, and operational excellence for high-throughput payment infrastructure, working with modern AWS services and cross-functional engineering teams.

Description

Role Summary We are seeking a proactive and hands-on Site Reliability Engineer (SRE) to join our Infrastructure & Reliability team. The successful candidate will own the reliability, observability, automation, and operational excellence of Liquid Group’s payment platforms. You will work closely with Infrastructure, DevOps, Security, and Engineering teams to ensure our payment services meet stringent availability, performance, and compliance requirements. In this role, you will bridge the gap between software development and infrastructure operations. You will leverage New Relic for deep APM tracing and synthetic simulation, manage our microservices across AWS ECS & EKS, oversee high-throughput message streaming with AWS MSK (Kafka), drive Infrastructure as Code with Terraform, and ensure high availability for AWS RDS Oracle & PostgreSQL databases. Key Responsibilities 1. Production Operations Participate in a 24/7 on-call rotation.Monitor production systems and respond to incidents.Troubleshoot application, infrastructure, networking, and database issues.Coordinate incident response and lead RCA.2. Observability & Proactive Monitoring Configure New Relic APM, Infrastructure Monitoring, Logs, Distributed Tracing, Browser Monitoring, and Synthetics.Build dashboards, NRQL queries, alerts, and workflows.Develop synthetic payment journeys.Improve MTTD/MTTR and manage SLIs/SLOs.3. Cloud Infrastructure, Orchestration & Automation Manage AWS ECS, EKS, EC2, RDS, and MSK.Drive Infrastructure as Code with Terraform and CI/CD. Requirements & Qualifications Must-Have Technical Expertise 4+ years of SRE/DevOps experience in high-volume 24/7 production environments.Strong AWS (ECS, EKS, EC2, RDS, MSK), New Relic, Terraform, CI/CD, and Python/Go/Bash experience. Regulatory, Security & Mindset Knowledge of MAS TRM, DR, and PCI DSS.24/7 on-call mindset.Automation-first culture with continuous improvement.