Summary
Join an Infrastructure & Reliability team responsible for highly available, high-performance payment platforms. You will bridge software development and infrastructure operations, working with AWS container platforms, Kafka-based streaming, Terraform, database services, and advanced observability.
Highlights
Own reliability, observability, automation, and operational excellence for high-throughput payment infrastructure, working with modern AWS services and cross-functional engineering teams.
Description
Role Summary
We are seeking a proactive and hands-on Site Reliability Engineer (SRE) to join our Infrastructure & Reliability team.
The successful candidate will own the reliability, observability, automation, and operational excellence of Liquid Group’s payment platforms.
You will work closely with Infrastructure, DevOps, Security, and Engineering teams to ensure our payment services meet stringent availability, performance, and compliance requirements.
In this role, you will bridge the gap between software development and infrastructure operations.
You will leverage New Relic for deep APM tracing and synthetic simulation, manage our microservices across AWS ECS & EKS, oversee high-throughput message streaming with AWS MSK (Kafka), drive Infrastructure as Code with Terraform, and ensure high availability for AWS RDS Oracle & PostgreSQL databases.
Key Responsibilities
1.
Production Operations
Participate in a 24/7 on-call rotation.Monitor production systems and respond to incidents.Troubleshoot application, infrastructure, networking, and database issues.Coordinate incident response and lead RCA.2.
Observability & Proactive Monitoring
Configure New Relic APM, Infrastructure Monitoring, Logs, Distributed Tracing, Browser Monitoring, and Synthetics.Build dashboards, NRQL queries, alerts, and workflows.Develop synthetic payment journeys.Improve MTTD/MTTR and manage SLIs/SLOs.3.
Cloud Infrastructure, Orchestration & Automation
Manage AWS ECS, EKS, EC2, RDS, and MSK.Drive Infrastructure as Code with Terraform and CI/CD.
Requirements & Qualifications
Must-Have Technical Expertise
4+ years of SRE/DevOps experience in high-volume 24/7 production environments.Strong AWS (ECS, EKS, EC2, RDS, MSK), New Relic, Terraform, CI/CD, and Python/Go/Bash experience.
Regulatory, Security & Mindset
Knowledge of MAS TRM, DR, and PCI DSS.24/7 on-call mindset.Automation-first culture with continuous improvement.