Summary
✨ AI‑Generated
Work as an SRE close to a high-performance electronic trading stack, helping maintain reliable, scalable infrastructure for latency-sensitive production systems operating across global markets. You will design intelligent monitoring and automated operational workflows, identify abnormal system behavior before it affects production, and partner closely with software engineers. Sponsorship is available.
Highlights
On-site SRE opportunity in a high-performance global trading environment, focused on resilient production infrastructure, intelligent monitoring, automation, rapid incident response, and latency-sensitive systems, with sponsorship available.
Description
Site Reliability Engineer (Trading Infrastructure) Multiple organisations
Sponsorship Available
Sydney or Hong Kong | Onsite | Global Trading Environment
Join a high-performance engineering team responsible for the reliability, scalability and operational excellence of the infrastructure powering a global electronic trading platform.
This is a hands-on Site Reliability Engineering role sitting close to the trading stack, where you'll help build resilient production systems, improve automation, and work alongside software engineers to support latency-sensitive applications operating across global financial markets.
Unlike traditional SRE environments focused primarily on SLI/SLO metrics, this team takes an event-driven approach to reliability engineering.
You'll design intelligent monitoring and automated operational workflows that identify abnormal system behaviour, infrastructure anomalies and production events before they impact trading.
The focus is on actionable signals, rapid diagnosis and engineering-led remediation rather than simply measuring service health.
What You'll Be Doing
Design, build and maintain highly reliable Linux-based production infrastructure.Develop Infrastructure as Code using Terraform and Ansible.Build observability platforms using Prometheus, Grafana and Splunk.Create event-driven monitoring, intelligent alerting and automated remediation workflows.Improve operational tooling and incident response across business-critical trading systems.Work with PostgreSQL and InfluxDB supporting production data platforms.Support workflow orchestration using Prefect.Administer virtualisation and storage platforms including CEPH, VMware, KVM and Proxmox.Partner closely with software engineers developing C++, Java, C# and Python applications.Improve CI/CD, deployment tooling and overall platform reliability through automation.
What We're Looking For
Strong Linux systems administration and troubleshooting experience.Python software development skills, specifically building tools for internal teamsExperience with Infrastructure as Code using Terraform, Ansible or similar tools.Hands-on experience with observability platforms such as Prometheus, Grafana or Splunk.Experience designing monitoring strategies that focus on operational events, anomaly detection and actionable alerts rather than purely SLI/SLO-driven metrics.Familiarity with PostgreSQL, InfluxDB or other production databases.Strong networking fundamentals including TCP/IP, routing and firewalls.Experience with Docker and modern development tooling.Exposure to virtualisation platforms such as VMware, KVM, Proxmox or CEPH.Experience supporting low-latency, distributed or mission-critical production environments is highly regarded.
Why This Role?
Work on infrastructure that directly supports real-time trading.Solve complex reliability challenges where milliseconds matter.Influence how monitoring, automation and operational engineering are built from the ground up.Collaborate with experienced infrastructure and software engineers in a highly technical environment.Work in a culture that values engineering ownership, continuous improvement and pragmatic problem solving.
If you're passionate about Linux, automation, observability and building resilient production platforms, we'd love to hear from you.