Summary
✨ AI‑Generated
A reliability engineering role responsible for operating and improving massive-scale data infrastructure. The position focuses on automation, production operations, monitoring, and maintaining highly available distributed platforms.
Highlights
Work on large-scale data infrastructure, production reliability, automation, and distributed systems supporting high-volume environments.
Description
Site Reliability Engineer – Data Infrastructure
Sydney | Quantitative Trading | Linux, Kafka, Python
We are working with a leading global quantitative trading firm that is growing its Data Engineering team in Sydney.
This is an SRE role focused on the infrastructure that powers large-scale data platforms.
You will help operate a multi-petabyte environment supporting millions of queries each day, working across Linux, distributed data systems, automation and production reliability.
This is not a traditional data engineering role focused on writing pipelines, DAGs or analytics workloads.
The team owns and operates the underlying platforms themselves.
What you'll be doing
Operate and improve large-scale distributed data infrastructureWork with technologies including Kafka, HDFS, Kubernetes and distributed query platformsOwn monitoring, alerting, incident response and production reliabilityAutomate infrastructure deployment, upgrades and operational processes using PythonTroubleshoot Linux, networking, storage and distributed systems issuesImprove capacity, resilience, failure handling and deployment processesWork directly with traders, researchers and developers to solve data infrastructure problemsParticipate in an on-call rotation and engineer out recurring issues
What we're looking for
Hands-on Linux systems administration and troubleshooting experienceExperience owning production systems, including monitoring, incidents and on-callOperator-side experience with at least one of Kafka, HDFS or KubernetesPython experience, ideally for infrastructure automation or operational toolingAn understanding of networking, storage, processes, memory and system performanceExperience with infrastructure automation, CI/CD or configuration management
Experience with Kafka or HDFS administration is particularly valuable, but you do not need to know the entire technology stack.
Engineers coming from SRE, infrastructure, platform engineering, systems engineering or production engineering backgrounds are encouraged to apply.