Site Reliability Engineer - HFT

Selby Jennings โ€” Hong Kong Sar ยท Posted ~4 hours ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

Role Overview The firm is seeking a Mid-Level Site Reliability Engineer to support mission-critical production systems and observability platforms. This role combines SRE responsibilities with data platform engineering, focusing on system reliability, monitoring, automation, cloud infrastructure, and real-time data pipelines. You will work closely with engineering teams to improve platform stability, operational visibility, and scalability across a complex distributed environmen Key Responsibilities Own and enhance end-to-end observability across production systems.Design actionable monitoring dashboards, alerts, and operational metrics.Investigate production incidents, diagnose root causes, and implement permanent fixes.Manage deployments through CI/CD pipelines and GitOps workflows.Operate and optimize Kubernetes-based workloads in cloud environments.Provision and maintain infrastructure using Infrastructure-as-Code practices.Build and support near real-time data pipelines and telemetry platforms.Design data quality controls and optimize analytical databases for large-scale querying.Develop and improve SQL-based analytics and reporting capabilities.Participate in a shared on-call rotation supporting critical production services. Key Requirements Degree in Computer Science, Software Engineering, or equivalent practical experience.Experience building and maintaining observability and monitoring solutions within production environments.Strong understanding of incident management, troubleshooting, and root cause analysis.Hands-on experience with monitoring tools such as Prometheus, Grafana, ELK, Datadog, CloudWatch, or similar.Experience deploying software through CI/CD platforms such as GitLab CI, Jenkins, or GitHub Actions.Exposure to distributed stream processing technologies such as Flink, Spark Structured Streaming, or comparable frameworks.Experience working with OLAP or columnar databases such as ClickHouse, BigQuery, Druid, or Redshift.Practical Kubernetes experience, including deployments, resource management, and application configuration.Strong programming skills in Python, Java, Scala, Rust, or similar languages.Advanced SQL skills with experience building and optimising analytical workloads.Knowledge of messaging platforms and event-driven architectures such as Kafka or SQS Nice to Have Experience with OpenTelemetry and distributed tracing.Knowledge of Airflow or workflow orchestration platforms.Advanced Terraform and Infrastructure-as-Code expertise.Familiarity with JVM build tooling or large-scale software delivery pipelines.Prior exposure to trading, market data, financial technology, or other low-latency environments.