Site Reliability Engineer

Pinpoint Asia — Hong Kong Sar · Posted ~4 hours ago

Mid

Skills

Kubernetes Amazon EKS AWS Terraform Grafana Prometheus Loki ELK Datadog CI/CD GitOps SQL Observability Data pipelines

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A financial technology organization is seeking a mid-level SRE to own production reliability and observability. You will operate Kubernetes and cloud infrastructure, build actionable monitoring and alerts, participate in on-call rotations, automate deployments, and engineer reliable near-real-time data pipelines.

Highlights

Own end-to-end observability and reliability for production platforms, work with cloud infrastructure and Kubernetes, and build high-performance data pipelines for large-scale systems.

Description

Our client, a financial trading company, is looking for a mid-level SRE who owns system reliability and observability. What You’ll Do SRE & Observability: Own platform observability end-to-end. Design Grafana dashboards and actionable alerts that pinpoint root causes. Manage production systems on Kubernetes (EKS), diagnose silent failures, and run AWS infrastructure via Terraform. Participate in on-call rotations for the data platform and drive deployments through CI/CD (GitOps).Data Pipeline Engineering: Build and maintain near-real-time pipelines for telemetry and trading logs. Implement robust delivery guarantees and data quality checks to quarantine bad records. Optimize columnar OLAP tables for fast, billion-row queries using advanced partitioning, deduplication, and tuned SQL. What You Need Degree in CS, Software Engineering, or equivalent practical experience.Proven experience with production observability stacks (Prometheus, Grafana, Loki, ELK, Datadog), CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins), and Kubernetes operations (pod management, resource limits).Hands-on experience with distributed stream processing engines (Flink, Spark Streaming) and columnar OLAP databases (ClickHouse, Redshift, BigQuery). Familiarity with message queues (Kafka, SQS).Strong skills in Python, Scala, Java, or Rust, alongside fluent analytical SQL. Nice to Have OpenTelemetry, distributed tracing, and handling high-cardinality telemetry.Workflow orchestration with Apache Airflow.Advanced Infrastructure as Code (Terraform modules, drift management).Modern build systems (Maven, Gradle, sbt, Cargo).