Site Reliability Engineer, Data & Observability Platform

Pulsar Crypto — Hong Kong Sar · Posted ~3 hours ago

Senior

Skills

SRE principles Observability Grafana Alerting Distributed Systems Streaming Data Pipelines CI/CD GitOps Production Operations Troubleshooting Documentation Monitoring

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Become a site reliability engineer in a cutting‑edge crypto platform, shaping end‑to‑end observability, creating insightful dashboards, and keeping mission‑critical systems running smoothly.

Highlights

Innovative site reliability engineer role focusing on end‑to‑end observability, building actionable Grafana dashboards, and managing production systems in a fast‑moving crypto environment.

Description

Target profile You will love this job if: You are pioneering and innovative and want to be part of the cutting-edge and disruptive crypto-currency worldYou are eager to learn new knowledge in both financial and technical fieldsYou thrive in a non-hierarchical organization with a casual working environmentYou enjoy solving complex distributed systems challenges and optimizing streaming data pipelinesYou value comprehensive documentation and collaborative problem-solving Team / Role As an Engineer you will: SRE & Observability Own the observability of the platform end to end: decide what to measure, instrument it, collect it, store it, and put it in front of peopleBuild Grafana dashboards and alerts that name what broke and what to check next — not alerts that merely report a number crossing a lineRun the production systems: diagnose silent failures, recover missed data, and fix the cause so it doesn’t recurShip changes through CI/CD and GitOps, and know the difference between “merged” and “live”Operate streaming workloads on Kubernetes (EKS) — resource budgeting, config delivery, capacity, and why an OOM-kill can also be a data-loss eventProvision the AWS resources behind the platform as code (we use Terraform; we’ll show you ours)Participate in on-call for the data and observability platform Data Pipeline Engineering Build and maintain the near-realtime pipelines that carry telemetry and logs from source to quarriable tableReason about delivery guarantees (at-most-once vs at-least-once) and what they mean for the numbers on a dashboardDesign data quality checks — route bad records for investigation rather than dropping them silentlyDesign and optimize OLAP tables so billion-row queries return fast: sort keys, partitioning, deduplication, retentionWrite and tune the SQL behind dashboards and analyticsTranslate a request like “show me p99 latency per venue” into a pipeline change and a table design Required Skillset University degree in Computer Science, Software Engineering or related disciplinesObservability and monitoring in production — instrumenting services, building dashboards, writing alerts that are worth acting on, and diagnosing a live incident from metrics and logs together (Prometheus, Grafana, Loki, ELK, Datadog, CloudWatch — any equivalent stack)CI/CD — delivering your own changes through a pipeline (GitLab CI, Jenkins, GitHub Actions, or similar), not handing them to someone elseA distributed stream processing engine in production (Flink, Spark Structured Streaming, or similar)An OLAP / columnar database — designing tables and tuning queries, not just reading from them (ClickHouse, BigQuery, Druid, Redshift, or similar)Kubernetes working knowledge — deployments, resource requests and limits, config delivery, reading pod stateStrong programming in at least one of Python, Scala, Java or Rust, plus fluent analytical SQLAbility to troubleshoot production systems methodically, and the instinct to fix the cause rather than the symptom It is a plus if: OpenTelemetry, distributed tracing, or high-cardinality telemetry at scaleAirflow (DAGs, scheduling, task orchestration)Infrastructure as code beyond the basics (Terraform modules, drift, targeted applies)Build tools for JVM or multi-module projects (Maven, Gradle, sbt, Cargo)Familiarity with message queues (SQS, Kafka, or similar) and event-driven architec We encourage applicants to read through our Privacy Notice for Applicants before submitting your applications - https://pulsar.com/careers/