Summary
✨ AI‑Generated
Take a lead role in building and operating observability capabilities for high-stakes, high-throughput, low-latency technology systems. You will work in an environment where reliability, precision and performance are critical, with opportunities to apply experience from demanding technical sectors.
Highlights
Work on high-frequency, high-throughput and low-latency systems in a well-resourced environment, with attractive compensation, bonuses and rewards.
Description
Observability Platform Engineer
Sydney | Full-time
Seek Selling Points
You already know when the telemetry is lying before the dashboard ever admits itAttractive salary, real bonus, and a rewards package built for someone this goodHigh-frequency, low-latency, mission-critical work, built for genuinely serious minds
About the Client
You'll be joining an organisation operating at the sharp end of technology: high-frequency, high-throughput, low-latency systems where speed, precision and reliability aren't aspirations, they're the job.
Here, downtime or delay has a real, measurable cost, not just reputational noise.
It's the kind of environment that suits people who've worked across high-stakes sectors, from finance, trading and gaming through to telecommunications and mission-critical government and defence technology, and who already know the difference between a system that's fast and one that's actually trustworthy.
Sydney based, corporate, well-resourced, and built for people who want the hardest version of this job, not the easiest.
About the Role
You're the one who gets pulled into the incident channel first, not last.
You've already built or fixed the observability platform other engineers rely on without thinking about it, and you know that's the whole point: it should be invisible until the moment it isn't.
This role puts you in charge of that platform end to end, metrics, logs, traces, events, alerting and diagnostics, the layer that tells everyone else what's actually happening under load.
You already treat observability as a trust problem, not a tooling problem.
That instinct is exactly why this role exists.
Duties
You'll design, build and operate a shared observability platform spanning telemetry collection, ingestion, storage, query, visualisation, alerting and diagnostics.You'll build the services, APIs, integrations and dashboards that make observability easier to adopt and more reliable to operate.You'll improve the scalability, reliability, performance and cost-effectiveness of high-volume telemetry systems.You'll improve developer and operator experience through self-service workflows, golden paths and practical platform abstractions.You'll own the reliability of what you build, including failure modes, monitoring, incident learnings and continuous improvement.
Requirements
You've got strong engineering experience in SRE, platform engineering, infrastructure, observability, developer tooling or distributed systems.You already reason about failure modes, debugging workflows and service reliability under pressure, because you've had to.You have technical depth across logs, metrics, traces, events, alerting, dashboards and telemetry pipelines, not just familiarity.You've designed, built or operated services, pipelines or tools that other engineering teams depend on.You're comfortable with modern observability and telemetry tooling across metrics, logging, tracing and time-series systems, and you keep up with where it's heading.
How to Apply
If most of this reads like a description of how you already work, not what you're hoping to grow into, we should talk.
Desired Skills and Experience
Observability Platform Engineer, Kafka, Grafana, ELK, OpenSearch, ClickHouse, VictoriaMetrics, InfluxDB, Telegraf, Vector, OpenTelemetry, Prometheus