Lead Observability Platform Engineer

Coxpurtell — Australia · Posted ~2 hours ago

Lead Full-time Onsite

Skills

observability platform engineering high-throughput systems low-latency systems reliability engineering telemetry

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Take a lead role in building and operating observability capabilities for high-stakes, high-throughput, low-latency technology systems. You will work in an environment where reliability, precision and performance are critical, with opportunities to apply experience from demanding technical sectors.

Highlights

Work on high-frequency, high-throughput and low-latency systems in a well-resourced environment, with attractive compensation, bonuses and rewards.

Description

Observability Platform Engineer Sydney | Full-time Seek Selling Points You already know when the telemetry is lying before the dashboard ever admits itAttractive salary, real bonus, and a rewards package built for someone this goodHigh-frequency, low-latency, mission-critical work, built for genuinely serious minds About the Client You'll be joining an organisation operating at the sharp end of technology: high-frequency, high-throughput, low-latency systems where speed, precision and reliability aren't aspirations, they're the job. Here, downtime or delay has a real, measurable cost, not just reputational noise. It's the kind of environment that suits people who've worked across high-stakes sectors, from finance, trading and gaming through to telecommunications and mission-critical government and defence technology, and who already know the difference between a system that's fast and one that's actually trustworthy. Sydney based, corporate, well-resourced, and built for people who want the hardest version of this job, not the easiest. About the Role You're the one who gets pulled into the incident channel first, not last. You've already built or fixed the observability platform other engineers rely on without thinking about it, and you know that's the whole point: it should be invisible until the moment it isn't. This role puts you in charge of that platform end to end, metrics, logs, traces, events, alerting and diagnostics, the layer that tells everyone else what's actually happening under load. You already treat observability as a trust problem, not a tooling problem. That instinct is exactly why this role exists. Duties You'll design, build and operate a shared observability platform spanning telemetry collection, ingestion, storage, query, visualisation, alerting and diagnostics.You'll build the services, APIs, integrations and dashboards that make observability easier to adopt and more reliable to operate.You'll improve the scalability, reliability, performance and cost-effectiveness of high-volume telemetry systems.You'll improve developer and operator experience through self-service workflows, golden paths and practical platform abstractions.You'll own the reliability of what you build, including failure modes, monitoring, incident learnings and continuous improvement. Requirements You've got strong engineering experience in SRE, platform engineering, infrastructure, observability, developer tooling or distributed systems.You already reason about failure modes, debugging workflows and service reliability under pressure, because you've had to.You have technical depth across logs, metrics, traces, events, alerting, dashboards and telemetry pipelines, not just familiarity.You've designed, built or operated services, pipelines or tools that other engineering teams depend on.You're comfortable with modern observability and telemetry tooling across metrics, logging, tracing and time-series systems, and you keep up with where it's heading. How to Apply If most of this reads like a description of how you already work, not what you're hoping to grow into, we should talk. Desired Skills and Experience Observability Platform Engineer, Kafka, Grafana, ELK, OpenSearch, ClickHouse, VictoriaMetrics, InfluxDB, Telegraf, Vector, OpenTelemetry, Prometheus