Site Reliability Engineer
Salient Group — Australia · Posted ~20 hours ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
Observability & Reliability Engineer | Melbourne
💰 $150k–$170k + Super + Bonus📍
Melbourne | Hybrid
When you’re building a platform that sits behind high-stakes decisions, knowing what to work on and when is important.
As the platform grows, so has the complexity underneath it - more services, more dependencies, more signals and, inevitably, more noise.
It’s easy to spend too much time working out which alerts actually matter.
That's the problem you're coming in to solve.
This isn’t just configuring another alert.
It’s designing what should be monitored in the first place, understanding what happens as systems scale, dealing with alert fatigue, tracing problems across distributed services and learning which signals genuinely predict something going wrong.
If you've built or significantly evolved an observability system before, you'll probably have some scars from doing it.
That's exactly what they're looking for.
So what could a week actually look like?
Monday, you're looking at the existing environment and working out why engineers are getting so much noise - what can disappear, what needs changing and what's currently missing altogether.
Tuesday, you're tracing a request across a distributed AWS environment, finding a degradation that previously would have been difficult to spot until it became an incident.
Wednesday, you're working directly with software engineers on instrumentation, getting deeper visibility into what's happening inside the application rather than simply watching the infrastructure around it.
Thursday, you're looking at where static thresholds stop being useful and where anomaly detection or automation could identify and resolve problems earlier.
Friday, something you've changed means an engineer doesn't get pulled away from their work by another meaningless alert.
Over time, that compounds.
More engineering time spent actually building the product.And that's what I think makes the role interesting from a career perspective.
You're not joining a huge reliability organisation and inheriting one small component of somebody else's system.
You get to shape the capability.
You'll take ownership of questions like:
What are the signals that actually tell us something is wrong?How do we identify degradation before it becomes an incident?How do we trace what's happening across distributed services?How do we stop one underlying problem generating a wall of noise?What should wake an engineer up, and what shouldn't?What needs to change as the platform and traffic scale?What can we automate away entirely?
You'll work across AWS, distributed and event-driven systems, OpenTelemetry, metrics, logging, tracing, anomaly detection and automation, with enough ownership to make decisions about how it all fits together.
And if you do it well, you'll be able to point to a production platform and say: “I built how we know this thing is healthy.”
You might currently be called an SRE, Reliability Engineer, Observability Engineer, Platform Engineer or DevOps Engineer.
I'm much more interested in what you've built than what you're called.
If you've tackled this problem before and want more ownership of it next time around, apply here or drop me Jon Holland a message on Linkedin.
We have 92,799 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 92,799 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume