Senior Site Reliability Engineer
Bell Financial Group Au — Australia · Posted ~2 hours ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
We are looking for a Senior Site Reliability Engineer to join our Engineering team in Melbourne.
Reporting to the Head of Engineering you will work across Bell Financial Group to build and operate the systems that keep our critical platforms available, reliable and efficient.
The role owns our delivery engineering toolchain and CI/CD platform, which we treat as production infrastructure in its own right.
We operate a hybrid environment, with cloud workloads on ECS Fargate and AWS native datastores alongside an on-premise footprint.
You will be comfortable engineering reliability across both environments and managing the dependencies between them.
This is a hands-on role working closely with the product engineering teams, with the opportunity to set reliability standards and coach developing engineers.
Key Responsibilities
Implement and maintain SLIs, SLOs and error budgets for critical services.Build resilience into production systems through redundancy, graceful degradation and safe recovery practices.Develop observability across metrics, logs, traces and synthetic monitoring.Participate in the on-call rotation and major incident response.Lead blameless post-incident reviews and ensure remediation actions are completed.Design and run controlled chaos engineering and disaster recovery exercises.Validate recovery time and recovery point objectives against actual system behaviour.Own and enhance CI/CD platforms, deployment pipelines and supporting build infrastructure.Implement safe deployment practices, including blue/green and canary releases, automated rollback and deployment gating.Develop reusable infrastructure-as-code modules and pipeline templates.Design highly available solutions across AWS and on-premise environments.Identify and eliminate operational toil through automation.Support and coach other engineers in reliability, observability and operational practices.
Skills & Experience
Bachelors degree in Information Technology, Computer Science or similar is desirable.AWS Professional-level certifications are also desirable.More than six years’ hands-on experience in software, infrastructure or site reliability engineering.Strong AWS experience, including ECS Fargate, Lambda, S3, DynamoDB, Aurora, CloudFront and IAM.Experience operating hybrid cloud and on-premise environments.Strong Linux administration, networking and troubleshooting skills.Experience owning or supporting CI/CD platforms such as GitHub Actions, GitLab CI, Jenkins, Azure DevOps or Buildkite.Hands-on infrastructure-as-code experience using AWS CDK, Terraform and/or CloudFormation.Proficiency in TypeScript, with experience in Python or Go for automation and tooling.Experience with observability tools such as New Relic, Datadog, Grafana, Prometheus, OpenTelemetry, Splunk or AWS-native tooling.Demonstrated experience with SLIs, SLOs, error budgets, disaster recovery testing and incident response.Strong written and verbal communication skills, with the ability to communicate effectively with technical and non-technical stakeholders.
About You
You are a pragmatic, hands-on engineer who enjoys solving complex reliability problems across cloud and on-premise environments.You take ownership of critical systems, remain calm during major incidents and look for opportunities to replace repetitive operational work with automation.You are collaborative and improvement-focused, with the ability to balance delivery speed, reliability, security and cost.You also enjoy sharing your knowledge and helping other engineers strengthen their technical and operational capabilities.Experience in financial services, AWS migrations, platform engineering, software supply-chain security or APRA-regulated environments will be highly regarded.
Why Join Us
Work in a modern AWS and hybrid technology environment.Take ownership of critical reliability and delivery engineering capabilities.Help shape reliability standards and engineering practices across the organization.Work in a stable and professional financial services environment.Join a supportive and collaborative technology team.
Applications for this role such be submitted using the “Apply” button.
Please note that only successful candidates will be notified.
We have 141,913 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 141,913 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume