Lead Site Reliability Engineer

Mcmillan Shakespeare — Australia · Posted ~1 day ago

Lead Full-time

Skills

Site Reliability Engineering observability Azure cloud infrastructure hybrid environments automation platform engineering technical leadership engineering standards mentoring cloud hybrid cloud AI-assisted operations

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary

An established organization is seeking a hands-on technical leader to build an SRE capability from the ground up. You will design an enterprise observability platform across cloud and hybrid environments, automate operations, establish engineering standards, mentor a team, and influence technology strategy.

Highlights

High-autonomy technical leadership opportunity to establish an SRE practice from the ground up, shape enterprise observability strategy, mentor engineers, and influence technology direction.

Description

This is a great opportunity to establish and lead Site Reliability Engineering at Macmillan Shakespeare. Here’s how you will make a difference in this role… You'll start as a hands-on technical leader, designing and building our enterprise observability platform across Azure, cloud and hybrid environments. As the capability grows, you'll transition into leading and building our SRE practice, setting engineering standards, mentoring a team and influencing technology strategy across the business. If you're excited by creating platforms, automating operations, improving reliability and using modern technologies—including the adoption of AI-assisted operations—this role offers the autonomy and support to make a lasting impact. At MMS clear expectations help our people Be the Difference. Your key responsibilities in this role will include: Build our SRE capability from the ground up.Shape the enterprise observability platform strategy and technology roadmapInfluence architecture and engineering practices across the organisationWork with modern Azure cloud technologies, automation and AI-enabled operationsEnsure Platform Robustness, Compliance & Cyber SecurityGrow into a strategic leadership role while staying close to technologyPartnering with Product Owners, Engineering Managers and Platform teams to embed reliability engineering practices throughout the software development lifecycle.Join a collaborative team that values innovation, continuous improvement and engineering excellence To be considered for this role you will have: Demonstrated experience designing and operating enterprise monitoring, observability and operational platforms in complex hybrid/cloud environments.Experience with Infrastructure as Code (IaC), CI/CD pipelines, platform automation, and Microsoft Azure integrations not just AWS, including logging, monitoring and alerting.Experience implementing monitoring, logging and distributed tracing using modern observability practices.Experience designing automated operational workflows using APIs and event-driven architectures.Strong understanding of Site Reliability Engineering (SRE), operational automation and continuous improvement.Experience leading major incident management, root cause analysis and post-incident reviews.Ability to simplify complex technical concepts and influence engineering teams through technical leadership.Strong stakeholder engagement skills across engineering, product, operations, cyber security and executive leadership.Experience developing engineering standards, governance frameworks and operational practices. Desirable Experience establishing or leading SRE, Platform Engineering or DevOps functions.Experience with enterprise IT Service Management (ITSM) platforms and automated incident management.Experience integrating Jira, Microsoft Teams and collaboration platforms into operational workflows.Experience with Azure AI Foundry, Azure Open AI or AI-assisted operational toolingExperience in regulated industries with strong governance and compliance requirements.Knowledge of modern cloud architecture, resilience engineering and distributed systems. Essential Bachelor's degree in Information Technology, Computer Science, Software Engineering or a related discipline, or equivalent industry experience. Desirable But Not Essential ITIL® 4 Foundation or Managing Professional certification.Relevant Microsoft Cloud certifications.Certified Site Reliability Engineer (SRE) qualification or equivalent industry training.Agile and/or Scrum certifications. What we can offer you: Novated leasing benefits and discounts12 weeks paid parental leave and access to our Parents PortalComprehensive learning and development opportunities to support your career growthSonder digital wellbeing platform, providing personalised support 24/7, plus annual flu vaccinationsDefault Income Protection Insurance reimbursed for members of the MMS Default Super FundExempt Employee Share PlanEmbracing our value of Everyone Matters we hold a collective commitment to foster an environment where all differences are valued and respected.We encourage individuals from all backgrounds including Aboriginal and Torres Strait Islander peoples, those caring for someone or living with a disability, LGBTQIA+ and culturally diverse applicants to apply.We embrace hybrid working and welcome conversations about flexibility.We value the skills and attributes veterans can bring to our organisation. Please note all successful candidates will undergo background checks (including criminal history and ASIC checks) and an NDIS Workers Screening Check if appropriate. All information provided will be treated confidentially. We acknowledge Aboriginal and Torres Strait Islander peoples as the Traditional Custodians of the lands where we live, learn and work. If you identify as a person living with disability and require adjustments to our recruitment process, please contact us at mmsgrouprecruitment@mmsg.com.au