Senior Site Reliability Engineer - Data Platforms

Asb Bank — New Zealand · Posted ~6 days ago

Senior Full-time

Skills

SRE cloud infrastructure automation observability platform operations

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior reliability engineering role focused on automating and improving enterprise data platform operations. The position combines infrastructure engineering, monitoring, and continuous improvement.

Highlights

Solve large-scale infrastructure challenges through automation, reliability engineering, and modern operational practices.

Description

Build the automation that keeps data platforms running Keeping enterprise-scale data platforms secure, stable and available takes more than operational support. It requires smart automation, modern engineering practices and people who can solve problems at scale. We're looking for a Site Reliability Engineer (Data & Analytics Platform) to help modernise and automate the infrastructure and operational processes that support some of ASB's most important data and analytics platforms. You'll play a key role in transforming how patching, upgrades and platform operations are delivered, building automation that reduces manual effort, improves reliability and helps the team meet increasingly ambitious operational timeframes. This is an opportunity to work on complex infrastructure challenges at enterprise scale, combining hands-on engineering with automation, observability and continuous improvement. If you enjoy solving operational problems, improving systems and seeing the impact of your work in production environments, you'll find plenty to get stuck into here. About The Team You'll join ASB's Data & Analytics Platform team, responsible for supporting and evolving the platforms that enable data ingestion, reporting, analytics and business insight across the bank. The team supports a broad range of technologies and services, including data ingestion platforms, reporting environments and analytics tooling. These platforms underpin critical business decision-making and require high levels of reliability, security and operational excellence. As expectations continue to grow and operational windows become increasingly compressed, the team is investing heavily in automation to improve scalability, resilience and efficiency. You'll work alongside experienced Site Reliability Engineers, platform specialists and data professionals who value curiosity, ownership, practical problem-solving and continuous improvement. About The Role This role exists to help modernise and automate the management of a large-scale platform environment that supports critical data and analytics services. A key focus will be building and implementing Ansible-driven automation to improve how infrastructure patching, upgrades and operational tasks are executed across both Windows and Linux environments. The team already has strong foundations in place, and this role will help accelerate the move from manual processes towards highly automated, repeatable and reliable operating models. You'll work across server management, patching automation, observability, CI/CD pipelines and operational tooling, helping to improve platform reliability while reducing operational effort and risk. You'll also contribute to the ongoing evolution of cloud-based services and help build more sustainable ways of operating as the team's platforms continue to evolve. Key Responsibilities Design, build and maintain automation solutions supporting platform operations and infrastructure managementDevelop and implement Ansible automation across Windows and Linux environmentsAutomate patching, upgrades and operational workflows to improve reliability and efficiencySupport the ongoing management, maintenance and optimisation of critical data and analytics platformsImprove operational processes by reducing manual intervention through automationBuild and maintain CI/CD pipelines and deployment workflows using GitHubInvestigate, troubleshoot and resolve infrastructure and operational issuesDevelop monitoring, reporting and observability capabilities that improve operational visibilityPartner with platform specialists and engineers to identify opportunities for continuous improvementSupport both on-premise and cloud-based infrastructure environmentsContribute to modern engineering practices that improve resilience, scalability and service reliability About You You're an engineer who enjoys solving infrastructure and operational problems through automation. You understand that maintaining reliable, secure and scalable platforms requires more than simply keeping systems running. You're motivated by identifying inefficiencies, improving processes and building automation that makes a meaningful difference. You'll likely come from a Site Reliability Engineering, Platform Engineering, Systems Engineering, Infrastructure Engineering or Automation Engineering background and will be comfortable working across both Windows and Linux environments. You're curious, hands-on and pragmatic, with strong troubleshooting skills and a passion for continuous improvement. You'll enjoy working in a collaborative environment where you'll have the opportunity to introduce new ideas, improve existing processes and help shape the future of automation within a critical technology domain. Essential Key Skills & Experience Strong experience with Ansible and infrastructure automationExperience managing and supporting both Windows and Linux server environmentsExperience building and supporting automation solutions within enterprise environmentsExperience with GitHub and CI/CD pipeline developmentExperience with infrastructure operations, Site Reliability Engineering, Platform Engineering or Automation EngineeringExperience with Infrastructure as Code (IaC) principles and practicesExperience working within Agile and DevOps environmentsStrong troubleshooting and problem-solving skillsExcellent communication, collaboration and documentation skills Desirable PowerShell scripting and automation experiencePython development or scripting experienceAzure experienceTerraform experienceExperience with observability and monitoring platformsExperience supporting large-scale enterprise infrastructure environmentsKnowledge of data and analytics platforms or servicesExperience improving patching, compliance or operational management processesExperience using AI-assisted development or problem-solving tools to improve efficiency and outcomes Working for ASB At ASB, we're committed to helping our people thrive both professionally and personally. We offer: Flexible and hybrid ways of workingOngoing learning and career development opportunitiesAccess to banking, insurance and financial wellbeing benefitsHealth, wellbeing and employee assistance programmesGenerous parental leave and family support initiativesAn inclusive and supportive culture where different perspectives are valuedThe opportunity to work with modern technologies and solve complex problems at enterprise scale Make an impact where it matters This is an opportunity to help shape the future of automation and reliability within one of New Zealand's largest technology environments. You'll work on meaningful engineering challenges, help modernise critical platform operations and improve the reliability of the services that power data and analytics across the bank. If you're passionate about automation, infrastructure engineering and building smarter ways of operating at scale, we'd love to hear from you. Apply now Ready to help modernise how critical platform services are managed and automated? Apply today and join a team where your ideas, expertise and impact will be recognised.