Manager, Site Reliability Engineering

Akkodis — Canada · Posted ~4 hours ago

Lead Full-time Hybrid Visa History ✓ $140K-$165K per annum plus bonus and benefits

Skills

Site Reliability Engineering Cloud infrastructure Automation Observability Scalability Production operations Cloud SRE

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A leading organization is seeking an SRE leader to improve reliability, scalability, monitoring, and resilience across critical technology platforms while building engineering practices and culture.

Highlights

Leadership role shaping reliability engineering practices, automation strategies, and operational standards for high-availability technology platforms.

Description

Title: Manager, Site Reliability Engineering (SRE) Position Type: Full-time, Permanent Location: Etobicoke, ON Working Model: Hybrid - 3 days per week in office Compensation: $140K - 165K per annum plus bonus and benefits About Our Client Our client is a leading Canadian organization operating highly available, mission-critical technology platforms that support millions of transactions and customer interactions annually. As part of a broader technology transformation and operational excellence initiative, the organization is continuing to invest in its Site Reliability Engineering (SRE) capabilities to improve reliability, scalability, observability, and production resilience across a large portfolio of business-critical applications. This is an exciting opportunity to join a growing SRE function at an important stage of its evolution, helping shape reliability practices, operational standards, automation strategies, and engineering culture. About the Opportunity We are seeking a hands-on Manager, Site Reliability Engineering (SRE) to lead a team of SRE engineers responsible for improving the reliability, availability, performance, and operational excellence of critical enterprise applications and platforms. This role offers a balance of people leadership and technical leadership, making it ideal for an experienced SRE, Platform Engineering, DevOps, or Production Engineering leader who enjoys building high-performing teams while remaining engaged in technical strategy and operational improvements. You will partner closely with Application Development, DevOps, Infrastructure, Security, and Incident Management teams to drive observability, automation, incident response, resiliency, and continuous reliability improvements across both cloud and on-premises environments. What You'll Do Lead and develop a team of Site Reliability EngineersDrive adoption of SRE principles, including SLOs, SLIs, error budgets, and reliability best practicesImprove observability, monitoring, alerting, and operational health across critical servicesReduce operational toil through automation and self-healing capabilitiesOversee production support, incident management, escalation processes, and on-call operationsPartner with engineering teams to embed reliability into application design and delivery pipelinesLead capacity planning, resiliency testing, and disaster recovery readiness initiativesAct as a senior escalation point during major incidentsImprove deployment reliability and reduce operational riskEstablish runbook standards, documentation practices, and operational governanceCollaborate with leadership teams to execute strategic SRE initiatives and roadmapsFoster a blameless culture focused on continuous improvement and operational excellence What You Bring Must-Have Qualifications 8+ years of experience supporting large-scale production environments, distributed systems, or cloud platforms3+ years of people leadership experience managing technical teamsExperience leading SRE, Platform Engineering, DevOps, Production Operations, or Reliability Engineering teamsStrong understanding of Site Reliability Engineering principles and practicesExperience with observability and monitoring platforms such as Dynatrace, Datadog, New Relic, AppDynamics, or similarHands-on experience with Azure (preferred) or AWSStrong Kubernetes and container platform experienceExpertise with Infrastructure as Code and automation tools such as Terraform and AnsibleStrong Linux administration experienceExperience leading incident response, problem management, and production operationsStrong scripting or programming skills (Python, PowerShell, Bash, etc.)Excellent communication, stakeholder management, and leadership skills Nice-to-Have Qualifications Experience in payments, fintech, banking, or highly regulated environmentsKnowledge of PCI-DSS, NIST, or related compliance frameworksExperience with change management and release governance processesExperience implementing SLOs, SLIs, and error budget frameworksBackground building or maturing an SRE functionExperience supporting enterprise-scale cloud transformation initiatives What You'll Love About This Opportunity Opportunity to help shape and mature a growing SRE organizationLead a team responsible for business-critical services and platformsDrive meaningful improvements in reliability, observability, and automationBalance people leadership with hands-on technical influenceWork alongside highly skilled Engineering, Infrastructure, Security, and DevOps teamsHigh visibility role with broad organizational impactCollaborative culture focused on innovation, resilience, and continuous improvementCompetitive compensation, bonus, and benefits packageStrong career growth opportunities within a modern technology organization How to Apply If you are interested in learning more, don't hesitate to apply today or check out Akkodis Canada website for more opportunities. Important We thank all applicants for their interest in this opportunity. Only candidates meeting the above qualifications will be contacted for further discussions. Akkodis Canada will never share your resume or any personal details without your explicit consent. We thank all applicants for their interest in this opportunity. Only candidates meeting the above qualifications will be contacted for further discussions. Our Commitment: At Akkodis, part of The Adecco Group, our purpose is simple: to make the future work for everyone. We live our values, Passion, Collaboration, Inclusion, Courage, and Customers at Heart, by fostering a workplace where diversity is celebrated and every voice matters. We encourage applications from individuals of all backgrounds and identities. Together, we’re making the future work for everyone.