Site Reliability Engineering Manager
Akkodis — Canada · Posted ~5 hours ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
Manager, Site Reliability Engineering (SRE)
Toronto, Ontario (Hybrid)
Permanent Opportunity
Location: Toronto, Ontario
Work Model: Hybrid
About the Opportunity
Akkodis is seeking an experienced Manager, Site Reliability Engineering (SRE) to lead a team responsible for the reliability, availability, performance, and scalability of enterprise applications and platforms.
The successful candidate will drive Site Reliability Engineering best practices, observability initiatives, automation, incident management, and continuous improvement across cloud and on-premises environments.
This opportunity is ideal for a technical leader with strong experience in cloud operations, production support, automation, and people leadership.
What Will the Successful Candidate Do?
The Manager, Site Reliability Engineering (SRE)'s primary responsibilities include, but are not limited to:
Lead, mentor, and develop a team of Site Reliability Engineers.Drive reliability, availability, performance, and scalability across critical applications and platforms.Implement and operationalize SRE practices, including SLIs, SLOs, error budgets, and post-incident reviews.Oversee production support operations, on-call processes, incident management, and escalations.Partner with Development, Infrastructure, Security, and DevOps teams to improve service reliability.Drive observability and monitoring initiatives across the organization.Reduce operational effort through automation and self-healing solutions.Lead major incident response activities and support problem management processes.Support capacity planning, resiliency testing, and disaster recovery initiatives.Develop operational standards, runbooks, and knowledge management practices.Recruit, hire, onboard, and develop SRE talent.Promote a culture of continuous improvement, collaboration, and operational excellence.
What the Successful Candidate Needs to Succeed
Must-Have Skills
Site Reliability Engineering (SRE)Azure CloudKubernetesInfrastructure as Code (IaC)LinuxAutomation & ScriptingObservability & Monitoring ToolsIncident ManagementProduction Support OperationsRequired Experience
8+ years supporting enterprise applications and distributed systems.3+ years of experience leading technical teams.Experience implementing Site Reliability Engineering practices.Experience supporting mission-critical production environments.Experience with incident response and problem management.Experience driving automation and operational improvements.Strong stakeholder management and collaboration skills.Required Qualifications
Bachelor’s Degree in Computer Science, Software Engineering, or equivalent experience.Strong understanding of SRE principles, SLIs, SLOs, and error budgets.Hands-on experience with Azure and Kubernetes.Experience with Infrastructure as Code and automation frameworks.Experience with observability platforms such as Dynatrace, Datadog, New Relic, or AppDynamics.Strong scripting or programming skills.Excellent communication and leadership abilities.Preferred Qualifications
Experience in fintech, payment processing, or regulated environments.Experience with change management and compliance practices.Knowledge of SDLC and DevOps best practices.Experience with disaster recovery and resiliency testing.Technical Skills
Site Reliability Engineering (SRE) - Must HaveAzure Cloud PlatformKubernetesInfrastructure as Code (Terraform, ARM, Bicep, etc.)LinuxDynatrace, Datadog, New Relic, or AppDynamicsAutomation & ScriptingIncident ManagementDevOps PracticesSoft Skills
Strong leadership and coaching abilitiesExcellent communication skillsStrong problem-solving capabilitiesAbility to work effectively with cross-functional teamsContinuous improvement mindsetAdditional Requirements
Ability to work in a hybrid environment in Toronto.Experience managing production support and operational teams.Strong stakeholder-facing experience.Ability to lead through major incidents and operational challenges.
How to Apply
Submit your resume in confidence or apply through the Akkodis Canada website.
We thank all applicants for their interest in this opportunity.
Only candidates who meet the qualifications outlined above will be contacted for further discussions.
Accessibility
At Akkodis, part of The Adecco Group, our purpose is simple: to make the future work for everyone.
We foster a workplace where diversity is celebrated and every voice matters.
We encourage applications from individuals of all backgrounds and identities.
#SRE #SiteReliabilityEngineering #Azure #Kubernetes #CloudEngineering #EngineeringManager #DevOps #PlatformEngineering #Observability #Automation #InfrastructureAsCode #TorontoJobs #HybridJobs #TechnologyJobs #HiringNow #Akkodis #CloudOperations #ProductionSupport #LeadershipJobs
We have 144,028 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 144,028 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume