Summary
✨ AI‑Generated
A newly created SRE management role leading a small team responsible for the reliability, availability, performance, and resilience of a large portfolio of business-critical applications. The position combines team development with hands-on technical leadership across cloud infrastructure, automation, observability, and production operations.
Highlights
Competitive salary of $155,000–$165,000 and an opportunity to build and mature an SRE function. The role combines people leadership with hands-on technical work across cloud, automation, observability, infrastructure, and production reliability.
Description
Position: Manager, Site Reliability Engineering (SRE)
Location: Toronto
Salary: $155,000 - $165,000
Posting Type: Open vacancy
Our client is looking for an experienced Manager, Site Reliability Engineering (SRE) to lead a team responsible for the reliability, availability, performance, and resiliency of business-critical applications and platforms.
This is a newly created position reporting to the Director, SRE & DevOps.
The successful candidate will combine people leadership with strong hands-on technical expertise, helping mature an evolving SRE function while supporting a large portfolio of critical applications.
Our client is looking for someone who can lead and develop an SRE team while remaining technically credible and hands-on across cloud, automation, observability, infrastructure and production reliability.
The SRE portfolio may ultimately support approximately 100 critical applications, with the Manager expected to lead a team of approximately 5-8 SRE Engineers and manage future backfills as required.
Description:
- Lead, coach and develop a team of Site Reliability Engineers.
- Own the reliability, availability, scalability and performance of assigned business-critical applications and platforms.
- Help mature SRE practices including SLIs, SLOs, error budgets, incident response, post-incident reviews and continuous reliability improvement.
- Oversee production operations, incident management, escalation handling and problem management.
- Drive improvements in observability, monitoring and operational visibility.
- Partner with Development, DevOps, Infrastructure, Security and Incident Management teams.
- Reduce operational toil through automation and self-healing capabilities.
- Lead capacity planning, resiliency testing and disaster recovery readiness.
- Act as a senior technical escalation point during major production incidents.
- Establish effective runbooks, documentation standards and knowledge-management practices.
- Help develop reliability roadmaps for the applications within your portfolio.
- Support hiring and workforce planning as the SRE organization continues to evolve.
Skills:
- 8+ years of experience in senior technical roles supporting complex or distributed systems.
- 3+ years of people leadership experience.
- Strong hands-on knowledge of at least one major cloud platform, with Azure preferred; AWS is also acceptable.
- Strong knowledge of Kubernetes.
- Near-expert capability with Infrastructure as Code, automation and observability.
- Strong experience with enterprise observability tools; Dynatrace is preferred, with Datadog, New Relic, AppDynamics or similar also relevant.
- Linux knowledge - this is an important technical requirement.
- Strong scripting and/or programming capabilities.
- Strong understanding of SRE principles including SLOs, SLIs, incident management, resiliency and toil reduction.
- Ability to operate as both a people leader and hands-on technical leader.
- Excellent communication and stakeholder-management skills.
Nice to Have:
- Experience within payments, fintech or other highly regulated environments.
- Exposure to PCI DSS and/or NIST frameworks.
- Experience working with change-management and compliance processes.
- Familiarity with SDLC best practices.
Please send your resume to cristina.arsene@quantum-qtr.com.
REFER A NEW HIRE AND EARN A CASH BONUS! For details, click here.
All applications are reviewed by our recruitment team, and hiring decisions are made by people.
We may also use AI-enabled tools to support parts of the application review process.