Description
Job Description
What is the opportunity?
Join our Credit Technology team as a Lead Site Reliability Engineer, where you'll play a key role in drive operational excellence through technology, process optimization, and cross-functional collaboration for the Personal Credit SRE & Ops team.
This exciting opportunity will challenge you to work with cutting-edge technologies, including AI and emerging innovations, and collaborate closely with development teams to deliver embedded SRE solutions.
As a vital link between QE, DevOps, Development, Infrastructure, and Support teams, you'll leverage your strong technical skills to solve complex problems and drive success across multiple components and technologies.
If you're passionate about tackling new challenges and developing innovative solutions, we invite you to join our team and take your career to the next level.
What will you do?
Act as one of the final escalation points for critical outages and lead 24/7 incident response with rapid resolution of customer-impacting issuesLead strategic direction and continuous improvement initiatives across Credit TechnologyManage cross-functional teams and stakeholders to execute upgrade and operational change managementOversee end-to-end reliability of the ecosystem (hardware, software, network) ensuring 99.9% availabilityGenerate performance metrics (SLOs/SLAs/SLIs) and maintain regulatory compliance and security standardsServe as primary relationship owner for vendor services, maintenance, and internal technology teamsIdentify, design, write and test automation procedures using AI, Ansible and other relevant technologiesSupport applications running on many platforms including OpenShift and distributed systemsImplement Chaos Engineering experiments and Disaster Recovery procedures to test and validate system resilience and reliabilityImplement monitoring and alerting, anomaly detection and reliability testing for applications in scope
What do you need to succeed?
To be successful in this role, you will need:
Must have:
5-7 years of experience as a Site Reliability Engineer or a Cloud Developer Proven leadership experience managing cross-functional teams and stakeholdersGood understanding of Kubernetes and Cloud working knowledge with experience and understanding of CICD pipeline and DevOps / Agile MethodologyDecent knowledge of the following SRE practices and technologies: Python, YAML, Shell scripting, OpenShift, Linux, MongoDB, Dynatrace, Prometheus, PagerDuty, Moog, Splunk, Elastic, Ansible, Grafana, Chaos Engineering, MQ, KafkaPerform production support role, including off-hours supportExcellent communication skills
Nice-to-have
Experience in the context of SRE and/or Application Development, Test Automation teams
What’s in it for you?
We thrive on the challenge to be our best, progressive thinking to keep growing, and working together to deliver trusted advice to help our clients thrive and communities prosper.
We care about each other, reaching our potential, making a difference to our communities, and achieving success that is mutual.
A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicable.Leaders who support your development through coaching and managing opportunities.Ability to make a difference and lasting impactWork in a dynamic, collaborative, progressive, and high-performing teamA world-class training program in financial servicesFlexible work/life balance options.Opportunities to do challenging work.Opportunities to take on progressively greater accountabilities.
Opportunities to building close relationships with clients.
Job Skills
Agile Methodology, Agile Methodology, Apache Kafka, Application Development, CI/CD, Collaboration, Cross-Functional Teamwork, Decision Making, DevOps, Dynatrace APM, Elastic Stack (ELK), GitHub, Grafana, Group Problem Solving, Incident Communications, IT Systems Integration, Linux, Mainframe Support, Organizational Leadership, Problem Solving, Production Support, Product Services, Red Hat Ansible, Red Hat OpenShift, Scrum (Agile) {+ 7 more}
Additional Job Details
Address:
RBC WATERPARK PLACE, 88 QUEENS QUAY W:TORONTO
City:
Toronto
Country:
Canada
Work hours/week:
37.5
Employment Type:
Full time
Platform:
TECHNOLOGY AND OPERATIONS
Job Type:
Regular
Pay Type:
Salaried
Posted Date:
2026-08-06
Application Deadline:
2026-08-27
Note: Applications will be accepted until 11:59 PM on the day prior to the application deadline date above
Our Employment Opportunities
At RBC, we are guided by living shared values of Client First, Integrity, Collaboration, Respect and Excellence and winning together as One RBC.
We believe an inclusive workplace that has diverse perspectives is core to our continued growth as one of the largest and most successful banks in the world.
Maintaining a workplace where our employees feel supported to perform at their best, effectively collaborate, drive innovation, and grow professionally helps to bring our Purpose to life and create value for our clients and communities.
RBC strives to deliver this through policies and programs intended to foster a workplace based on respect, belonging and opportunity for all.
Join our Talent Community
Stay in-the-know about great career opportunities at RBC.
Sign up and get customized info on our latest jobs, career tips and Recruitment events that matter to you.
Expand your limits and create a new future together at RBC.
Find out how we use our passion and drive to enhance the well-being of our clients and communities at jobs.rbc.com
RBC is presently inviting candidates to apply for this existing vacancy.
Applying to this posting allows you to express your interest in this current career opportunity at RBC.
Qualified applicants may be contacted to review their resume in more detail.