Staff, Site Reliability Engineer(Global Security)

Rbc β€” Canada Β· Posted ~5 hours ago

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Description

Job Description What is the opportunity? We are seeking an experienced and hands-on Staff Site Reliability Engineer (SRE) who is passionate about building scalable systems, automating infrastructure, and improving the reliability of production environments. You will work closely with system support, engineering and infrastructure teams to ensure the availability, performance, and efficiency of our IAM systems and services. This is a technical, execution-focused role with deep engineering work. What will you do? Serve as the senior-most technical voice for IAM reliability β€” setting architecture direction and reliability standards, and leading by example through hands-on design, coding, and operationsOwn service reliability for IAM systems end-to-end by defining and maintaining SLOs, SLIs, and error budgets, and using them to drive prioritization and continuous improvementDesign and implement resilient, highly available IAM infrastructure and services β€” spanning authentication, authorization, identity lifecycle, and privileged access β€” across multi-region and hybrid-cloud architecturesWrite and review production-grade code (services, APIs, automation frameworks, and internal tools), applying software engineering rigor β€” testing, code review, version control, modular design β€” to reliability work rather than treating it as throwaway scriptingBuild self-service platforms, reusable modules, and golden-path automation that let application teams provision, deploy, and safely operate IAM services with less hand-holding from the production support teamChampion Infrastructure as Code and GitOps practices (Terraform, Ansible, Puppet, Kubernetes/Helm) to eliminate manual toil, enforce consistency, and enable safe, repeatable deployments at scaleOwn and continuously improve CI/CD pipelines and release engineering for IAM services, embedding reliability, security, and rollback safety directly into the delivery pipelineLead incident response and on-call operations for high-severity IAM outages and performance degradations, driving root cause analysis, blameless postmortems, and long-term structural remediationBuild and evolve observability, monitoring, and alerting pipelines (metrics, logs, traces) using modern tooling to proactively detect, investigate, and resolve availability and security issues before they impact usersDevelop and test failover strategies and recovery procedures, including chaos engineering exercises, backup validation, and disaster recovery simulations to validate IAM system readinessOrchestrate workload automation, scheduling, and release pipelines across enterprise systems (e.g., Stonebranch, CI/CD platforms) to streamline delivery and reduce manual operational overheadPartner with security, infrastructure, application, and compliance teams to embed IAM into enterprise-wide business continuity and resilience strategy, ensuring alignment with risk and regulatory mandatesMentor and coach engineers on reliability, automation-first thinking, and sound software engineering practice; help raise the bar for how the broader team builds and operates IAM services What do you need to succeed? Must Have: 5+ years of experience in Site Reliability Engineering, DevOps or Platform Engineering with demonstrated staff/senior-level technical leadership across SLOs/SLIs, error budgets, and driving continuous reliability improvement at scaleSolid software engineering fundamentals β€” proficient in at least one modern language (Python, Go, Java, or similar), with the ability to design and build production-quality services and tooling, not just automation scriptsProven experience designing, implementing, and operating highly available, fault-tolerant, and scalable systems in production, including hybrid and multi-cloud environmentsStrong DevOps foundation β€” experience owning CI/CD pipelines (e.g., Jenkins, GitLab CI, GitHub Actions) and release engineering practices that build reliability into the delivery process itselfPlatform-engineering mindset β€” a track record of turning recurring operational work into self-service tooling, reusable modules, or golden-path automation that other engineering teams can adopt independentlyExperience with containerization and orchestration (Docker, Kubernetes) in production environmentsDeep knowledge of building and operating monitoring, alerting, and observability platforms (e.g., Prometheus, Grafana, Dynatrace, ELK, Splunk, SIEM) to enable proactive incident detection and responseProven incident management skills β€” able to lead high-severity incident response, root cause analysis, and postmortem processes, and to drive long-term fixesProficient in disaster recovery, failover strategies, and resilience testing (e.g., chaos engineering, tabletop exercises) to validate system readinessSolid understanding of cloud platforms (AWS, Azure) and hybrid environments, including experience supporting production workloads at scaleExcellent collaboration and communication skills β€” comfortable working across security, infrastructure, application, and compliance teams, and influencing technical direction without direct authority Nice to Have: Experience with IAM platforms and security-focused systems (Microsoft Entra, Okta, Idira, SailPoint, Ping, HashiCorp Vault, etc.)Expertise with Infrastructure as Code and configuration management (Terraform, Ansible, Puppet) and scripting (Python, PowerShell, Bash) to reduce manual toil and streamline deploymentsFamiliarity with identity, authentication, authorization, and privileged access management concepts and protocols (OAuth2, OIDC, SAML, LDAP, SCIM)Knowledge of enterprise security architecture and compliance frameworks as they relate to IAM and service resilienceExposure to AIOps or ML-based anomaly detection for proactive reliability management What’s in it for you? We thrive on the challenge to be our best, progressive thinking to keep growing, and working together to deliver trusted advice to help our clients thrive and communities prosper. We care about each other, reaching our potential, making a difference to our communities, and achieving success that is mutual. A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicableLeaders who support your development through coaching and managing opportunitiesAbility to make a difference and lasting impactWork in a dynamic, collaborative, progressive, and high-performing teamA world-class training program in financial servicesOpportunities to do challenging work #TECHPJ Job Skills Agile Working, Application Security, Automation Tools, Business Continuity, Cloud Platform, Cyber Security Management, Decision Making, Enterprise Security Architecture, High Reliability, Hybrid Systems, Identity Access Management (IAM), Information Security Management, Information Technology Security, Infrastructure Penetration Testing, Interpersonal Communication, IT Security Architecture, IT Systems Integration, Microsoft PowerShell, Python (Programming Language), Security Information and Event Management (SIEM), Security Standards, Security Tools, Strategic Thinking, Systems Development Lifecycle (SDLC), Terraform (Software) Additional Job Details Address: 16 YORK ST:TORONTO City: Toronto Country: Canada Work hours/week: 37.5 Employment Type: Full time Platform: TECHNOLOGY AND OPERATIONS Job Type: Regular Pay Type: Salaried Posted Date: 2026-08-25 Application Deadline: 2026-09-08 Note: Applications will be accepted until 11:59 PM on the day prior to the application deadline date above Our Employment Opportunities At RBC, we are guided by living shared values of Client First, Integrity, Collaboration, Respect and Excellence and winning together as One RBC. We believe an inclusive workplace that has diverse perspectives is core to our continued growth as one of the largest and most successful banks in the world. Maintaining a workplace where our employees feel supported to perform at their best, effectively collaborate, drive innovation, and grow professionally helps to bring our Purpose to life and create value for our clients and communities. RBC strives to deliver this through policies and programs intended to foster a workplace based on respect, belonging and opportunity for all. Join our Talent Community Stay in-the-know about great career opportunities at RBC. Sign up and get customized info on our latest jobs, career tips and Recruitment events that matter to you. Expand your limits and create a new future together at RBC. Find out how we use our passion and drive to enhance the well-being of our clients and communities at jobs.rbc.com RBC is presently inviting candidates to apply for this existing vacancy. Applying to this posting allows you to express your interest in this current career opportunity at RBC. Qualified applicants may be contacted to review their resume in more detail.