Senior SRE/AIOps Engineer

Rbc — Canada · Posted ~7 hours ago

Senior

Skills

Site Reliability Engineering AIOps Monitoring Automation Incident management Dependency management Service ID management Certificate management SSO Vulnerability remediation Operational support SRE Snowflake SaaS

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior SRE/AIOps position responsible for designing and maintaining reliability, monitoring, automation, and self-healing capabilities across a complex enterprise technology environment. You will coordinate incidents, manage dependencies and escalations, address security and compliance needs, and provide end-to-end operational support for integrated data-processing systems.

Highlights

Senior SRE/AIOps role with end-to-end operational ownership across complex enterprise systems. The position combines reliability engineering, proactive monitoring, self-healing automation, incident coordination, security and compliance responsibilities, and support for integrated data pipelines.

Description

Job Description WHAT IS THE OPPORTUNITY? This role is responsible for designing, implementing, and maintaining SRE (Site Reliability Engineering) and AIOps (Artificial Intelligence for IT Operations) capabilities to ensure system reliability, proactive monitoring, and automation of self-healing operations. In addition to day-to-day support, the position provides end-to-end operational ownership across systems managed by multiple enterprise teams, including incident coordination, dependency management, and escalation. The role is also responsible for key security and compliance functions such as service ID and certificate management, SSO updates, vulnerability remediation, and lifecycle management of end-of-life components. Our team supports a portfolio of multi-platform HR data pipelines that move and process data into Snowflake through multiple integrated components, requiring end-to-end monitoring, coordination, and support across systems. In addition, we support SaaS-based applications that are primarily vendor-managed, while we retain responsibility for integration, access management, monitoring, and operational oversight. WHAT WILL YOU DO? Key Responsibilities: SRE & Reliability Engineering Define and operationalize SLIs, SLOs, and error budgetsOwn the incident management lifecycle (detection → triage → resolution → RCA → prevention)Lead problem management and eliminate recurring issuesDevelop and maintain runbooks, playbooks, and recovery proceduresDrive resilience engineering, including failover testing and capacity planning AIOps, Observability & Logging Implement and optimize AIOps capabilities using platforms such as MoogsoftLeverage Dynatrace for deep APM insightsIntegrate alerting and escalation workflows with PagerDutyUtilize synthetic monitoring via CatchpointDesign and maintain centralized logging solutions using the ELK stack (Elasticsearch, Logstash, Kibana) and enterprise Logging as a Service (LaaS) platformsPerform log analysis, correlation, and anomaly detection to support proactive issue identificationDrive event noise reduction and intelligent alerting strategiesBuild dashboards and observability KPIs for operational insights Automation & Self-Healing Systems Design and implement automation-first solutions using Ansible and scripting (Python, Bash)Enable self-healing capabilities (auto-remediation, restart logic, workflow recovery)Orchestrate workflows using StonebranchReduce operational toil through automation and continuous improvement Application & Data Platform Support Provide L2/L3 support for data pipelines and integration workflows, including Snowflake ingestion and transformation processesSupport Snowflake pipelines, ETL workflows, and orchestration dependenciesTroubleshoot across distributed systems including:Object storage (e.g., S3)APIs, messaging, and file transfer systemsSupport containerized workloads on OpenShift Infrastructure & Platform Expertise Administer and support Windows Server and IIS-based applicationsManage certificate lifecycle (TLS, SAML, OAuth)Support SSO integrations via Microsoft Entra IDWork with relational databases (SQL Server, PostgreSQL) for troubleshooting and performance tuning Disaster Recovery & Resiliency Lead and coordinate Disaster Recovery (DR) planning and executionValidate failover processes across dependent systemsEnsure end-to-end DR readiness across integrated platformsDocument recovery strategies and participate in DR exercises Governance, Compliance & Collaboration Ensure adherence to enterprise security, compliance, and audit requirementsCollaborate with IAM, PAM, logging, and cloud platform teamsSupport change management, CAB processes, and release coordinationProvide reporting, KPIs, and executive summaries on system health WHAT DO YOU NEED TO SUCCEED? Must Have: 3+ years of SRE or Systems Engineering experience with strong technical expertise.Experience with ServiceNow, ITSM processes including incident, problem, change, and release management.Demonstrated ability to work independently, take ownership, and drive projects to completion.Knowledge of containerized platforms such as Kubernetes or OpenShift.Experience working with Ansible Automation Platform or strong willingness to learn.In-depth knowledge of monitoring and observability tools such as Dynatrace, Elasticsearch, Moogsoft, Catchpoint, and PagerDuty.Knowledge and experience with scripting languages such as BASH, Python and PowerShell.Experience with Linux and Windows Server administration.Experience with or strong interest in intelligent monitoring, anomaly detection, and automation technologiesExcellent problem-solving skills and attention to detail.Experience with API's and secure file transfer procedures. Nice to Have: Experience with telemetry standardization (OpenTelemetry) and observability data correlation.Understanding of AI/ML concepts and their application to observability and operations (AIOps).Experience with tools such as Ansible, Stonebranch, Kafka, and their role in system reliability.Experience with CI/CD and developer platform tools such as Jenkins, GitHub, GitHub Actions, Artifactory, and Vault.Familiarity with Snowflake and MongoDB Atlas, including experience with writing and executing basic queries, validating data, and troubleshooting production issues.Understanding of Single Sign-On(SSO) technologies, including Microsoft Entra ID, SAML and OAuth.Experience supporting enterprise applications in banking or financial services industry with understanding of regulatory, security and compliance requirements. WHAT'S IN IT FOR YOU? We thrive on the challenge to be our best, progressive thinking to keep growing, and working together to deliver trusted advice to help our clients thrive and communities prosper. We care about each other, reaching our potential, making a difference to our communities, and achieving success that is mutual. A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicableLeaders who support your development through coaching and managing opportunitiesAbility to make a difference and lasting impactWork in a dynamic, collaborative, progressive, and high-performing teamA world-class training program in financial servicesOpportunities to do challenging work #TECHPJ Job Skills Critical Thinking, Customer Support Systems, Group Problem Solving, Installation Support, IT Service Level Management, IT Service Management (ITSM), IT Standards, Technical Troubleshooting Additional Job Details Address: 20 KING ST W:TORONTO City: Toronto Country: Canada Work hours/week: 37.5 Employment Type: Full time Platform: TECHNOLOGY AND OPERATIONS Job Type: Regular Pay Type: Salaried Posted Date: 2026-07-23 Application Deadline: 2026-09-28 Note: Applications will be accepted until 11:59 PM on the day prior to the application deadline date above Our Employment Opportunities At RBC, we are guided by living shared values of Client First, Integrity, Collaboration, Respect and Excellence and winning together as One RBC. We believe an inclusive workplace that has diverse perspectives is core to our continued growth as one of the largest and most successful banks in the world. Maintaining a workplace where our employees feel supported to perform at their best, effectively collaborate, drive innovation, and grow professionally helps to bring our Purpose to life and create value for our clients and communities. RBC strives to deliver this through policies and programs intended to foster a workplace based on respect, belonging and opportunity for all. Join our Talent Community Stay in-the-know about great career opportunities at RBC. Sign up and get customized info on our latest jobs, career tips and Recruitment events that matter to you. Expand your limits and create a new future together at RBC. Find out how we use our passion and drive to enhance the well-being of our clients and communities at jobs.rbc.com RBC is presently inviting candidates to apply for this existing vacancy. Applying to this posting allows you to express your interest in this current career opportunity at RBC. Qualified applicants may be contacted to review their resume in more detail.