Site Reliability Engineering Manager

Infoway Solutions — Canada · Posted ~4 hours ago

Lead

Skills

SLI SLO SLA Error budgets Reliability governance Dynatrace Application Performance Monitoring Distributed tracing Real User Monitoring Synthetic monitoring OpenTelemetry Kubernetes Docker OpenShift High availability Disaster recovery Resilience Business continuity Incident management Problem management Root cause analysis Continuous service improvement Operational automation Capacity planning Performance engineering Chaos engineering Observability strategy

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Lead site reliability and observability strategy for complex enterprise environments. You will establish reliability governance, design highly available and resilient architectures, oversee incident and problem management, reduce operational toil, and advance monitoring, automation, capacity, performance, and chaos engineering practices.

Highlights

Lead enterprise reliability and observability initiatives spanning governance, monitoring, resilience, disaster recovery, incident management, automation, and advanced performance engineering.

Description

SRE Architect Vancouver, Canada JD- Experience defining and implementing SLI, SLO, SLA, Error Budget, and Reliability Governance frameworks. Expertise in Dynatrace and other enterprise observability platforms, including APM, Distributed Tracing, RUM, Synthetic Monitoring, and OpenTelemetry. Strong experience with Kubernetes, Docker, OpenShift, and cloud-native architectures. Experience designing High Availability (HA), Disaster Recovery (DR), Resilience, and Business Continuity solutions. Proven track record in Incident Management, Problem Management, RCA, and Continuous Service Improvement. Experience driving toil reduction, self-healing, auto-remediation, and operational automation initiatives. Strong understanding of Capacity Planning, Performance Engineering, and Chaos Engineering practices. Experience establishing enterprise observability strategy, standards, governance, and monitoring frameworks. Familiarity with AIOps, predictive analytics, event correlation, and intelligent alert management. Experience with CI/CD, DevSecOps, GitOps, and Infrastructure as Code practices. Ability to lead production readiness reviews, operational readiness reviews, and reliability assessments. Strong stakeholder management skills with the ability to influence engineering, architecture, and business leadership teams. Experience leading SRE transformation and reliability engineering adoption across large-scale enterprise environments. Relevant certifications in Azure, Kubernetes, Dynatrace, SRE, DevOps, or Cloud Architecture preferred. Most important addition: The architect should have proven experience leading enterprise SRE transformation, observability strategy, SLO governance, reliability engineering adoption, and operational excellence initiatives, not just monitoring tool implementation