Summary
✨ AI‑Generated
Lead site reliability and observability strategy for complex enterprise environments. You will establish reliability governance, design highly available and resilient architectures, oversee incident and problem management, reduce operational toil, and advance monitoring, automation, capacity, performance, and chaos engineering practices.
Highlights
Lead enterprise reliability and observability initiatives spanning governance, monitoring, resilience, disaster recovery, incident management, automation, and advanced performance engineering.
Description
SRE Architect
Vancouver, Canada
JD-
Experience defining and implementing SLI, SLO, SLA, Error Budget, and Reliability Governance frameworks.
Expertise in Dynatrace and other enterprise observability platforms, including APM, Distributed Tracing, RUM, Synthetic Monitoring, and OpenTelemetry.
Strong experience with Kubernetes, Docker, OpenShift, and cloud-native architectures.
Experience designing High Availability (HA), Disaster Recovery (DR), Resilience, and Business Continuity solutions.
Proven track record in Incident Management, Problem Management, RCA, and Continuous Service Improvement.
Experience driving toil reduction, self-healing, auto-remediation, and operational automation initiatives.
Strong understanding of Capacity Planning, Performance Engineering, and Chaos Engineering practices.
Experience establishing enterprise observability strategy, standards, governance, and monitoring frameworks.
Familiarity with AIOps, predictive analytics, event correlation, and intelligent alert management.
Experience with CI/CD, DevSecOps, GitOps, and Infrastructure as Code practices.
Ability to lead production readiness reviews, operational readiness reviews, and reliability assessments.
Strong stakeholder management skills with the ability to influence engineering, architecture, and business leadership teams.
Experience leading SRE transformation and reliability engineering adoption across large-scale enterprise environments.
Relevant certifications in Azure, Kubernetes, Dynatrace, SRE, DevOps, or Cloud Architecture preferred.
Most important addition: The architect should have proven experience leading enterprise SRE transformation, observability strategy, SLO governance, reliability engineering adoption, and operational excellence initiatives, not just monitoring tool implementation