Description
Title – Sr Support Engineer
Duration – 12+ months (With possibility of extension)
Location: Montreal, QC- Hybrid (2-4 Days WFO)
Job Description:
We are seeking an experienced Senior Observability / SRE Engineer to design, implement, and enhance observability and reliability solutions across complex, distributed environments.
The ideal candidate will have strong expertise in observability, SRE/DevOps, Azure, Kubernetes, Terraform, and application monitoring, with hands-on experience supporting mission-critical production systems.
The successful candidate will build observability-as-code solutions, improve system visibility and reliability, troubleshoot complex distributed-system incidents, and establish effective SLI/SLO practices across microservices and event-driven architectures.
Key Responsibilities
• Design and implement Observability-as-Code solutions using Terraform for monitoring pipelines, dashboards, alerts, and operational tooling.
• Drive observability improvements using Dynatrace, ELK, Splunk, and PagerDuty to provide real-time performance insights and comprehensive system visibility.
• Instrument applications for end-to-end observability, including distributed tracing, metrics collection, log aggregation, and correlation IDs.
• Support observability across Node.js and .NET microservices and event-driven architectures.
• Troubleshoot complex production incidents and perform root cause analysis across services, databases, caches, APIs, and infrastructure.
• Design, implement, and troubleshoot Azure Kubernetes Service (AKS) infrastructure and containerized workloads.
• Leverage Terraform and Azure managed services, including Azure SQL Managed Instance (SQL MI), Azure Redis, Azure Functions, and Azure Event Grid.
• Translate business and technical requirements into resilient, observable systems aligned with defined SLIs and SLOs.
• Automate operational processes using Infrastructure-as-Code (IaC) and CI/CD best practices to reduce operational toil.
• Lead incident response and remediation for high-severity and mission-critical production issues.
• Conduct blameless postmortems, identify corrective actions, and drive reliability improvements.
• Apply chaos engineering and tabletop exercises to strengthen system resilience and operational readiness.
• Collaborate with development, platform, infrastructure, and business teams to improve availability, scalability, performance, and operational excellence.
Required Qualifications
• 8+ years of hands-on experience in Observability, Site Reliability Engineering (SRE), DevOps, or related roles.
• Deep expertise with observability platforms and tools such as:
o Dynatrace
o ELK / Elasticsearch / Logstash / Kibana
o Splunk
o PagerDuty
• Strong understanding of observability principles, including:
o Instrumentation
o Distributed tracing
o Metrics
o Log aggregation
o Correlation IDs
o SLIs / SLOs
• Advanced hands-on experience with:
o Azure Kubernetes Service (AKS)
o Terraform
o Azure managed services
• Strong experience with Azure SQL Managed Instance, Redis, Azure Functions, and Azure Event Grid.
• Proven experience instrumenting Node.js and .NET applications for end-to-end observability.
• Strong troubleshooting skills across distributed systems, databases, caches, APIs, microservices, and infrastructure.
• Hands-on experience with PagerDuty and ServiceNow for incident management.
• Strong understanding of Incident, Problem, and Change Management processes.
• Experience applying SRE principles, blameless postmortems, and chaos engineering practices.
• Strong understanding of Infrastructure-as-Code and CI/CD practices.
• Excellent communication, collaboration, problem-solving, and technical leadership skills.
Preferred Experience
• Experience with event-driven architectures and distributed systems.
• Experience developing and implementing observability standards across enterprise environments.
• Experience with reliability engineering and production readiness practices.
• Experience conducting chaos engineering exercises and operational tabletop sessions.
• Experience designing resilient and highly scalable Azure-based architectures.
• Experience working in cross-functional Agile/DevOps environments.
Core Technical Skills
Observability: Dynatrace, Splunk, ELK, PagerDuty
Cloud: Microsoft Azure, AKS, Azure SQL MI, Redis, Functions, Event Grid
Infrastructure: Terraform, Infrastructure-as-Code, CI/CD
Application: Node.js, .NET, Microservices, APIs
Reliability: SRE, SLI/SLO, Distributed Tracing, Metrics, Logging, RCA
Operations: ServiceNow, Incident Management, Problem Management, Change Management
Resilience: Chaos Engineering, Blameless Postmortems, Tabletop Exercises