Summary
✨ AI‑Generated
A senior observability engineering role in a hybrid environment, focused on building monitoring and alerting capabilities as code. You will work with Terraform and leading observability platforms, instrument microservices and event-driven systems for tracing, metrics, and logs, and troubleshoot complex production incidents across infrastructure, databases, caches, and APIs.
Highlights
Design observability-as-code solutions across distributed systems, gain deep exposure to leading monitoring platforms, and solve complex production incidents using modern reliability practices.
Description
Job Description
Senior Observability Engineer
Key Skills: AKS + Terraform, Dynatrace/Splunk/ELK, PagerDuty/ServiceNow
Montreal, QC - Hybrid (2-4 Days WFO
Role Descriptions:
What will you do?
• Design and implement observability-as-code solutions using Terraform to deploy monitoring pipelines, dashboards, and alerting strategies across distributed systems.
• Drive observability improvements leveraging industry-leading tools (Dynatrace, ELK, Splunk, PagerDuty) to achieve real-time performance insights and comprehensive system visibility.
• Instrument applications for end-to-end observability implementing distributed tracing, metrics collection, and log aggregation across Node.js and .NET microservices and event-driven architectures.
• Troubleshoot complex incidents in production environments, diagnosing root causes across multiple service layers, databases, caches, and APIs under load using SLI/SLO frameworks.
• Investigate and resolve Azure Kubernetes Service (AKS) infrastructure, ensuring reliability and scalability of containerized workloads with deep proficiency in Terraform and Azure managed services (SQL MI, Redis, Functions, Event Grid).
• Translate business requirements into observable, resilient systems that meet defined SLIs/SLOs and drive reliability improvements.
• Automate operational tasks to reduce toil and improve system resilience through infrastructure-as-code and CI/CD best practices.
• Lead incident response and remediation for mission-critical systems, conducting blameless postmortems and building resilience through chaos engineering and tabletop exercises.
• Collaborate cross-functionally with development, platform, and business teams to improve service availability, scalability, and operational excellence.
What do you need to succeed?
Must-have:
• 8+ years hands-on experience in observability, SRE, or DevOps roles with proven expertise across infrastructure and application-level reliability.
• Deep expertise in observability tooling: Dynatrace, ELK, Splunk, and PagerDuty; demonstrated understanding of observability principles (instrumentation, correlation IDs, SLI/SLO frameworks).
• Advanced proficiency with Azure Kubernetes Service (AKS), Terraform, and Azure managed services (SQL MI, Redis, Functions, Event Grid); proven ability to design and implement infrastructure-as-code solutions.
• Strong hands-on experience instrumenting applications for comprehensive observability: distributed tracing, metrics collection, and log aggregation across Node.js and .NET applications in microservices and event-driven architectures.
• Proven troubleshooting expertise in distributed systems diagnosing root causes across multiple service layers, databases, caches, and APIs in production environments.
• Excellent incident management skills: hands-on experience with PagerDuty and ServiceNow; ability to resolve high-severity incidents rapidly and conduct effective root cause analysis.
• Knowledge of incident, problem, and change management processes, including SRE principles, blameless postmortems, and chaos engineering practices.
• Exceptional communication and leadership.