Senior Observability Engineer

Astra North Infoteck Inc — Canada · Posted ~17 hours ago

Senior Hybrid

Skills

AKS Terraform Dynatrace Splunk ELK PagerDuty ServiceNow observability as code distributed tracing metrics collection log aggregation incident troubleshooting SLI/SLO Azure Node.js .NET SLI SLO

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior observability engineering role in a hybrid environment, focused on building monitoring and alerting capabilities as code. You will work with Terraform and leading observability platforms, instrument microservices and event-driven systems for tracing, metrics, and logs, and troubleshoot complex production incidents across infrastructure, databases, caches, and APIs.

Highlights

Design observability-as-code solutions across distributed systems, gain deep exposure to leading monitoring platforms, and solve complex production incidents using modern reliability practices.

Description

Job Description Senior Observability Engineer Key Skills: AKS + Terraform, Dynatrace/Splunk/ELK, PagerDuty/ServiceNow Montreal, QC - Hybrid (2-4 Days WFO Role Descriptions: What will you do? • Design and implement observability-as-code solutions using Terraform to deploy monitoring pipelines, dashboards, and alerting strategies across distributed systems. • Drive observability improvements leveraging industry-leading tools (Dynatrace, ELK, Splunk, PagerDuty) to achieve real-time performance insights and comprehensive system visibility. • Instrument applications for end-to-end observability implementing distributed tracing, metrics collection, and log aggregation across Node.js and .NET microservices and event-driven architectures. • Troubleshoot complex incidents in production environments, diagnosing root causes across multiple service layers, databases, caches, and APIs under load using SLI/SLO frameworks. • Investigate and resolve Azure Kubernetes Service (AKS) infrastructure, ensuring reliability and scalability of containerized workloads with deep proficiency in Terraform and Azure managed services (SQL MI, Redis, Functions, Event Grid). • Translate business requirements into observable, resilient systems that meet defined SLIs/SLOs and drive reliability improvements. • Automate operational tasks to reduce toil and improve system resilience through infrastructure-as-code and CI/CD best practices. • Lead incident response and remediation for mission-critical systems, conducting blameless postmortems and building resilience through chaos engineering and tabletop exercises. • Collaborate cross-functionally with development, platform, and business teams to improve service availability, scalability, and operational excellence. What do you need to succeed? Must-have: • 8+ years hands-on experience in observability, SRE, or DevOps roles with proven expertise across infrastructure and application-level reliability. • Deep expertise in observability tooling: Dynatrace, ELK, Splunk, and PagerDuty; demonstrated understanding of observability principles (instrumentation, correlation IDs, SLI/SLO frameworks). • Advanced proficiency with Azure Kubernetes Service (AKS), Terraform, and Azure managed services (SQL MI, Redis, Functions, Event Grid); proven ability to design and implement infrastructure-as-code solutions. • Strong hands-on experience instrumenting applications for comprehensive observability: distributed tracing, metrics collection, and log aggregation across Node.js and .NET applications in microservices and event-driven architectures. • Proven troubleshooting expertise in distributed systems diagnosing root causes across multiple service layers, databases, caches, and APIs in production environments. • Excellent incident management skills: hands-on experience with PagerDuty and ServiceNow; ability to resolve high-severity incidents rapidly and conduct effective root cause analysis. • Knowledge of incident, problem, and change management processes, including SRE principles, blameless postmortems, and chaos engineering practices. • Exceptional communication and leadership.