Senior Observability / SRE Engineer

Lancesoft — Canada · Posted ~3 hours ago

Senior Contract Hybrid

Skills

observability SRE DevOps Azure Kubernetes Terraform application monitoring distributed systems SLI/SLO incident troubleshooting microservices

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A complex technology environment is seeking a senior SRE/observability engineer to improve reliability and visibility across distributed production systems. You will build observability-as-code solutions with Terraform, operate Kubernetes and cloud infrastructure, troubleshoot difficult incidents, and establish meaningful SLI/SLO practices across microservices and event-driven architectures.

Highlights

Long-term engineering engagement focused on observability, reliability, and mission-critical distributed systems, with hands-on work across cloud infrastructure, Kubernetes, Terraform, and SRE practices.

Description

Title – Sr Support Engineer Duration – 12+ months (With possibility of extension) Location: Montreal, QC- Hybrid (2-4 Days WFO) Job Description: We are seeking an experienced Senior Observability / SRE Engineer to design, implement, and enhance observability and reliability solutions across complex, distributed environments. The ideal candidate will have strong expertise in observability, SRE/DevOps, Azure, Kubernetes, Terraform, and application monitoring, with hands-on experience supporting mission-critical production systems. The successful candidate will build observability-as-code solutions, improve system visibility and reliability, troubleshoot complex distributed-system incidents, and establish effective SLI/SLO practices across microservices and event-driven architectures. Key Responsibilities • Design and implement Observability-as-Code solutions using Terraform for monitoring pipelines, dashboards, alerts, and operational tooling. • Drive observability improvements using Dynatrace, ELK, Splunk, and PagerDuty to provide real-time performance insights and comprehensive system visibility. • Instrument applications for end-to-end observability, including distributed tracing, metrics collection, log aggregation, and correlation IDs. • Support observability across Node.js and .NET microservices and event-driven architectures. • Troubleshoot complex production incidents and perform root cause analysis across services, databases, caches, APIs, and infrastructure. • Design, implement, and troubleshoot Azure Kubernetes Service (AKS) infrastructure and containerized workloads. • Leverage Terraform and Azure managed services, including Azure SQL Managed Instance (SQL MI), Azure Redis, Azure Functions, and Azure Event Grid. • Translate business and technical requirements into resilient, observable systems aligned with defined SLIs and SLOs. • Automate operational processes using Infrastructure-as-Code (IaC) and CI/CD best practices to reduce operational toil. • Lead incident response and remediation for high-severity and mission-critical production issues. • Conduct blameless postmortems, identify corrective actions, and drive reliability improvements. • Apply chaos engineering and tabletop exercises to strengthen system resilience and operational readiness. • Collaborate with development, platform, infrastructure, and business teams to improve availability, scalability, performance, and operational excellence. Required Qualifications • 8+ years of hands-on experience in Observability, Site Reliability Engineering (SRE), DevOps, or related roles. • Deep expertise with observability platforms and tools such as: o Dynatrace o ELK / Elasticsearch / Logstash / Kibana o Splunk o PagerDuty • Strong understanding of observability principles, including: o Instrumentation o Distributed tracing o Metrics o Log aggregation o Correlation IDs o SLIs / SLOs • Advanced hands-on experience with: o Azure Kubernetes Service (AKS) o Terraform o Azure managed services • Strong experience with Azure SQL Managed Instance, Redis, Azure Functions, and Azure Event Grid. • Proven experience instrumenting Node.js and .NET applications for end-to-end observability. • Strong troubleshooting skills across distributed systems, databases, caches, APIs, microservices, and infrastructure. • Hands-on experience with PagerDuty and ServiceNow for incident management. • Strong understanding of Incident, Problem, and Change Management processes. • Experience applying SRE principles, blameless postmortems, and chaos engineering practices. • Strong understanding of Infrastructure-as-Code and CI/CD practices. • Excellent communication, collaboration, problem-solving, and technical leadership skills. Preferred Experience • Experience with event-driven architectures and distributed systems. • Experience developing and implementing observability standards across enterprise environments. • Experience with reliability engineering and production readiness practices. • Experience conducting chaos engineering exercises and operational tabletop sessions. • Experience designing resilient and highly scalable Azure-based architectures. • Experience working in cross-functional Agile/DevOps environments. Core Technical Skills Observability: Dynatrace, Splunk, ELK, PagerDuty Cloud: Microsoft Azure, AKS, Azure SQL MI, Redis, Functions, Event Grid Infrastructure: Terraform, Infrastructure-as-Code, CI/CD Application: Node.js, .NET, Microservices, APIs Reliability: SRE, SLI/SLO, Distributed Tracing, Metrics, Logging, RCA Operations: ServiceNow, Incident Management, Problem Management, Change Management Resilience: Chaos Engineering, Blameless Postmortems, Tabletop Exercises