Site Reliability Engineer

Lancesoft — Canada · Posted ~3 hours ago

Senior CAD 60-65/hr

Skills

8+ years of hands-on experience in observability, SRE, or DevOps Dynatrace ELK Splunk PagerDuty Observability principles SLI/SLO frameworks Azure Kubernetes Service (AKS) Terraform Azure managed services Distributed tracing Metrics collection Log aggregation Node.js .NET Microservices Event-driven architectures Distributed systems troubleshooting Production incident troubleshooting Azure AKS Azure SQL MI Redis Azure Functions Azure Event Grid

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A well-compensated Site Reliability Engineering role for an experienced engineer who can improve reliability across infrastructure and applications. You will work with cloud infrastructure, Kubernetes, infrastructure as code, observability platforms, distributed tracing, metrics, logs, and production incident response across modern microservices and event-driven systems.

Highlights

High-impact reliability engineering role with strong compensation, deep exposure to modern cloud infrastructure and observability, and opportunities to solve complex production issues across distributed systems.

Description

Pay Range: CAD 60-65/hr Must-have: 8+ years hands-on experience in observability| SRE| or DevOps roles with proven expertise across infrastructure and application-level reliability.Deep expertise in observability tooling: Dynatrace| ELK| Splunk| and PagerDuty; demonstrated understanding of observability principles (instrumentation| correlation IDs| SLI/SLO frameworks).Advanced proficiency with Azure Kubernetes Service (AKS)| Terraform| and Azure managed services (SQL MI| Redis| Functions| Event Grid); proven ability to design and implement infrastructure-as-code solutions.Strong hands-on experience instrumenting applications for comprehensive observability: distributed tracing| metrics collection| and log aggregation across Node.js and .NET applications in microservices and event-driven architectures.Proven troubleshooting expertise in distributed systemsdiagnosing root causes across multiple service layers| databases| caches| and APIs in production environments.Excellent incident management skills: hands-on experience with PagerDuty and ServiceNow; ability to resolve high-severity incidents rapidly and conduct effective root cause analysis.Knowledge of incident| problem| and change management processes| including SRE principles| blameless postmortems| and chaos engineering practices.Exceptional communication and leadership