Senior Support Engineer

Apptoza Inc — Canada · Posted ~2 hours ago

Senior

Skills

observability SRE DevOps Dynatrace ELK Splunk PagerDuty Azure Kubernetes Service Terraform Azure managed services distributed tracing metrics collection log aggregation Node.js .NET microservices event-driven architectures distributed systems troubleshooting incident management Azure AKS Azure SQL Managed Instance Redis Azure Functions Azure Event Grid

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior reliability engineering opportunity for an experienced observability, SRE, or DevOps professional. You will design and improve cloud infrastructure, instrument distributed applications, troubleshoot complex production issues, and lead incident response across microservices and event-driven systems. Strong expertise with Azure, Kubernetes, Terraform, observability platforms, and application telemetry is central to the role.

Highlights

Senior reliability-focused role requiring deep observability expertise and strong production troubleshooting capabilities. The position combines cloud infrastructure, infrastructure as code, distributed systems, incident management, and application observability across modern microservice environments.

Description

8+ years hands-on experience in observability| SRE| or DevOps roles with proven expertise across infrastructure and application-level reliability.Deep expertise in observability tooling: Dynatrace| ELK| Splunk| and PagerDuty; demonstrated understanding of observability principles (instrumentation| correlation IDs| SLI/SLO frameworks).Advanced proficiency with Azure Kubernetes Service (AKS)| Terraform| and Azure managed services (SQL MI| Redis| Functions| Event Grid); proven ability to design and implement infrastructure-as-code solutions.Strong hands-on experience instrumenting applications for comprehensive observability: distributed tracing| metrics collection| and log aggregation across Node.js and .NET applications in microservices and event-driven architectures.Proven troubleshooting expertise in distributed systemsdiagnosing root causes across multiple service layers| databases| caches| and APIs in production environments.Excellent incident management skills: hands-on experience with PagerDuty and ServiceNow; ability to resolve high-severity incidents rapidly and conduct effective root cause analysis.Knowledge of incident| problem| and change management processes| including SRE principles| blameless postmortems| and chaos engineering practices.Exceptional communication and leadership