Site Reliability Engineer

Envision Tech Sol — Canada · Posted ~8 hours ago

Mid

Skills

Splunk Dynatrace Grafana Datadog incident response root cause analysis ServiceNow reliability engineering Python Go Bash

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking a reliability-focused engineer to operate enterprise platforms, improve system resilience, manage incidents, and drive automation. The role requires strong observability experience, troubleshooting skills, and collaboration across technical teams.

Highlights

Opportunity to work on large-scale enterprise reliability challenges with a focus on monitoring, automation, incident management, and cross-functional collaboration.

Description

Hands-on experience with monitoring, observability, and alerting tools, specifically Splunk, Dynatrace, Grafana and Datadog.Proven experience operating and supporting a large-scale enterprise platform environment.Demonstrated experience with incident response and leading or contributing to root cause analysis (RCA) processes.Strong understanding of reliability engineering principles, including availability, resiliency, monitoring and alerting best practices.Experience with ticketing and incident management workflows (e.g., ServiceNow).Excellent communication skills, with the ability to drive remediation efforts and collaborate across technical, product and risk teams. What would be great to have: Experience in financial services, payments, or embedded finance environments.Proficiency with scripting or programming languages for automation (e.g., Python, Go, Bash).Familiarity with cloud platforms, containerization, and CI/CD pipelines.Experience defining and managing SLOs, SLIs and error budgets