Site Reliability Engineer

Yochana — Canada · Posted ~3 hours ago

Senior Contract Hybrid

Skills

Site Reliability Engineering Production support DevOps Cloud operations Monitoring Alerting Logging Observability Kubernetes Docker CI/CD Incident management Root cause analysis Disaster recovery Cloud platforms

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A Site Reliability Engineer is sought for a hybrid role focused on keeping applications and infrastructure healthy, available, and resilient. You will manage monitoring, alerting, logging, observability, cloud and container platforms, automate operational processes, support CI/CD, troubleshoot production incidents, and contribute to capacity planning and disaster recovery.

Highlights

Own reliability and availability across applications and infrastructure in a hybrid environment. The role offers broad exposure to observability, cloud platforms, Kubernetes, Docker, automation, incident response, disaster recovery, and platform resiliency.

Description

Position Name – SRE Type of hiring – Subcon Location – Brampton, ON or Toronto, ON (Hybrid) Job Description: Key Responsibilities & Skills: Monitor application and infrastructure health, performance, and availability.Troubleshoot production incidents, perform root cause analysis, and drive issue resolution.Implement and manage monitoring, alerting, logging, and observability solutions.Support cloud platforms, Kubernetes, Docker, and infrastructure operations.Automate operational processes and improve system reliability and efficiency.Support CI/CD pipelines, application deployments, and release activities.Perform capacity planning, performance optimization, and disaster recovery planning.Collaborate with engineering, cloud, and security teams to improve platform resiliency.Create operational dashboards, runbooks, and technical documentation. Required Skills: 6+ Years of experience in Site Reliability Engineering, Production Support, DevOps, or Cloud Operations.Experience supporting production applications and cloud environments (Azure preferred).Strong knowledge of Kubernetes, Docker, Linux, networking, and system administration.Experience with monitoring tools such as Dynatrace, Grafana, or Azure Monitor.Scripting and automation experience using Python, PowerShell, or Bash.