Senior Site Reliability Engineer

Jmgroupinc — Canada · Posted ~3 hours ago

Senior Full-time

Skills

SRE observability automation Kubernetes incident management cloud-native systems OpenShift Monitoring Logging Automation AI Tools

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Lead reliability engineering practices for enterprise applications by improving monitoring, automation, incident response, and cloud-native operations. Work with advanced tools to enhance system stability and efficiency.

Highlights

Opportunity to improve enterprise reliability through automation, AI-powered operations, and modern cloud practices.

Description

Seeking a Senior SRE to drive system reliability, observability, automation and AI-powered operations across enterprise applications. Key Responsibilities Lead SRE practices covering monitoring, alerting, logging, self-healing and reliability testing.Implement observability, automation and AIOps to improve incident detection, RCA and reduce MTTR.Use Generative AI / AI-powered tools for troubleshooting, runbook automation, knowledge management and production support.Support cloud-native applications, Kubernetes/OpenShift and application deployments.Lead Incident & Problem Management, production troubleshooting and RCA.Partner with development teams to ensure releases meet reliability and performance standards.Automate SRE processes and improve operational efficiency.