Site Reliability Engineer

Open Systems Technologies — Canada · Posted ~21 hours ago

Mid Full-time

Skills

Kubernetes Linux command line system debugging troubleshooting public cloud incident response root cause analysis automation observability Azure AWS Service Mesh

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Operate and improve a Kubernetes-based platform across public and private cloud environments. You will own incident response, troubleshoot complex infrastructure, build automation and diagnostic tools, improve observability and documentation, manage capacity, and help keep platform components reliable and up to date.

Highlights

Hands-on reliability role covering production Kubernetes, cloud infrastructure, observability, incident management, automation, capacity planning, and continuous improvement. Strong exposure to complex production systems.

Description

Responsibilities: Operate and maintain the Kubernetes based platform across public and private cloud environments.Work with observability tooling to ensure alerts are actionable through up-to-date runbooks and documentation.Manage incidents from end-to-end, including incident response, root cause analysis, and post incident reviews.Onboard and support new clients onto the platform.Build automation and diagnostic tooling that cuts manual effort and evaluates performance.Identify and deliver process improvements.Support software and hardware upgrades and keep components up to date.Proactively manage capacity so the platform scales with demand.Required Skills: Hands on experience operating Kubernetes in production, ideally with Service Mesh.Strong Linux and command line fundamentals.Confident debugging & troubleshooting complex systems, from the application layer through to lower-level infrastructure.Experience working with a public cloud provider, preferable Azure or AWS.Working knowledge of Grafana, Prometheus, Loki and Tempo is a plus.Scripting or coding in Python or Java is a strong plus.CI/CD, infrastructure as code such as Helm or Terraform is a plusA financial services background is not required.