Site Reliability Engineer

Axelon Services Corp. — Canada · Posted ~21 hours ago

Senior Contract Hybrid

Skills

Kubernetes Linux Command line Incident response Root cause analysis Observability Automation Troubleshooting Service Mesh Public Cloud Private Cloud

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A 12-month hybrid Site Reliability Engineering contract focused on operating a Kubernetes-based platform across public and private cloud environments. The role covers incident response, troubleshooting, automation, observability, client onboarding, upgrades, and capacity management.

Highlights

A 12-month hybrid SRE contract offering hands-on responsibility for a Kubernetes platform, incident management, automation, observability, capacity planning, and continuous reliability improvements.

Description

Duration: 12 Months contract Work Mode: Hybrid Location: Montreal (Day 1 onboarding onsite/in office presence 3x/week) Responsibilities Operate and maintain the Kubernetes-based platform across public and private cloud environments.Work with observability tooling to ensure alerts are actionable through up-to-date runbooks and documentation.Manage incidents from end-to-end, including incident response, root cause analysis, and post-incident reviews.Onboard and support new clients onto the platform.Build automation and diagnostic tooling that cuts manual effort and evaluates performance.Identify and deliver process improvements.Support software and hardware upgrades and keep components up to date.Proactively manage capacity so the platform scales with demand. Requirements Minimum 5 years of hands-on experience operating Kubernetes in production, ideally with Service Mesh.Strong Linux and command line fundamentals.Confident in debugging and troubleshooting complex systems, from the application layer through to lower-level infrastructure.Experience working with a public cloud provider, preferably Azure or AWS. Preferred Skills Working knowledge of Grafana, Prometheus, Loki, and Tempo.Scripting or coding in Python or Java.Experience with CI/CD and infrastructure as code such as Helm or Terraform.A financial services background is not required. This role is for an existing vacancy.