Senior Site Reliability Engineer

Gala Solutions Inc — Canada · Posted ~22 hours ago

Senior Contract Onsite

Skills

Kubernetes platform operations cloud infrastructure automation troubleshooting system reliability API platforms public cloud hybrid cloud on-premises infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Take ownership of the day-to-day reliability and improvement of a Kubernetes-based API platform. You will operate infrastructure across public cloud, hybrid, and on-premises environments while engineering automation, diagnostic tests, and tooling that improve stability and developer experience. This senior contract role is ideal for an engineer who enjoys combining hands-on operations with practical software engineering.

Highlights

Senior 12-month contract focused on both engineering and operations. Work with Kubernetes-based platforms across public cloud, hybrid cloud, and on-premises environments while building automation, diagnostic tooling, and practical reliability improvements.

Description

Job Title: Site Reliability Engineer Experience Level: Level 3 (senior): 5-7 years Open positions : 1 (potentially 2) Job Level: FTC 12 Months contract Location: M ontreal (Day 1 onboarding onsite/in office presence 3x/week) ***'s API Platform team champions an API first approach across the Firm's development teams, with a consistent focus on delivering an outstanding Developer Experience. We deploy and support solutions across a range of architectures, including cloud, hybrid cloud and on-premises environments. Operational Engineers ensure the platform remains stable, performant and well supported. The role is as much about engineering as it is operations: beyond keeping the platform healthy, it involves building automation and tooling, designing diagnostic tests and engineering practical improvements to how the platform runs. The focus is squarely on running and improving the platform day to day. Responsibilities: Operate and maintain the Kubernetes based platform across public and private cloud environments. Work with observability tooling to ensure alerts are actionable through up-to-date runbooks and documentation. Manage incidents from end-to-end, including incident response, root cause analysis, and post incident reviews. Onboard and support new clients onto the platform. Build automation and diagnostic tooling that cuts manual effort and evaluates performance. Identify and deliver process improvements. Support software and hardware upgrades and keep components up to date. Proactively manage capacity so the platform scales with demand. Required Skills: Hands on experience operating Kubernetes in production, ideally with Service Mesh. Strong Linux and command line fundamentals. Confident debugging & troubleshooting complex systems, from the application layer through to lower-level infrastructure. Experience working with a public cloud provider, preferable Azure or AWS. Working knowledge of Grafana, Prometheus, Loki and Tempo is a plus. Scripting or coding in Python or Java is a strong plus. CI/CD, infrastructure as code such as Helm or Terraform is a plus A financial services background is not required.