DevOps, Kubernetes and Site Reliability Engineer

Gala Solutions Inc — Canada · Posted ~4 hours ago

Senior Full-time Onsite

Skills

DevOps Kubernetes Site reliability engineering CI/CD Cloud platforms Infrastructure automation Platform operations

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Work as an experienced DevOps, Kubernetes, and Site Reliability Engineer supporting reliable application platforms and delivery capabilities. You will design and automate infrastructure, improve CI/CD, collaborate across development and security teams, and strengthen platform stability and operational resilience.

Highlights

Hands-on engineering role focused on reliable application platforms, CI/CD, automation, operational resilience, and engineering efficiency, with cross-functional collaboration and advance notice for scheduled weekend deployments.

Description

Job Title: DevOps, Kubernetes, and Site Reliability Engineer Location: Montreal (day 1 onboarding / onsite presence required 3x/week) Years of experience: 5-7 Schedule: On-call weekend support might be required (1-2 hours for weekend deployments). It is on a rotation and always communicated in advance DESCRIPTION Position Description We are looking for an experienced DevOps, Kubernetes, and Site Reliability Engineer with a minimum of five years of relevant industry experience. Experience within financial services or another highly regulated technology environment is preferred. The successful candidate will join the Operations Technology team and help design, automate, deploy, and support reliable application platforms and CI/CD capabilities. The role will work closely with application development, infrastructure, network, security, and production support teams to improve software delivery, platform stability, operational resilience, and engineering efficiency. The position requires strong hands-on experience with Kubernetes, Docker/Podman, Linux, networking, CI/CD pipelines, GitHub, Python, and shell scripting. The candidate should understand modern DevOps and SRE practices and be comfortable supporting production systems, troubleshooting complex application and infrastructure issues, and automating repetitive operational tasks. The ideal candidate will be passionate about automation, production reliability, and continuous improvement. The candidate should be organized, disciplined, detail-oriented, self-motivated, collaborative, and focused on delivering measurable engineering outcomes. Required Skills Minimum of 5 years of relevant experience in DevOps, SRE, production engineering, platform engineering, infrastructure engineering, or a related discipline. Strong hands-on experience with Kubernetes, including application deployment, configuration, troubleshooting, scaling, services, ingress, secrets, and operational support. Strong hands-on experience with Docker/Podman and containerized application environments. Strong Linux and UNIX system administration and troubleshooting skills. Strong experience with Linux shell scripting, such as Bash or KornShell. Hands-on programming and automation experience using Python or a comparable language. Strong understanding of CI/CD concepts, software delivery lifecycles, release automation, and deployment strategies. Hands-on experience with CI/CD and artifact-management tools such as Jenkins and Artifactory, or equivalent platforms. Hands-on experience with Git and GitHub, including repository management, pull requests, branching strategies, release workflows, and automated checks. Experience developing automation using YAML, Ansible, or an equivalent automation framework. Strong understanding of networking concepts, including DNS, TCP/IP, HTTP/HTTPS, TLS, proxies, firewalls, routing, load balancing, and network troubleshooting. Experience supporting applications and resolving production issues in a fast-paced environment. Understanding of SRE practices, including monitoring, incident response, root-cause analysis, service reliability, operational readiness, and continuous improvement. Experience integrating code-quality tools, security scanning, automated testing, and policy controls into CI/CD pipelines. Ability to troubleshoot issues across application, operating system, container, network, infrastructure, and database layers. Experience working in an Agile development environment and using tools such as Jira. Strong written and verbal communication skills. Ability to collaborate effectively with globally distributed engineering and support teams. Ability to prioritize work, manage multiple tasks, and deliver results with limited supervision. Ability to participate in an after-hours on-call support rotation. Desired Skills Experience with enterprise Kubernetes platforms such as OpenShift or another managed Kubernetes environment. Experience provisioning on-demand environments using virtual machines and containers. Experience working with Azure, AWS, GCP, or another cloud platform. Experience with infrastructure-as-code and configuration-management technologies. Experience with Kubernetes package-management and deployment tools such as Helm. Experience with GitOps deployment models and tools. Knowledge of service mesh technologies, container networking, ingress controllers, and API gateways. Experience with observability and telemetry platforms, including metrics, logs, traces, dashboards, and alerting. Experience developing or maintaining Grafana dashboards. Familiarity with production incident management, problem management, change management, service operations, and release management. Understanding of high availability, disaster recovery, capacity management, and production resiliency. Experience with relational database technologies such as DB2, Sybase, or Oracle. Experience working in financial services or another regulated enterprise environment. Familiarity with secure software supply-chain practices, secrets management, certificate management, and vulnerability remediation. Education: Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related discipline, or equivalent practical industry experience.