Description
Summary
Location: Montreal (day 1 onboarding / onsite presence required 3x/week)Duration: 12 Months ContractSchedule: On-call weekend support might be required (1-2 hours for weekend deployments).
It is on a rotation and always communicated in advance.
Responsibilities
Design, build, maintain, and enhance CI/CD pipelines and supporting build, test, release, and deployment infrastructure.Develop and maintain automated deployment solutions for applications running in Kubernetes and containerized environments.Partner with application development teams to improve build, test, deployment, and release processes.Support the deployment of applications, configuration changes, patches, and platform upgrades across development, testing, and production environments.Build and maintain Kubernetes deployment artifacts, including YAML configuration, Helm charts, or equivalent packaging and configuration mechanisms.Develop automation using Python, Linux shell scripting, and related tools.Manage and improve GitHub repositories, branching strategies, pull-request workflows, access controls, and automated repository processes.Integrate automated testing, code-quality validation, security scanning, dependency checks, and linting into CI/CD pipelines.Troubleshoot application, infrastructure, container, Kubernetes, network, and deployment issues in complex environments.Investigate production incidents, identify root causes, and implement preventive or corrective engineering solutions.Apply SRE principles to improve system reliability, availability, scalability, observability, and operational readiness.Define and improve monitoring, alerting, dashboards, operational metrics, and production support procedures.Automate routine operational activities to reduce manual effort and operational risk.Collaborate with infrastructure, network, database, cybersecurity, release management, and application teams.Create and maintain technical documentation, operational runbooks, deployment procedures, and troubleshooting guides.Participate in design reviews, production-readiness reviews, incident reviews, and continuous-improvement initiatives.Participate in an after-hours production support and on-call rotation when required.
Requirements
Minimum of 5 years of relevant experience in DevOps, SRE, production engineering, platform engineering, infrastructure engineering, or a related discipline.Strong hands-on experience with Kubernetes, including application deployment, configuration, troubleshooting, scaling, services, ingress, secrets, and operational support.Strong hands-on experience with Docker/Podman and containerized application environments.Strong Linux and UNIX system administration and troubleshooting skills.Strong experience with Linux shell scripting, such as Bash or KornShell.Hands-on programming and automation experience using Python or a comparable language.Strong understanding of CI/CD concepts, software delivery lifecycles, release automation, and deployment strategies.Hands-on experience with CI/CD and artifact-management tools such as Jenkins and Artifactory, or equivalent platforms.Hands-on experience with Git and GitHub, including repository management, pull requests, branching strategies, release workflows, and automated checks.Experience developing automation using YAML, Ansible, or an equivalent automation framework.Strong understanding of networking concepts, including DNS, TCP/IP, HTTP/HTTPS, TLS, proxies, firewalls, routing, load balancing, and network troubleshooting.Experience supporting applications and resolving production issues in a fast-paced environment.Understanding of SRE practices, including monitoring, incident response, root-cause analysis, service reliability, operational readiness, and continuous improvement.Experience integrating code-quality tools, security scanning, automated testing, and policy controls into CI/CD pipelines.Ability to troubleshoot issues across application, operating system, container, network, infrastructure, and database layers.Experience working in an Agile development environment and using tools such as Jira.Strong written and verbal communication skills.Ability to collaborate effectively with globally distributed engineering and support teams.Ability to prioritize work, manage multiple tasks, and deliver results with limited supervision.Ability to participate in an after-hours on-call support rotation.
Preferred Skills
Experience with enterprise Kubernetes platforms such as OpenShift or another managed Kubernetes environment.Experience provisioning on-demand environments using virtual machines and containers.Experience working with Azure, AWS, GCP, or another cloud platform.Experience with infrastructure-as-code and configuration-management technologies.Experience with Kubernetes package-management and deployment tools such as Helm.Experience with GitOps deployment models and tools.Knowledge of service mesh technologies, container networking, ingress controllers, and API gateways.Experience with observability and telemetry platforms, including metrics, logs, traces, dashboards, and alerting.Experience developing or maintaining Grafana dashboards.Familiarity with production incident management, problem management, change management, service operations, and release management.Understanding of high availability, disaster recovery, capacity management, and production resiliency.Experience with relational database technologies such as DB2, Sybase, or Oracle.Experience working in financial services or another regulated enterprise environment.Familiarity with secure software supply-chain practices, secrets management, certificate management, and vulnerability remediation.Education: Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related discipline, or equivalent practical industry experience.
This role is for an existing vacancy.