Summary
✨ AI‑Generated
A technology organization is seeking an experienced Platform and Site Reliability Engineer to design, deploy, and manage scalable containerized applications. You will own observability, CI/CD, authentication integrations, production troubleshooting, and infrastructure automation while partnering closely with application teams.
Highlights
Design and operate scalable cloud-native platforms, improve observability and reliability, automate infrastructure operations, and collaborate with application teams in a hybrid environment.
Description
Position Name – Platform & SRE Engineer
Type of hiring – Fulltime
Location – Montreal, QC (Hybrid – 3 days Onsite)
Job Description:
Seeking a Platform & Site Reliability Engineer (SRE) with a strong background in modern cloud-native technologies and infrastructure automation.
Looking for a highly skilled professional with 6+ Years of experience in platform engineering, container orchestration, and site reliability practices.
Key Responsibilities
Design, deploy, and manage scalable, containerized applications on Kubernetes clusters.Build and maintain CI/CD pipelines to enable reliable and repeatable deployments.Own and improve platform observability using tools like Prometheus, Grafana, and alerts.Work closely with application teams to onboard services with proper authentication and authorization (OAuth2, SSO, JWT, etc.).Troubleshoot production issues, improve availability, and drive root cause analysis.Standardize and automate routine tasks using Python scripting or Ansible.Support web-based deployments including APIs, UI apps, and their configurations.Collaborate on system design, capacity planning, and performance tuning.
Must-Have Skills
Solid hands-on experience with Kubernetes and Docker.Strong understanding of containerization, deployment strategies, and orchestration.Deep familiarity with observability stacks: Prometheus, Grafana, alerts, and metrics.Working knowledge of web deployment and web component architecture.Experience in setting up or working with authN/authZ mechanisms (OAuth, SSO)Good understanding of CI/CD pipelines (Jenkins, GitLab CI, etc.).Comfort with Linux and basic database operations (SQL)
Nice to Have
Proficiency in Python scripting for automation.Experience with Ansible or other config management tools.Exposure to infrastructure as code (Terraform, Helm).Familiarity with cloud-native concepts and services.
Who You Are
A systems thinker who understands both the development and operations side.Someone who thrives in a fast-paced, cross-functional team.Passionate about automation, stability, and continuous improvement.