Platform and Site Reliability Engineer

Yochana — Canada · Posted ~3 hours ago

Senior Full-time Hybrid

Skills

Platform Engineering SRE Kubernetes Container orchestration CI/CD Prometheus Grafana Production troubleshooting Python Ansible Cloud-native infrastructure OAuth2 SSO JWT

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking an experienced Platform and SRE Engineer to design and operate scalable Kubernetes-based platforms. You will build CI/CD pipelines, improve observability, onboard application services securely, troubleshoot production issues, conduct root-cause analysis, and automate operational tasks.

Highlights

Senior platform role focused on scalable cloud-native infrastructure, reliability, observability, automation, and production operations. Offers the opportunity to improve developer platforms and service availability.

Description

Position Name – Platform & SRE Engineer Type of hiring – Fulltime Location – Montreal, QC (Hybrid – 3 days Onsite) Job Description: Seeking a Platform & Site Reliability Engineer (SRE) with a strong background in modern cloud-native technologies and infrastructure automation. Looking for a highly skilled professional with 6+ Years of experience in platform engineering, container orchestration, and site reliability practices. Key Responsibilities Design, deploy, and manage scalable, containerized applications on Kubernetes clusters.Build and maintain CI/CD pipelines to enable reliable and repeatable deployments.Own and improve platform observability using tools like Prometheus, Grafana, and alerts.Work closely with application teams to onboard services with proper authentication and authorization (OAuth2, SSO, JWT, etc.).Troubleshoot production issues, improve availability, and drive root cause analysis.Standardize and automate routine tasks using Python scripting or Ansible.Support web-based deployments including APIs, UI apps, and their configurations.Collaborate on system design, capacity planning, and performance tuning. Must-Have Skills Solid hands-on experience with Kubernetes and Docker.Strong understanding of containerization, deployment strategies, and orchestration.Deep familiarity with observability stacks: Prometheus, Grafana, alerts, and metrics.Working knowledge of web deployment and web component architecture.Experience in setting up or working with authN/authZ mechanisms (OAuth, SSO)Good understanding of CI/CD pipelines (Jenkins, GitLab CI, etc.).Comfort with Linux and basic database operations (SQL) Nice to Have Proficiency in Python scripting for automation.Experience with Ansible or other config management tools.Exposure to infrastructure as code (Terraform, Helm).Familiarity with cloud-native concepts and services. Who You Are A systems thinker who understands both the development and operations side.Someone who thrives in a fast-paced, cross-functional team.Passionate about automation, stability, and continuous improvement.