Platform & Site Reliability Engineer

Techdoquest — Canada · Posted ~2 hours ago

Skills

Kubernetes Docker CI/CD Prometheus Grafana Containerization Deployment strategies Authentication and authorization OAuth2 SSO JWT Python Ansible Production troubleshooting Root cause analysis Capacity planning Performance tuning

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a technical platform engineering team responsible for building and operating scalable containerized environments. You will manage Kubernetes and Docker workloads, create dependable CI/CD pipelines, establish strong observability with metrics and dashboards, automate operational tasks with Python or Ansible, and help application teams securely onboard APIs and web services. The role also involves production troubleshooting, root cause analysis, capacity planning, and performance optimization.

Highlights

Opportunity to own and improve a modern cloud-native platform, build reliable deployment automation, strengthen observability, and work closely with application teams on secure and scalable services. The role offers broad exposure to infrastructure, reliability, automation, and system performance.

Description

Key Responsibilities • Design, deploy, and manage scalable, containerized applications on Kubernetes clusters. • Build and maintain CI/CD pipelines to enable reliable and repeatable deployments. • Own and improve platform observability using tools like Prometheus, Grafana, and alerts. • Work closely with application teams to onboard services with proper authentication and authorization (OAuth2, SSO, JWT, etc.). • Troubleshoot production issues, improve availability, and drive root cause analysis. • Standardize and automate routine tasks using Python scripting or Ansible. • Support web-based deployments including APIs, UI apps, and their configurations. • Collaborate on system design, capacity planning, and performance tuning. Must-Have Skills • Solid hands-on experience with Kubernetes and Docker. • Strong understanding of containerization, deployment strategies, and orchestration. • Deep familiarity with observability stacks: Prometheus, Grafana, alerts, and metrics. • Working knowledge of web deployment and web component architecture. • Experience in setting up or working with authN/authZ mechanisms (OAuth, SSO) • Good understanding of CI/CD pipelines (Jenkins, GitLab CI, etc.). • Comfort with Linux and basic database operations (SQL) Nice to Have • Proficiency in Python scripting for automation. • Experience with Ansible or other config management tools. • Exposure to infrastructure as code (Terraform, Helm). • Familiarity with cloud-native concepts and services. Who You Are • A systems thinker who understands both the development and operations side. • Someone who thrives in a fast-paced, cross-functional team. • Passionate about automation, stability, and continuous improvement.