Site Reliability Engineer
Mantugroup โ Canada ยท Posted ~22 hours ago
๐ Log in to save this job, tailor your resume & track your apply process โ 7 days free, no card needed.
Log in to add to target listDescription
About the Role
Join an API Platform team responsible for operating and continuously improving a critical enterprise platform used by development teams across the organization.
The team follows an API-first approach and supports a diverse technology landscape spanning public cloud, hybrid cloud, and on-premises environments.
This is a hands-on SRE / Platform Engineering opportunity that combines software engineering and operations.
You will help keep the Kubernetes platform stable, secure, performant, and scalable while building automation, improving observability, supporting production incidents, and enhancing the overall developer experience.
The role is focused on the day-to-day operation and continuous improvement of the platform, rather than traditional application development.
Main Responsibilities
Operate and maintain a Kubernetes-based platform across public cloud, private cloud, and on-premises environments.Monitor platform health using observability tools, ensuring alerts are actionable and supported by accurate runbooks and documentation.Manage incidents end-to-end, including incident response, troubleshooting, root cause analysis, and post-incident reviews.Build automation, diagnostic tools, and performance tests to reduce manual effort and improve platform reliability.Support client onboarding, platform upgrades, infrastructure changes, and capacity management to ensure the platform scales with demand.Identify and implement continuous improvements across operations, automation, deployment, monitoring, and support processes.
Qualifications
5โ7 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or a similar infrastructure-focused role.Strong hands-on experience operating Kubernetes in production, ideally with Service Mesh exposure.Excellent Linux and command-line skills with the ability to troubleshoot complex systems from application to infrastructure layers.Experience with a major public cloud platform, preferably Azure or AWS.Familiarity with Grafana, Prometheus, Loki, and Tempo; Python or Java scripting experience is an asset.Knowledge of CI/CD and Infrastructure as Code, particularly Helm or Terraform, combined with strong analytical and problem-solving skills.
We have 91,923 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume โ in under a minute we'll analyze all 91,923 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume