Senior Site Reliability Engineer

Bamboorose — Canada · Posted ~2 hours ago

Senior

Skills

SRE site reliability engineering system performance monitoring automation infrastructure system reliability availability optimization monitoring cloud infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Take ownership of reliability engineering across a large-scale technology environment. You’ll design and evolve SRE practices, monitor and optimize availability and performance, automate operational work, and collaborate across engineering teams to strengthen reliability and reduce operational toil.

Highlights

Senior SRE role combining hands-on engineering, automation, performance optimization, and organizational influence, with an opportunity to establish and mature reliability practices.

Description

About Bamboo Rose At Bamboo Rose, we’re building the world’s leading collaborative product development platform for global retail. Our technology helps retailers and brands bring great products to market faster, smarter, and more sustainably. We value curiosity, innovation, and solving real problems across global supply chains About the Role We’re looking for an experienced Sr. Site Reliability Engineer (SRE) to help design, implement, and scale reliability practices across our systems and infrastructure. This role blends hands-on engineering with strong ownership, collaboration, and influence, and plays a key part in establishing SRE ways of working as the organization continues to mature its reliability posture. What You’ll Do Design, implement, and evolve reliability practices aligned with SRE principles across the TotalPLM stack.Monitor, analyze, and optimize system performance, availability, and reliability.Build and improve automation to reduce toil and increase operational efficiency.Partner with Software Engineering and Customer Support teams to ensure reliable delivery and operation of services.Define, track and analyze SLI/SLO metrics.Participate in incident response, post-incident reviews, and root cause analysis, driving remediation.Contribute to defining and rolling out SRE standards, patterns, and best practices.Mentor and support junior engineers through knowledge sharing and hands-on guidance.Manage and drive technical projects related to reliability, automation, and infrastructure improvements. What You Bring 6+ years of progressive experience in SRE, DevOps, Platform, or equivalent roles.Demonstrated ownership of production systems and experience operating in on-call environments.Strong knowledge of DevOps practices, GitOps, CI/CD pipelines, and Infrastructure as Code (IaC) tools.Experience with automation tools such as Ansible, Puppet or equivalent, monitoring and observability tools like NewRelic, DataDog, Prometheus/Grafana; and incident management.Experience using AI to build skills, automations, AI assisted code generation and agentic software engineering, and the curiosity to continue to explore new ways of using AI.Strong project execution skills, with the ability to drive initiatives from idea through delivery.Clear communication skills and the ability to work effectively with both technical and non-technical stakeholders.A mindset of accountability, continuous improvement, and learning.Strong Scripting skills using Bash/Shell, Python, JS, etc. Why You’ll Love Working Here You’ll help shape and define SRE practices, not just follow them.You’ll work on meaningful reliability challenges that directly impact the business.You’ll have the opportunity to mentor others and grow your technical leadership skills.You’ll collaborate with engaged engineering partners who value reliability and operational excellence.You’ll be empowered to innovate, automate, and improve how systems are built and operated.A chance to work on technology that makes global retail more connected, sustainable, and resilient.Competitive compensation and benefits, with flexibility for remote work