Summary
✨ AI‑Generated
Join an engineering team focused on creating highly reliable and scalable cloud platforms. The role involves improving production systems, automating operational processes, managing containerized workloads, and implementing repeatable infrastructure solutions.
Highlights
Opportunity to build scalable cloud infrastructure, improve reliability of critical systems, automate operations, and work with modern infrastructure practices in a fully remote engineering environment.
Description
Site Reliability Engineer, Infrastructure Platforms
📍 Location: Canada (Remote)
🏢 Industry: IT Services and IT Consulting
💼 Work Setting: Remote
Are you passionate about building highly reliable, scalable, and efficient cloud infrastructure that powers mission-critical applications? Join a team of talented engineers dedicated to enhancing system reliability, automating operations, and driving continuous improvement across large-scale production environments.
What You'll Do
Ensure the reliability, performance, and scalability of production systems and customer-facing services.Design and develop automation solutions that eliminate repetitive manual tasks and improve operational efficiency.Manage, maintain, and troubleshoot containerized environments, including deployments, upgrades, scaling, and performance optimization.Develop and maintain Infrastructure as Code (IaC) solutions to enable consistent, secure, and repeatable infrastructure delivery.Support CI/CD and GitOps practices to streamline software and infrastructure deployments.Participate in incident management, root cause analysis, and post-incident reviews to improve system resilience.Monitor system health using metrics, logs, traces, and service-level objectives to proactively identify and resolve issues.Create and maintain technical documentation, operational runbooks, and architectural guidelines.Collaborate with cross-functional teams to improve platform reliability, security, and operational excellence.What We're Looking For
Experience supporting and maintaining highly available production environments with a strong focus on reliability engineering.Proven ability to design and build infrastructure automation, custom tooling, and engineering solutions from the ground up.Strong programming and debugging skills, with the ability to analyze code behavior, performance, and failure scenarios.Hands-on experience with Infrastructure as Code tools and modern container orchestration platforms.Experience working with major cloud platforms such as AWS, GCP, or similar cloud environments.Solid understanding of observability practices, including monitoring, logging, alerting, SLIs, and SLOs.Experience participating in on-call rotations and effectively handling production incidents under pressure.Strong problem-solving, troubleshooting, and analytical skills.Excellent written and verbal communication skills, with the ability to work independently in distributed teams.Demonstrated success leveraging automation and emerging technologies, including AI-driven solutions, to increase productivity and reduce operational overhead.A collaborative mindset, commitment to continuous learning, and passion for improving system reliability at scale.