Site Reliability Engineer

Sotalentjobs — Canada · Posted ~7 hours ago

Mid Full-time Remote

Skills

cloud infrastructure system reliability automation containerized environments Infrastructure as Code CI/CD GitOps cloud Docker Kubernetes IaC

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join an engineering team focused on creating highly reliable and scalable cloud platforms. The role involves improving production systems, automating operational processes, managing containerized workloads, and implementing repeatable infrastructure solutions.

Highlights

Opportunity to build scalable cloud infrastructure, improve reliability of critical systems, automate operations, and work with modern infrastructure practices in a fully remote engineering environment.

Description

Site Reliability Engineer, Infrastructure Platforms 📍 Location: Canada (Remote) 🏢 Industry: IT Services and IT Consulting 💼 Work Setting: Remote Are you passionate about building highly reliable, scalable, and efficient cloud infrastructure that powers mission-critical applications? Join a team of talented engineers dedicated to enhancing system reliability, automating operations, and driving continuous improvement across large-scale production environments. What You'll Do Ensure the reliability, performance, and scalability of production systems and customer-facing services.Design and develop automation solutions that eliminate repetitive manual tasks and improve operational efficiency.Manage, maintain, and troubleshoot containerized environments, including deployments, upgrades, scaling, and performance optimization.Develop and maintain Infrastructure as Code (IaC) solutions to enable consistent, secure, and repeatable infrastructure delivery.Support CI/CD and GitOps practices to streamline software and infrastructure deployments.Participate in incident management, root cause analysis, and post-incident reviews to improve system resilience.Monitor system health using metrics, logs, traces, and service-level objectives to proactively identify and resolve issues.Create and maintain technical documentation, operational runbooks, and architectural guidelines.Collaborate with cross-functional teams to improve platform reliability, security, and operational excellence.What We're Looking For Experience supporting and maintaining highly available production environments with a strong focus on reliability engineering.Proven ability to design and build infrastructure automation, custom tooling, and engineering solutions from the ground up.Strong programming and debugging skills, with the ability to analyze code behavior, performance, and failure scenarios.Hands-on experience with Infrastructure as Code tools and modern container orchestration platforms.Experience working with major cloud platforms such as AWS, GCP, or similar cloud environments.Solid understanding of observability practices, including monitoring, logging, alerting, SLIs, and SLOs.Experience participating in on-call rotations and effectively handling production incidents under pressure.Strong problem-solving, troubleshooting, and analytical skills.Excellent written and verbal communication skills, with the ability to work independently in distributed teams.Demonstrated success leveraging automation and emerging technologies, including AI-driven solutions, to increase productivity and reduce operational overhead.A collaborative mindset, commitment to continuous learning, and passion for improving system reliability at scale.