Summary
✨ AI‑Generated
Work remotely as a Site Reliability Engineer helping operate and improve large-scale production environments. You will build reliable and highly available systems, automate operational processes, manage Kubernetes workloads, develop Infrastructure as Code, and continuously improve reliability while reducing operational toil.
Highlights
Remote SRE role focused on highly reliable, scalable, and automated production infrastructure. Responsibilities combine software engineering with operational excellence, Kubernetes, Infrastructure as Code, automation, and reduction of manual operational work.
Description
Site Reliability Engineer, Infrastructure Platforms
📍 Location: Canada (Remote)
🏢 Industry: IT Services and IT Consulting
💼 Work Setting: Remote
Are you passionate about building highly reliable, scalable, and automated infrastructure that powers mission-critical applications and services? We are seeking a Site Reliability Engineer (SRE) to help design, operate, and continuously improve large-scale production environments.
In this role, you will combine software engineering expertise with operational excellence to enhance system reliability, reduce operational toil, and drive infrastructure automation across cloud-native platforms.
Key Responsibilities
Design, build, and maintain reliable, scalable, and highly available production systems and services.Develop automation and tooling that eliminate manual processes and improve operational efficiency.Manage, deploy, and troubleshoot containerized workloads within Kubernetes environments.Build and maintain Infrastructure as Code (IaC) solutions to support consistent, repeatable, and secure infrastructure deployments.Implement and optimize CI/CD and GitOps workflows to enable safe, automated software delivery.Participate in on-call rotations, monitor production environments, respond to alerts, and resolve service incidents.Improve observability through metrics, logging, monitoring, alerting, and service-level objectives (SLOs).Lead and contribute to incident response activities, root cause analysis, post-incident reviews, and preventative improvements.Document architecture decisions, operational procedures, troubleshooting guides, and best practices.Collaborate with engineering teams to improve service reliability, performance, scalability, and operational readiness.Drive continuous improvement initiatives focused on resilience, automation, efficiency, and platform stability.Required Qualifications
Experience supporting and maintaining large-scale production systems with a focus on reliability, performance, and operational excellence.Strong software engineering background with the ability to read, analyze, debug, and troubleshoot application code.Experience designing and building infrastructure automation solutions rather than solely administering existing tools.Hands-on expertise with Infrastructure as Code technologies and cloud infrastructure management.Strong experience with Kubernetes and cloud-native technologies, including deployment, scaling, and operational management.Experience working with at least one major public cloud platform such as AWS, Google Cloud Platform (GCP), or Azure.Knowledge of CI/CD pipelines, GitOps methodologies, and automated deployment practices.Experience implementing observability solutions, including monitoring, logging, alerting, metrics, SLIs, and SLOs.Ability to diagnose and resolve complex production issues in high-pressure environments.Experience participating in incident response, root cause investigations, and operational reviews.Strong problem-solving, analytical, and troubleshooting skills.Excellent written and verbal communication skills, with the ability to work effectively in distributed and remote-first environments.Preferred Qualifications
Experience building custom infrastructure tools, automation frameworks, operators, controllers, or platform services.Strong programming skills in languages such as Go, Python, Ruby, Java, or similar.Experience developing and managing Kubernetes operators, controllers, or advanced platform automation.Knowledge of distributed systems architecture, scalability patterns, and resilience engineering principles.Experience with cloud security, networking, and platform engineering best practices.Familiarity with service reliability engineering concepts, error budgets, capacity planning, and performance optimization.Experience leveraging AI-enabled tools and automation to improve productivity, operational efficiency, and engineering workflows.Demonstrated ability to lead reliability initiatives and influence technical direction across teams.Experience mentoring engineers and establishing operational best practices.