Description
This position is listed on behalf of a partner company, who manages all applications and next steps.
Our partner is looking for a Senior Site Reliability Engineer based in Canada.
This is an impactful opportunity for an experienced SRE to strengthen the reliability, scalability, security, and operational efficiency of modern cloud platforms.
You will take ownership of production systems across AWS and Kubernetes while establishing measurable reliability standards.
The role combines software engineering, infrastructure automation, observability, incident management, and resilience engineering.
You will work closely with software, platform, security, QA, and product teams to make systems safer and easier to operate at scale.
A major focus will be reducing operational toil through automation, improving deployment reliability, and turning production incidents into lasting improvements.
You will also contribute to responsible cloud cost management while ensuring performance, availability, and security remain strong.
This role is well suited to a technically strong engineer who enjoys solving complex problems and shaping engineering practices for mission-critical systems.
Accountabilities
Define, measure, and continuously improve service reliability through SLIs, SLOs, error budgets, availability targets, and capacity planning.Build and maintain comprehensive observability across infrastructure and applications, including metrics, logs, traces, dashboards, and actionable alerting using tools such as CloudWatch, Prometheus, Grafana, New Relic, ELK, or OpenSearch.Participate in production incident response, troubleshooting, escalation, root-cause analysis, blameless post-incident reviews, and corrective-action tracking.Improve operational excellence through production readiness reviews, runbooks, documentation, change management practices, and engineering standards that reduce risk and improve maintainability.Optimize infrastructure for reliability and performance while managing cloud consumption responsibly and supporting FinOps initiatives.Analyze system performance, resource utilization, latency, throughput, and growth trends to identify bottlenecks and implement scalable solutions proactively.Identify repetitive operational tasks and replace them with reliable automation using Python, Bash, Go, CI/CD tooling, and platform APIs.Design and improve CI/CD pipelines that enable safe, repeatable deployments through automated testing, validation, progressive delivery, rollback strategies, and deployment observability.Apply security and compliance controls across cloud environments, including least-privilege IAM, encryption, network security, secrets management, patching, vulnerability management, and auditability.Design, develop, review, and maintain reusable Terraform modules and infrastructure-as-code patterns for AWS environments.Operate and improve Kubernetes platforms, including EKS clusters, workloads, Helm deployments, autoscaling, upgrades, resource management, networking, storage, and workload resilience.Architect, operate, and optimize AWS services including EC2, S3, RDS, EKS, Lambda, VPC, IAM, Route 53, and related cloud services.Design and validate fault-tolerant architectures, backup strategies, recovery procedures, and disaster recovery capabilities, including reliability testing and failure exercises where appropriate.Partner with development teams to improve application operability, instrumentation, deployment patterns, reliability, and production readiness.Contribute to incident response programs, on-call practices, resilience testing, operational standards, and continuous improvement initiatives across engineering teams.
Requirements
5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, cloud infrastructure, or a closely related discipline, with significant production ownership.Strong hands-on experience designing and operating production workloads in AWS, including networking, IAM, compute, storage, databases, DNS, and managed Kubernetes.Advanced Terraform expertise, including reusable modules, remote state, dependency management, environment design, code reviews, and infrastructure lifecycle management.Strong experience with observability and production telemetry, using technologies such as Prometheus, Grafana, CloudWatch, New Relic, ELK, OpenSearch, or comparable platforms.Practical understanding of SRE principles including SLIs, SLOs, error budgets, capacity planning, fault tolerance, graceful degradation, and operational toil reduction.Deep Kubernetes knowledge covering EKS, Helm, workload scheduling, networking, storage, autoscaling, upgrades, troubleshooting, and production operations.Strong Linux systems expertise, with the ability to diagnose issues involving CPU, memory, disk, networking, processes, DNS, and application dependencies.Proficiency in Python, Bash, Go, or another general-purpose programming language used for operational tooling and automation.Experience troubleshooting complex production incidents and contributing to incident response, root-cause analysis, postmortems, and corrective actions.Experience designing or operating CI/CD systems such as GitHub Actions, Jenkins, GitLab CI, Argo CD, or comparable technologies.Working knowledge of cloud security practices, including IAM, encryption, secrets management, network segmentation, vulnerability management, and audit controls.Strong written and verbal communication skills, with the ability to collaborate effectively across technical and business teams.Ability to work independently, take ownership of production systems, and make sound technical decisions in complex environments.AWS certifications such as Solutions Architect Professional or DevOps Engineer Professional are desirable.Experience with GitOps, Argo CD or Flux, formal on-call programs, chaos engineering, resilience testing, service meshes, distributed systems, microservices, RDS, DynamoDB, PostgreSQL, multi-account AWS environments, cloud governance, FinOps, or compliance frameworks such as SOC 2, ISO 27001, or PCI DSS is a plus.
Benefits
Flexible remote or hybrid work options.Comprehensive health and wellness benefits, including medical, dental, and vision coverage with 100% premium coverage for you.Generous paid time off and paid holidays.MyShare Employee Ownership Program.401(k) plan with up to a 4% employer match, with full vesting from day one.Opportunities to work alongside industry leaders and experienced engineering professionals.Professional growth opportunities in modern cloud, reliability, and energy technology environments.Opportunity to contribute to advanced energy intelligence and infrastructure supporting a more modern, secure, and resilient grid.Occasional travel of 10% or less may be required.
How Jobgether Works
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements.
Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company.
The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer.
This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR).
You may exercise your rights (access, rectification, erasure, objection) at any time.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information.
These tools assist our recruitment team but do not replace human judgment.
Final hiring decisions are ultimately made by humans.
If you would like more information about how your data is processed, please contact us.