Staff Site Reliability Engineer

Jobgether — Canada · Posted ~3 hours ago

Senior Full-time

Skills

Cloud architecture Distributed systems Infrastructure engineering DevOps Observability Cloud

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking a senior infrastructure specialist to design, operate, and improve scalable platforms while collaborating across engineering and security teams.

Highlights

Opportunity to own critical infrastructure systems, influence engineering practices, and solve large-scale reliability challenges.

Description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Infrastructure Engineer based in Canada. This role offers the opportunity to shape the infrastructure and platforms that support large-scale, customer-facing technology products. You will take end-to-end ownership of critical infrastructure systems, from technical design and development through deployment, operation, and continuous improvement. The position combines cloud architecture, distributed systems, reliability engineering, observability, performance optimization, and developer tooling. You will partner closely with product engineering, security, and DevOps teams to build secure, resilient, scalable, and cost-efficient systems. As a senior technical contributor, you will also influence engineering practices, mentor colleagues, and help resolve complex infrastructure challenges. This is a strong opportunity for a systems-minded engineer who thrives with significant autonomy in a globally distributed, fully remote environment. Accountabilities Take full ownership of a core infrastructure product or subsystem end to end, including design, development, deployment, operation, and production performance.Define project goals and success metrics, align technical work with organizational objectives, and proactively identify and mitigate risks.Translate product requirements and technical specifications into practical designs that address critical edge cases without unnecessary complexity.Build secure, reliable, resilient, high-performing, and cost-efficient infrastructure for diverse applications and workloads.Design, develop, and deploy production software and developer-facing tools that improve engineering workflows and reduce operational toil.Manage infrastructure through code and configuration, primarily using Terraform and established architectural patterns.Partner with product engineering teams to design services for scale and resolve ambiguous technical requirements with stakeholders.Participate in incident response and apply systematic debugging techniques to diagnose infrastructure and service issues.Develop and improve monitoring and observability practices, using operational data to identify stability, performance, and reliability improvements.Apply a security-focused mindset across infrastructure development, implementation, and peer reviews by proactively identifying potential vulnerabilities.Serve as a technical resource for complex infrastructure challenges and mentor engineers through code reviews, pairing, and design feedback.Drive collaboration across engineering and other stakeholder groups, facilitating discussions around technical decisions, processes, and infrastructure strategy. Requirements 6–10 years of experience in infrastructure, platform, or backend engineering, primarily within cloud-based environments; AWS experience is preferred.Proven experience owning significant infrastructure products or subsystems through their full lifecycle, including design, implementation, deployment, and production operations.T-shaped technical expertise, with deep specialization in one or two areas and sufficient breadth to navigate and contribute across wider systems with limited guidance.Strong hands-on experience managing infrastructure through code and configuration using Terraform or an equivalent technology.Deep understanding of cloud infrastructure fundamentals, including networking, load balancing, containerization, Kubernetes/EKS, and distributed systems.Strong programming skills in Go, Python, or a comparable language, with the ability to develop production-ready software.Hands-on production experience operating Redis or ElastiCache, including cluster and shard management, failover behavior, memory eviction policies, and scaling strategies.Experience with observability and monitoring technologies such as Prometheus, Grafana, OpenTelemetry, or similar tools.Strong understanding of performance tuning, incident management, reliability engineering, and production troubleshooting.Fluency in software engineering best practices, including source control, code reviews, comprehensive testing, edge-case handling, and safe deployment practices.High degree of ownership and autonomy, with demonstrated ability to make progress when requirements are ambiguous or not fully defined.Strong written and verbal English communication skills, including the ability to produce clear technical documentation, participate in effective code reviews, and communicate decisions across engineering teams.Ability to collaborate effectively with product engineering, security, DevOps, and other stakeholders in a distributed environment. Benefits Fully remote work arrangement.Opportunity to work on large-scale cloud infrastructure and systems supporting customer-facing technology products.Significant ownership over infrastructure products and subsystems from design through production operations.Exposure to cloud architecture, distributed systems, Kubernetes, Terraform, observability, reliability engineering, and developer tooling.Opportunity to work with technologies including AWS, Redis/ElastiCache, Prometheus, Grafana, OpenTelemetry, Go, and Python.Strong focus on engineering quality, security, scalability, and operational excellence.Opportunity to mentor engineers and influence technical standards and infrastructure practices.Collaboration with globally distributed engineering, product, security, and DevOps teams.High-autonomy environment suited to engineers who enjoy solving complex and ambiguous technical problems.US-based cash compensation range of $152,000–$205,000; compensation varies by hiring location and this range is not directly applicable to all locations.Visa sponsorship is not provided; candidates must be authorized to work from their home location.Specific India-based salary, healthcare, retirement, paid time off, and other benefits were not specified in the source description. How Jobgether Works We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.