Senior Site Reliability Engineer

Gemini Solutions India โ€” Canada ยท Posted ~1 day ago

Senior Full-time Remote

Skills

SRE DevOps Distributed systems Cloud infrastructure Automation Observability Cloud Distributed Systems Monitoring

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

A senior reliability engineering role responsible for designing resilient systems, improving availability, automating operations, and supporting large-scale distributed platforms.

Highlights

Remote opportunity focused on reliability engineering, scalable systems, automation, and technical leadership.

Description

Position: Senior SRE Engineer Job Location: Canada/Remote Job Type: Full Time Immediate Interview Role Overview We are looking for a Senior Site Reliability Engineer (SRE) with a strong platformownership mindset to drive reliability, scalability, and performance of mission-critical, distributed systems.This role sits at the intersection of software engineering, cloud infrastructure, and production operations, with a focus on building resilient systems, improving observability, automating operations, and driving reliability at scale.You will act as a technical lead for platform reliability, working closely with engineering and business stakeholders to ensure systems are highly available, performant, and continuously improving. Experience: 3+ years of experience in SRE, DevOps, or Production EngineeringExperience supporting large-scale distributed systemsExperience working in production-critical environments with high availability requirementsExposure to global systems and cross-team collaboration Key Responsibilities Platform Reliability & Ownership Own availability, performance, and scalability of production systemsDefine and implement SLIs, SLOs, and error budgetsDrive continuous improvements in system resilience and efficiencyIncident Management & Root Cause Analysis Lead end-to-end incident response and service restorationPerform deep root cause analysis across infrastructure, application, data, and network layersImplement long-term fixes and reduce recurrence through engineering improvementsObservability & Monitoring Design and enhance monitoring, logging, and alerting systemsDevelop actionable dashboards and improve alert qualityEnable proactive detection of system issuesAutomation & DevOps Practices Automate operational workflows to reduce manual effortBuild and maintain CI/CD pipelinesImplement Infrastructure as Code (IaC) for scalable infrastructure managementCloud & Distributed Systems Manage and optimize systems on modern cloud platformsTroubleshoot distributed systems across compute, storage, and network layersDiagnose latency, routing, and performance issues in globally distributed environmentsData & Workflow Reliability Troubleshoot data pipelines, job failures, and data inconsistenciesPerform data validation and analysisEnsure reliability across data dependencies and workflowsNetworking & Traffic Management Diagnose issues related to DNS, HTTP/S, proxies, and load balancingWork with CDN and edge delivery platforms (e.g., Akamai or similar) to optimize traffic routing and performanceStakeholder Collaboration Act as a liaison between engineering teams and business stakeholdersCommunicate system status, incidents, and risks with clarity and contextPartner with cross-functional teams to drive reliability improvementsAI-Driven Reliability (Emerging Focus) Apply AI/ML-driven techniques for anomaly detection, alert optimization, andpredictive issue identificationLeverage intelligent automation to improve incident response and operationalefficiencyCore Expectations Demonstrates strong ownership of production systems and outcomesIndependently drives incident resolution and follow-throughApplies structured, analytical thinking to complex technical problemsCommunicates effectively in high-impact, production-critical scenariosFocuses on long-term reliability and scalability improvements Technical Skills: Programming & Automation Strong experience in Python for automation and toolingProficiency in shell scripting (Bash)Experience with API-driven and event-driven automationCloud & Infrastructure Hands-on experience with AWS, Azure, or GCPStrong understanding of cloud architecture, networking, and security fundamentalsInfrastructure as Code using Terraform, CloudFormation, or AnsibleDevOps & CI/CD Experience with Jenkins, GitLab CI, or similar toolsStrong understanding of build, release, and deployment pipelinesObservability Experience with Datadog, Splunk, Prometheus, or GrafanaStrong logging, monitoring, and alerting practicesFamiliarity with incident management tools (e.g., PagerDuty)Data & Databases Strong SQL skills for troubleshooting and validationUnderstanding of data pipelines and system dependenciesSystems & Platform Strong Linux fundamentalsExperience with Docker and containerized environmentsExposure to Kubernetes and web servers (e.g., Nginx)Orchestration Experience with Airflow, Autosys, or similar scheduling toolsNetworking & CDN Strong understanding of DNS, HTTP/S, proxies, and load balancingExperience with CDN and edge delivery platforms (e.g., Akamai or similar)