Senior Site Reliability Engineer

Itc Infotech — Canada · Posted ~1 hour ago

Senior Full-time

Skills

DevOps AWS PostgreSQL Snowflake Airflow cloud infrastructure backup and recovery

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior reliability engineering position responsible for maintaining and scaling cloud-based data environments. The role focuses on infrastructure operations, automation, monitoring, resilience, and performance optimization.

Highlights

Work on scalable cloud infrastructure, reliability engineering, and data platform operations using modern cloud and database technologies.

Description

About Us: ITC Infotech is a leading global technology services and solutions provider, led by Business and Technology Consulting. ITC Infotech provides business-friendly solutions to help clients succeed and be future-ready, by seamlessly bringing together digital expertise, strong industry specific alliances and the unique ability to leverage deep domain expertise from ITC Group businesses. We provide technology solutions and services to enterprises across industries such as Banking & Financial Services, Healthcare, Manufacturing, Consumer Goods, Travel and Hospitality, through a combination of traditional and newer business models, as a long-term sustainable partner. Role Summary We’re looking for an experienced Sr DevOps and Site Reliability Engineer to maintain and scale AWS, PostgreSQL (ODS and transactional), Airflow and Snowflake environments. Should also be comfortable with replication, failover mechanisms and backup/recovery verifications mainly in the Snowflake space. Understanding various models and curated data sets that power Retail Planning analytics and decision-making (e.g., demand planning, assortment, allocation, replenishment, merchandise/financial planning, and performance reporting). You’ll partner closely with Planning stakeholders, Analytics/BI, and upstream source teams to ensure high-quality, governed, and reliable data products across the Retail Planning data ecosystem. Key Responsibilities DevOps & SRE Role Design, implement, and maintain scalable cloud infrastructure on AWS.Build and manage Infrastructure as Code (IaC) using Terraform.Develop and maintain CI/CD pipelines using GitLab CI/CD.Implement GitOps deployment strategies using ArgoCD.Automate operational processes and workflow integrations using Python.Manage containerized environments using Kubernetes, Helm, and Rancher.Establish infrastructure standards, reusable modules, and deployment best practices.Collaborate with development, architecture, and security teams to improve software delivery efficiency.Ensure platform availability, performance, scalability, and reliability.Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and SLAs.Implement observability and monitoring strategies using Datadog and Splunk.Lead incident management, root cause analysis, and post-incident reviews.Drive automation initiatives to reduce operational overhead and improve system resilience.Conduct capacity planning, performance optimization, and reliability assessments.Participate in on-call rotations and support critical production environments. Stakeholder Collaboration Work closely with Retail Planning, Merchandising, and Supply Chain stakeholders to evaluate upcoming implementations and relevant infrastructure impacts.Participate and contribute to BI/Analytics teams to automate recurring data preparation, KPI calculations, and reporting datasets.Participate and contribute in incident triage and root-cause analysis for platform related pipeline/data issues, drive prevention through durable fixes.Required Skills & Qualifications 8+ years of experience in DevOps, Cloud Engineering, Platform Engineering, or SRE roles.Strong hands-on experience with AWS cloud services.Expertise in GitLab CI/CD pipeline development and deployment automation.Extensive experience with Terraform for infrastructure provisioning and management.Strong Python scripting and automation skills.Deep understanding of Kubernetes architecture, administration, and troubleshooting.Experience managing Kubernetes clusters using Rancher.Hands-on experience with Helm chart development and deployment.Experience implementing GitOps workflows with ArgoCD.Strong experience with monitoring, logging, and observability platforms:Strong Linux administration and troubleshooting skills.Experience with networking concepts including DNS, Load Balancers, VPCs, Security Groups, and IAM. Retail Planning domain exposure (preferred) Familiarity with retail planning concepts and data: forecasting, inventory, allocation, replenishment, assortment, pricing/promo, sell-through, WOS/DOH, in-season vs. pre-season planning, etc.Exposure to common retail data sources like sales transactions, inventory snapshots, product hierarchies, store/channel attributes, supplier/lead times, and order flows.Education Qualification Bachelor / Master in Computer Science, Information Technology, Engineering, or related field (equivalent experience acceptable). ITC Infotech is an Equal Opportunity Employer. We believe that no one should be discriminated against because of their differences, such as age, disability, ethnicity, gender, gender identity and expression, religion, or sexual orientation. All employment decisions shall be made without regard to age, race, creed, color, religion, sex, national origin, ancestry, disability status, veteran status, sexual orientation, gender identity or expression, genetic information, marital status, citizenship status or any other basis as protected by federal, state, or local law. ITC infotech is committed to providing veteran employment opportunities to our service men and women.