Site Reliability Engineer

Tjx Canada Winners Marshalls Homesense — Canada · Posted ~3 hours ago

Skills

Site reliability engineering AIOps Generative AI Automation Observability Incident response Self-healing systems

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Help build intelligent enterprise operations as a Site Reliability Engineer. You will use AIOps, generative AI, automation, observability, and self-healing technologies to increase platform reliability, accelerate incident response, and reduce manual operational work.

Highlights

Opportunity to shape intelligent IT operations through AIOps, GenAI, automation, observability, and self-healing technologies, with an emphasis on improving reliability and incident response.

Description

TJX Companies At TJX Canada, every day brings new opportunities for growth, exploration, and achievement. You’ll be part of our vibrant team that embraces diversity, fosters collaboration, and prioritizes your development. Whether you’re working in our Distribution Centers, Corporate Offices, or Retail Stores—WINNERS, HomeSense, and Marshalls, you’ll find abundant opportunities to learn, thrive, and make an impact. Come join our TJX family—a Fortune 100 company and the world’s leading off-price retailer. Here at TJX Canada, we are an equal opportunity employer committed to the inclusion and accommodation of all individuals. Job Description We’re looking for a Site Reliability Engineer to help shape the future of intelligent operations by leveraging AIOps, GenAI, automation, and self-healing technologies across enterprise platforms. In this role, you’ll drive initiatives that enhance system reliability, improve observability, and accelerate incident response while reducing manual effort through innovative automation solutions. You'll have the opportunity to influence modern engineering practices, work with cutting-edge AI technologies, and make a measurable impact on the scalability and performance of critical business applications. Join a collaborative team where innovation, continuous improvement, and operational excellence are at the heart of everything we do. Why Work With Us? We value integrity, respect, and teamwork, encouraging a unique and inclusive culture.Enjoy Associate discounts at our stores, available to you and eligible family members.Immediate access to our Group Benefits package, including a Health Care Spending Account, Retirement Savings Program, Associate & Family Assistance Program, and various well-being resourcesA competitive vacation package, paired with a Vacation Trade Program that allows you to opt in for an extra week. Comprehensive training and development resources designed to help you learn, grow, and succeed.Exciting career paths with growth opportunities and tuition reimbursement to support your career progression. What You’ll Do Implement SRE best practices across production support, incident management, problem management, change management, release, and deployment processes.Build and maintain automation scripts, runbooks, self-healing workflows, and operational tools to reduce manual effort and improve MTTR.Support AI-driven operational capabilities such as incident summarization, alert enrichment, anomaly detection, log analysis, event correlation, and root cause recommendations.Configure and enhance observability across logs, metrics, traces, dashboards, alerts, and synthetic monitoring.Work with monitoring/APM tools such as Splunk, AppDynamics, Dynatrace, Datadog, New Relic, Azure Monitor, or similar tools.Support implementation of SLIs, SLOs, SLAs, service health metrics, reliability KPIs, and operational dashboards.Integrate monitoring, ITSM, CI/CD, cloud, and automation platforms using APIs, scripts, and workflows.Participate in incident response, troubleshooting, RCA, postmortems, and corrective/preventive action planning.Drive shift-left reliability by embedding monitoring, alerting, automation, and operational readiness into SDLC and CI/CD pipelines.Support reliable application and platform operations in Azure Cloud.Create and maintain documentation, runbooks, knowledge articles, and automation playbooks.Ensure AI-enabled operations follow enterprise security, compliance, data privacy, and responsible AI guidelines. About You 6+ years of experience in SRE, DevOps, Production Support Engineering, Cloud Operations, Automation Engineering, or related roles.Hands-on experience in application support, incident management, problem management, change/release management, deployment support, monitoring, and documentation.Hands-on experience with observability/APM tools such as Splunk, AppDynamics, Dynatrace, Datadog, New Relic, Azure Monitor, or similar tools.Experience with SLIs, SLOs, SLAs, operational KPIs, service health dashboards, and reliability metrics.Strong coding/scripting experience in one or more languages such as Python, Java, PowerShell, Go, or Shell scripting.Hands-on experience with automation and DevOps tools such as Azure DevOps, GitHub Actions, Jenkins, Ansible, Terraform, Power Automate, Rundeck, or similar tools.Knowledge of AIOps, Gen AI, or AI-enabled operations use cases such as anomaly detection, log analysis, ticket classification, event correlation, and incident summarization.Experience integrating enterprise tools using APIs, webhooks, scripts, and automation workflows.Hands-on experience supporting or implementing solutions in Azure Cloud.Understanding of distributed systems, APIs, microservices, cloud-native applications, and reliability engineering principles.Strong troubleshooting, analytical, communication, and stakeholder management skills.Ability to work across application, infrastructure, cloud, operations, security, and leadership teams. Preferred Qualifications / Good To Have Experience with Agentic AI, AI agents, or Gen AI-based automation for IT operations.Experience with Azure OpenAI, Microsoft Copilot Studio, LangChain, Semantic Kernel, vector databases, or RAG-based solutions.Experience with ITSM/incident response tools such as ServiceNow, Jira Service Management, PagerDuty, xMatters, Opsgenie, or similar platforms.Experience with Docker, Kubernetes, OpenShift, CI/CD pipelines, Infrastructure as Code, GitOps, or DevSecOps.Knowledge of machine learning concepts such as anomaly detection, classification, clustering, and time-series analysis.Understanding of Responsible AI, prompt engineering, model governance, data privacy, and security controls.Experience in the Retail domain.Relevant certifications in Azure, DevOps, SRE, AI, or ITIL. Posting Details Posting End Date: September 15, 2026 If you’re ready to bring your energy and passion, we’d love to hear from you. Join us and be part of a place where every day is a chance to make a difference. Additional Information: Candidates aged 18 and over will be required to undergo a criminal record check as part of the hiring process. This job posting is for an existing position vacancy within our organization. TJX Canada uses artificial intelligence (AI) to assist in screening and assessing applicants for this position. Internal TJX Associates must submit their applications via the Jobs Hub in Workday. Direct applications to this job posting will not be accepted. Address 60 Standish Court Location: CAN Home Office Mississauga ON Salary Range: $ 87,780.00-$122,892.00 /year *This represents the expected hiring range and may not represent the full pay range for the position. The salary offered may be higher than the posted range depending on several factors such as relevant skills, qualifications, and experience.