Production Support Engineer / SRE

Lorventech — Canada · Posted ~1 day ago

Contract Onsite

Skills

monitoring and alerting incident response troubleshooting automation infrastructure management capacity planning release engineering cloud infrastructure monitoring SRE

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking a Production Support Engineer/SRE for a long-term onsite engagement. You will proactively monitor systems, respond to incidents and outages, automate repetitive operational work, manage infrastructure and cloud resources, plan capacity, and maintain reliable software release processes. The role is well suited to engineers focused on system reliability and operational excellence.

Highlights

Long-term project opportunity focused on production reliability, automation, infrastructure, and incident management. The role provides broad exposure to monitoring, cloud resources, capacity planning, and software release processes.

Description

Our client is looking Production Support Engineer SRE for Long term project in Mississauga, ON – Canada (Onsite) Below is the detail requirement. Job Title: Production Support Engineer SRE Location: Mississauga, ON – Canada (Onsite) Job Description: Monitoring and Alerting: Implement and maintain monitoring systems to proactively identify potential issues and alert engineers to problems before they impact users.Incident Response: Respond to incidents and outages, diagnose problems, and implement solutions to minimize downtime and restore service.Automation: Automate repetitive tasks and processes to improve efficiency and reduce manual effort.Infrastructure Management: Manage and maintain the underlying infrastructure, including servers, networks, and cloud resources.Capacity Planning: Plan for future capacity needs to ensure systems can handle anticipated workloads.Release Engineering: Develop and maintain processes for deploying software updates and releases.Collaboration: Work closely with developers, operations teams, and other stakeholders to ensure system reliability and availability.Documentation: Maintain clear and concise documentation of systems, processes, and procedures.Continuous Improvement: Identify areas for improvement and implement changes to enhance system reliability and performance. Skills and Qualifications: Cloud Platform (OCP)8+ Years’ experience in production support handling Prod incidents.Excellent knowledge of OCP and windows environement..Monitoring tools ( Dynatrace )Operating System (Windows, Linux)Scripting (Shell Scripting, Python, Power Shell)Database (SQL database management, MongoDb)Container Services (Kubernetes)Disaster Recovery Planning and execution