Production Support Engineer

Epsilon Solutions Canada β€” Canada Β· Posted ~7 hours ago

Mid Full-time

Skills

Production support Monitoring and alerting Incident response Troubleshooting Automation Infrastructure management Capacity planning Release engineering Cloud infrastructure System reliability Servers Networks

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An organization is seeking a Production Support Engineer to maintain reliable and available systems. You will build monitoring and alerting capabilities, respond to incidents, automate repetitive operations, manage servers, networks, and cloud resources, plan capacity, support software releases, and collaborate closely with development and operations teams.

Highlights

Operations-focused engineering role centered on system reliability, availability, and efficient incident response. Responsibilities span monitoring, automation, infrastructure, capacity planning, release engineering, collaboration with development teams, and maintaining operational documentation.

Description

Role Name: Production Support Engineer Location: Toronto Job Description: Monitoring and Alerting:Implement and maintain monitoring systems to proactively identify potential issues and alert engineers to problems before they impact users.Incident Response:Respond to incidents and outages, diagnose problems, and implement solutions to minimize downtime and restore service.Automation:Automate repetitive tasks and processes to improve efficiency and reduce manual effort.Infrastructure Management:Manage and maintain the underlying infrastructure, including servers, networks, and cloud resources.Capacity Planning:Plan for future capacity needs to ensure systems can handle anticipated workloads.Release Engineering:Develop and maintain processes for deploying software updates and releases.Collaboration:Work closely with developers, operations teams, and other stakeholders to ensure system reliability and availability.Documentation:Maintain clear and concise documentation of systems, processes, and procedures.Continuous Improvement:Identify areas for improvement and implement changes to enhance system reliability and performance. Skills and Qualifications: ● Cloud Platform (OCP) experience in production support handling Prod incidents. ● Excellent knowledge of OCP with command line. ● Monitoring tools (Dynatrace) ● Operating System (Windows, Linux) ● Scripting (Shell Scripting, Python, Power Shell) ● Database (SQL database management, MongoDB) ● Container Services (Kubernetes) ● Disaster Recovery Planning and execution