Site Reliability Engineer
Capgemini — Canada · Posted ~21 hours ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d like, where you’ll be supported and inspired by a collaborative community of colleagues around the world, and where you’ll be able to reimagine what’s possible.
Join us and help the world’s leading organizations unlock the value of technology and build a more sustainable, more inclusive world.
Job Description
Application Support (Windows, UNIX/Linux, OpenShift, and PostgreSQL)
This role is part of the Production Support and Reliability Engineering team, responsible for ensuring the stability, availability, and performance of enterprise applications running across Windows, UNIX/Linux, OpenShift, and PostgreSQL environments.
Key Responsibilities:
Provide Level 2/Level 3 application and infrastructure support for enterprise applications hosted on Windows, UNIX/Linux, OpenShift, and PostgreSQL platforms.
Monitor application health, server performance, database availability, and OpenShift container workloads to ensure optimal system reliability and uptime.
Troubleshoot and resolve production incidents across operating systems, middleware, containers, and databases.
Perform root cause analysis (RCA) and implement preventative measures to reduce recurring incidents.
Support application deployments, environment management, and release activities across multiple environments.
Manage and support OpenShift platform operations, including container deployments, pod troubleshooting, scaling, and monitoring.
Monitor and support PostgreSQL databases, including performance tuning, query analysis, backup validation, and connectivity troubleshooting.
Collaborate with development, infrastructure, database, and cloud teams to drive operational excellence.
Develop and maintain automation scripts, operational runbooks, dashboards, and support documentation.
Participate in on-call support rotations and major incident management activities.
Identify and implement opportunities for process improvement, automation, and operational efficiency.
Required Skills & Experience:
Operating Systems:
Strong hands-on experience supporting and troubleshooting:
Windows Server
UNIX/Linux environments
Container & Platform Technologies:
Experience with OpenShift Container Platform (OCP), including:
Application deployment and support
Pod and container troubleshooting
Resource monitoring and performance analysis
OpenShift administration fundamentals
Database Technologies:
Strong experience with PostgreSQL, including:
Database monitoring and support
SQL query analysis and troubleshooting
Performance tuning and optimization
Backup and recovery validation
Database connectivity troubleshooting
Reliability & Support:
Experience in Application Support, Production Support, SRE, or Platform Operations roles.
Strong understanding of incident, problem, and change management processes.
Experience supporting mission-critical applications in large enterprise environments.
Ability to analyze logs, alerts, metrics, and system trends to rapidly resolve issues and improve service reliability.
Knowledge of monitoring and observability tools for proactive system management.
Strong analytical, troubleshooting, and communication skills.
Nice to Have:
Experience in banking, financial services, or other highly regulated industries.
Exposure to automation and scripting using PowerShell, Bash, Python, or Shell Scripting.
Experience with CI/CD pipelines and DevOps practices.
Knowledge of cloud-native technologies and hybrid cloud environments.
Experience with monitoring tools such as Splunk, Dynatrace, AppDynamics, Grafana, Prometheus, or similar platforms.
Understanding of ITIL processes, SLOs, SLIs, and reliability engineering principles.
Preferred Qualifications:
Bachelor's degree in Computer Science, Engineering, Information Technology, or related discipline.
Relevant certifications in OpenShift, Linux, PostgreSQL, Cloud, or DevOps technologies are considered an asset.
We have 83,775 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 83,775 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume