Summary
✨ AI‑Generated
Lead SRE and DevOps operations for a large-scale, business-critical digital platform serving multiple markets. You will operate containerized infrastructure, build automated delivery pipelines, implement controlled release strategies, strengthen observability and security, and partner with engineering teams on application deployment.
Highlights
Lead reliability and DevOps for a large-scale, business-critical digital platform, with strong ownership of automation, delivery, observability, security, and operational resilience.
Description
We are working with a fintech company supporting a large-scale, business-critical digital platform across multiple markets.
They are looking for an experienced SRE/ DevOps Lead to join their technology team.
Key ResponsibilitiesSupport and operate a containerized application platform within a distributed environment.Enhance software delivery processes by building and maintaining automated pipelines and deployment workflows across multiple environments.Implement modern release strategies to ensure stable and controlled application rollouts and recovery.Drive automation and continuous delivery practices to improve deployment efficiency and consistency.Maintain system observability, including monitoring, logging, and alerting capabilities to ensure platform health and early issue detection.Ensure platform security, including system hardening, patching, and vulnerability management.Work closely with development and engineering teams to support application onboarding and deployment.Respond to and manage production incidents, including investigation, root cause analysis, and follow-up improvements.Contribute to ongoing improvements in platform reliability, scalability, and operational standards.RequirementsDegree or Higher Diploma in Computer Science, Information Technology, or a related discipline.At least 5 years of experience in platform engineering, DevOps, or site reliability roles.Strong hands-on experience with enterprise Linux environments, particularly Red Hat-based platforms (mandatory).Solid experience with containerization and orchestration concepts in production environments.Experience in CI/CD practices and automation tools within modern software delivery environments.Familiarity with infrastructure automation and configuration management tools.Good understanding of networking concepts and distributed systems architecture.Experience with monitoring, logging, and system observability practices.Proficiency in scripting or programming for automation tasks.Exposure to system infrastructure components such as databases, operating systems, virtualization, and container technologies.Strong problem-solving skills with the ability to troubleshoot complex system issues.Understanding of security best practices and compliance requirements in enterprise environments.Relevant certifications and knowledge of industry standards are advantageous.Good communication skills and ability to collaborate across teams.Proficiency in English and Chinese (spoken and written).Willingness to support off-hours incident handling when required.If you believe you have the right skills, attitude and experience please click 'apply now' below and upload your resume.
Alternatively, for a confidential chat, please contact Kevin Ng by applying directly to email kng@captarpartners.com or reach out at +852 3901 8736.
We apologies that only shortlisted candidates will be contacted.