Site Reliability Engineer (DevOps)

Ant International — Malaysia · Posted ~1 day ago

Mid Full-time

Skills

Site Reliability Engineering Public cloud infrastructure AWS Google Cloud Platform Cloud operations Security patching Incident management Infrastructure reliability Cloud infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary

A global technology organization is hiring an SRE professional to improve infrastructure stability, scalability, and security. The role focuses on cloud operations, monitoring, risk reduction, and collaboration across international engineering teams.

Highlights

Work with global teams on large-scale infrastructure reliability, cloud operations, security improvements, and strategic platform enhancements.

Description

Ant International is a leading global digital payment, digitisation and financial technology provider. Through collaboration across the private and public sectors, our unified techfin platform supports financial institutions and merchants of all sizes to achieve inclusive growth through a comprehensive range of cutting-edge digital payment and financial services solutions. Our Mission Make it easy to do business anywhere, bringing small and beautiful changes to the world Our Vision To be the most trusted and innovative digital partner to bring inclusive growth to all Job Description: Collaborate with global teams to complete the daily ops and alarm handling.Identify and implement solutions on stability, scalability and security of business infrastructure using frameworks and industry best practices.Drive and manage technical and solution architecture discussions between global teams and partners to ensure timely delivery that meet customer needs.Plan and execute roadmap for strategic infrastructure improvement incorporating initiatives that align with the company goals. Requirement: Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.Experienced in site reliability engineering.Extensive experience in performing O&M activities which includes security patching, version upgrade, alarm management and handling in public cloud especially Google Cloud or AWS services.Advance proficiency and understanding in the factors and scenarios that generate technology risks in public cloud infrastructure.Have the know how to manage and prevent these risks, and be able to design general technology risk solutions/systems/products, etc. through systematic abstraction.Excellent communication and interpersonal skills with very pro-active attitude in solving difficult problems.Fresh Graduates are welcomed to apply.