SRE Automation Engineer (Site Reliability Engineering)

Horizontal Talent โ€” Malaysia ยท Posted ~10 hours ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

About The Role We are looking for an SRE Automation Engineer to join HAVI's Supply Chain Technology team. You will be responsible for building automation solutions that improve system reliability, reduce operational toil and increase platform resilience. The role focuses heavily on automation, self-healing capabilities, incident response and improving deployment reliability across enterprise platforms. You will work closely with SRE, Platform Engineering and Application Development teams in a global environment. What You'll Do Design and implement automation to eliminate repetitive manual operational tasksBuild and maintain auto-remediation and self-healing mechanismsDevelop automation for incident response and runbook executionImprove CI/CD pipeline reliability and deployment safeguardsIntegrate observability tooling programmatically to improve detection and recoveryStandardise automation frameworks across platformsReduce alert noise through automated correlation and responseWork with SRE and Platform Engineering teams to embed resilience and reliability patterns into system design What We're Looking For 3+ years of experience in Automation Engineering, DevOps or SREStrong programming/scripting experience with Python, Go, Bash or PowerShellExperience building automation frameworks and operational toolingKnowledge of Infrastructure as Code, such as Terraform, ARM or CloudFormationExperience with CI/CD and deployment automationUnderstanding of observability tools and telemetry integrationExperience designing auto-remediation workflowsCloud platform experience, with Azure preferredStrong troubleshooting, analytical and systems-thinking skills Why Consider This Role? You'll be joining a global SRE environment where your work will directly contribute to automation, resilience and reliability engineering across enterprise technology platforms.