Summary
Join an experienced platform engineering team responsible for reliable, scalable, and highly available enterprise container platforms. You will administer cloud-native platforms and Kubernetes, build CI/CD pipelines, develop Infrastructure-as-Code and automation, implement observability, lead incident response and root-cause analysis, and improve operational reliability. Strong experience across platform engineering, automation, production support, and container technologies is expected, with opportunities to provide technical leadership and mentor junior engineers.
Highlights
Senior platform reliability opportunity focused on enterprise-scale cloud-native infrastructure, automation, observability, production reliability, technical leadership, and mentoring junior engineers.
Description
Job Title: Platform Reliability Engineer (PRE) - Tanza
Location: Kuala Lumpur, Malaysia
Job Summary:
We are seeking an experienced Platform Reliability Engineer (PRE) / Site Reliability Engineer (SRE) with strong expertise in VMware Tanzu, Pivotal Cloud Foundry (PCF), Kubernetes, cloud platforms, and DevOps automation.
The ideal candidate will be responsible for ensuring the reliability, scalability, performance, and availability of enterprise container platforms and cloud-native applications.
This role requires hands-on experience in platform engineering, infrastructure automation, CI/CD, monitoring, and production support.
Key Responsibilities
Manage, administer, and support VMware Tanzu and/or Pivotal Cloud Foundry (PCF/PAS/PKS) platforms.Deploy, maintain, and troubleshoot Kubernetes clusters and containerized applications.Ensure platform reliability, availability, scalability, and performance across production environments.Design, build, and maintain CI/CD pipelines using tools such as Jenkins, Bamboo, GitHub, Bitbucket, and Nexus.Develop Infrastructure-as-Code (IaC) solutions using Terraform, Ansible, Chef, or similar automation tools.Create automation scripts and utilities using Python, Go, Java, or other programming languages.Implement and maintain monitoring, logging, and observability solutions using Prometheus, Grafana, ELK, and related tools.Support cloud-based workloads across AWS, Azure, VMware, or hybrid cloud environments.Lead incident response activities, root cause analysis (RCA), and postmortem reviews to improve platform stability.Drive operational excellence through automation, proactive monitoring, and continuous reliability improvements.Manage Docker container lifecycle and support container platform operations.Collaborate with Infrastructure, Cloud, DevOps, Security, and Application teams to deliver platform enhancements and resolve complex issues.Mentor junior engineers and provide technical leadership for platform operations and reliability initiatives.Required Qualifications
5+ years of experience in Platform Engineering, Production Reliability Engineering (PRE), or Site Reliability Engineering (SRE).3-5+ years of hands-on experience with VMware Tanzu and/or Pivotal Cloud Foundry (PAS, PKS, PCF).4-6+ years of Kubernetes and container platform administration experience.4-6+ years supporting CI/CD pipelines and DevOps toolchains.3-5+ years of Infrastructure-as-Code and automation experience using Terraform, Ansible, Chef, or similar tools.3-5+ years of programming/scripting experience using Python, Go, Java, or equivalent languages.Strong experience with monitoring, observability, incident management, and production support.Hands-on experience with Docker and container lifecycle management.Strong knowledge of Git-based version control and branching strategies.Preferred Qualifications
Experience supporting mission-critical enterprise platforms in highly available environments.Certifications in Kubernetes, VMware Tanzu, AWS, Azure, or related cloud technologies.Experience leading small engineering teams and engaging with cross-functional stakeholders.Strong understanding of SRE principles, platform automation, and cloud-native architecture.