Description
Cloud Senior DevOps Engineer
Location: Penang, MY
About the Role
Our client is a world-leading technology company operating at the intersection of high-performance computing and artificial intelligence, with a global footprint spanning multiple continents and a diversified energy portfolio exceeding 3 GW.
They design and manufacture their own computing hardware, operate large-scale datacenters, and provide advanced AI cloud services to enterprise customers with demanding workloads.
Headquartered in Singapore, the company is scaling rapidly and offers a rare opportunity to work on both Bitcoin mining infrastructure and cutting-edge AI/HPC systems at global scale.
This role is central to the reliability, security, and continuous delivery of mission-critical platforms — making it an exceptional opportunity for a senior infrastructure engineer looking to operate at the frontier of cloud-native and AI infrastructure.
Key Responsibilities
Design, implement, and maintain end-to-end CI/CD and MLOps pipelines for software applications and machine learning models, automating build, test, deployment, and rollback processes to ensure smooth transitions from development to production
Act as the technical lead during complex system incidents — driving rapid troubleshooting, thorough root cause analysis, and the implementation of preventative remediation plans
Establish and enforce robust security and compliance standards, including release workflow governance, Zero Trust access controls, secrets management, and adherence to frameworks such as SOC2 and ISO27001
Collaborate closely with R&D, Data Science, Security, and Business teams to eliminate workflow bottlenecks and drive engineering efficiency through Internal Developer Platform (IDP) and Platform Engineering initiatives
Architect and continuously improve comprehensive observability systems — including monitoring, logging, and alerting stacks (e.g., Prometheus, Grafana, ELK/EFK) — to maintain deep visibility into system health, application performance, and AI model metrics
Champion Infrastructure as Code (IaC) practices using tools such as Terraform, Ansible, and Helm to achieve fully automated, reproducible, and auditable infrastructure provisioning across multi-cloud environments
Own high-availability architecture design in production, including disaster recovery strategies, self-healing mechanisms, capacity planning, and performance tuning to meet stringent business SLAs
Build, optimise, and scale cloud-native infrastructure using Kubernetes and Docker, including provisioning and managing GPU clusters to support high-performance AI workloads and model inferencing
Requirements
Bachelor's degree or above in Computer Science, Engineering, or a related technical discipline, with at least 5 years of hands-on experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles
Solid domain knowledge across CI/CD methodologies, Infrastructure as Code, observability paradigms, and SRE principles, combined with strong problem-solving ability and cross-functional communication skills
Strong scripting and programming proficiency in at least one major language such as Go, Python, or Shell, with a clear focus on automation and internal tooling development
Demonstrated experience designing and managing infrastructure on major public or hybrid cloud platforms (e.g., AWS, GCP, Azure, Alibaba Cloud), including multi-cloud strategies
Deep hands-on expertise with Docker and Kubernetes orchestration, including cluster management and production-level best practices, alongside expert-level Linux and core networking knowledge (TCP/IP, DNS, HTTP, load balancing, VPCs)
Experience with MLOps practices and model serving frameworks (e.g., vLLM, TGI, Triton Inference Server), and managing GPU clusters for AI/ML workloads (is a bonus)
Prior experience in large-scale distributed or high-concurrency environments (e.g., fintech, real-time processing, or AI platforms), and/or experience building Internal Developer Platforms (IDP) (is a bonus)
Experience in a technical lead or mentoring capacity within a DevOps or SRE team (is a bonus)