Senior Cloud DevOps Engineer

Dadaconsultants Pte Ltd — Malaysia · Posted ~4 hours ago

Senior Full-time

Skills

DevOps cloud infrastructure CI/CD AI infrastructure high-performance computing datacenter operations automation Cloud AI/HPC

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A global technology company is looking for a Senior Cloud DevOps Engineer to design and maintain mission-critical cloud platforms. The position focuses on scalable infrastructure, automation, AI workloads, and operational excellence.

Highlights

Senior infrastructure role working on large-scale cloud platforms, advanced computing systems, reliability, and continuous delivery.

Description

Cloud Senior DevOps Engineer Location: Penang, MY About the Role Our client is a world-leading technology company operating at the intersection of high-performance computing and artificial intelligence, with a global footprint spanning multiple continents and a diversified energy portfolio exceeding 3 GW. They design and manufacture their own computing hardware, operate large-scale datacenters, and provide advanced AI cloud services to enterprise customers with demanding workloads. Headquartered in Singapore, the company is scaling rapidly and offers a rare opportunity to work on both Bitcoin mining infrastructure and cutting-edge AI/HPC systems at global scale. This role is central to the reliability, security, and continuous delivery of mission-critical platforms — making it an exceptional opportunity for a senior infrastructure engineer looking to operate at the frontier of cloud-native and AI infrastructure. Key Responsibilities Design, implement, and maintain end-to-end CI/CD and MLOps pipelines for software applications and machine learning models, automating build, test, deployment, and rollback processes to ensure smooth transitions from development to production Act as the technical lead during complex system incidents — driving rapid troubleshooting, thorough root cause analysis, and the implementation of preventative remediation plans Establish and enforce robust security and compliance standards, including release workflow governance, Zero Trust access controls, secrets management, and adherence to frameworks such as SOC2 and ISO27001 Collaborate closely with R&D, Data Science, Security, and Business teams to eliminate workflow bottlenecks and drive engineering efficiency through Internal Developer Platform (IDP) and Platform Engineering initiatives Architect and continuously improve comprehensive observability systems — including monitoring, logging, and alerting stacks (e.g., Prometheus, Grafana, ELK/EFK) — to maintain deep visibility into system health, application performance, and AI model metrics Champion Infrastructure as Code (IaC) practices using tools such as Terraform, Ansible, and Helm to achieve fully automated, reproducible, and auditable infrastructure provisioning across multi-cloud environments Own high-availability architecture design in production, including disaster recovery strategies, self-healing mechanisms, capacity planning, and performance tuning to meet stringent business SLAs Build, optimise, and scale cloud-native infrastructure using Kubernetes and Docker, including provisioning and managing GPU clusters to support high-performance AI workloads and model inferencing Requirements Bachelor's degree or above in Computer Science, Engineering, or a related technical discipline, with at least 5 years of hands-on experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles Solid domain knowledge across CI/CD methodologies, Infrastructure as Code, observability paradigms, and SRE principles, combined with strong problem-solving ability and cross-functional communication skills Strong scripting and programming proficiency in at least one major language such as Go, Python, or Shell, with a clear focus on automation and internal tooling development Demonstrated experience designing and managing infrastructure on major public or hybrid cloud platforms (e.g., AWS, GCP, Azure, Alibaba Cloud), including multi-cloud strategies Deep hands-on expertise with Docker and Kubernetes orchestration, including cluster management and production-level best practices, alongside expert-level Linux and core networking knowledge (TCP/IP, DNS, HTTP, load balancing, VPCs) Experience with MLOps practices and model serving frameworks (e.g., vLLM, TGI, Triton Inference Server), and managing GPU clusters for AI/ML workloads (is a bonus) Prior experience in large-scale distributed or high-concurrency environments (e.g., fintech, real-time processing, or AI platforms), and/or experience building Internal Developer Platforms (IDP) (is a bonus) Experience in a technical lead or mentoring capacity within a DevOps or SRE team (is a bonus)