Site Reliability Engineer

Slbglobal — Malaysia · Posted ~2 hours ago

Mid Full-time

Skills

cloud infrastructure Kubernetes monitoring logging CI/CD Azure DevOps security automation

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A site reliability engineering role focused on cloud platforms, observability, container systems, CI/CD automation, and secure production operations.

Highlights

Design reliable cloud infrastructure, monitoring systems, and delivery pipelines while improving availability and operational quality.

Description

Responsibilities: Cloud-based infrastructure · Teamwork to design a flexible cloud infrastructure with high availability and reliability. · Design containerization, Kubernetes and service mesh implementation plans. Full stack monitoring and logging system settings · Design/develop a full-stack monitoring system for good monitoring of development and production scenarios within the entire cloud resource. · Design/develop a centralized log collection system that collects logs and establishes clear traceability with the debugging system so that the team can quickly determine the cause of the incident. CI/CD workflow/policy/pipeline setup and maintenance · Team design artifact management and promotion strategy · Work closely with QA team to define quality standards for product delivery cycle · Set up full CI/CD workflows and pipelines on Azure DevOps to ensure continuous delivery of products with good quality and security/legal compliance. Security Compliance · Build security scans (SAST, DAST, licenses, 3rd party licenses…) from the early stages of the product delivery lifecycle · ü Visualize security scan reports and set up remediation tracking and knowledge base Required Knowledge & Technical Skills: Bachelor degree and above, 3~5 years working experience in large factories. · Cloud platform and technology (Microsoft Azure preferred). Not only understand the basic concepts, but also have hands-on and architect design experience, including but not limited to APIM, load balancers, virtual machines, Kubernetes, serverless services, microservices, monitoring, and log collection. · There is a large-scale microservice architecture monitoring, and Kubernetes operation and maintenance experience is preferred · Development/operation experience of one or more relational databases or no sql databases, MongoDB, Cassandra, Elastic Search operation and maintenance experience is preferred · Message bus usage/operation experience, RabbitMQ, kafka operation and maintenance experience is preferred Soft Skills: · Fast Learner · Proactive & self-driven . Good Communication and teamwork Special skills as plus: · Large and highly distributed system architect design considering a system holding hundreds of or even thousands of applications. · Hands on development/operation/administration experience of Cloud/Docker/Kubernetes/Service Mesh · web and cloud security-related aspects · Agile process