Summary
β¨ AIβGenerated
A hands-on DevOps position responsible for managing cloud infrastructure, improving system reliability, and building effective monitoring and alerting solutions. The role involves modern cloud platforms and production operations.
Highlights
Hands-on role owning cloud infrastructure, reliability, observability, and production operations.
Description
Reports to: CEO/ HEAD OF ENGINEERING
Location: Kuala Lumpur (or remote with regular on-site)
Works closely with: CISO (independent function β you'll support their evidence and control requests, not own them)
Role SummaryOwn and run Senang.io's cloud infrastructure and production operations end-to-end, including completing the in-progress monitoring stack consolidation.
This is a hands-on execution role focused on infrastructure reliability and observability β not a security-policy or compliance-ownership role.
Key ResponsibilitiesOwn Azure infrastructure day-to-day: AKS/Kubernetes, MySQL, Azure Firewall Premium, Application Gateway WAF V2, Microsoft Defender, Azure MonitorFinish consolidating the monitoring stack (Kibana/ELK, Prometheus, Grafana, Uptime Kuma, Wazuh, SonarQube) into a single Grafana dashboard with tiered alerting; clear the stalled integration ticketsTake the SIEM from proof-of-concept to fully operationalManage FusionAuth, Power Automate flows, JIRA-tracked infra work, and PagerDuty on-call rotationExecute incident response (detect, contain, remediate, escalate per plan) β policy and post-incident reporting sit with the CISOProvide technical documentation and evidence (configs, logs, change records) when the CISO or an external auditor requests it β you're the source of truth on the systems, not the one interpreting regulatory requirementsRequired Experience4β6 years in a DevOps/infrastructure/SRE role, with real production ownership (not just project-based cloud work)Hands-on Azure experience is strongly preferred; solid AWS/GCP experience is acceptable if willing to convertHas taken at least one monitoring tool or SIEM from setup/POC through to production useComfortable owning a live production environment with genuine on-call responsibilityMust-Know Technical AreasKubernetes (AKS or equivalent) β deployments, scaling, troubleshooting at a working levelCloud networking and security groups/firewalls (Azure Firewall, WAF, or equivalent)At least one observability stack (ELK, Prometheus/Grafana, or similar) configured from scratch, not just dashboards inherited from someone elseInfrastructure-as-code basics (Terraform, ARM/Bicep, or equivalent)Identity/auth systems (FusionAuth, Auth0, Keycloak, or similar)Nice-to-Have (not required)Exposure to a regulated industry (banking, insurance, fintech) β helpful context, not a filterAny BNM RMiT / MAS TRM / ISO 27001 familiarity β bonus, but the CISO covers thisCKA or similar certification