Team Lead - Application DevSecOps & SRE
Dialog Group Berhad β Malaysia Β· Posted ~1 day ago
π Log in to save this job, tailor your resume & track your apply process β 7 days free, no card needed.
Log in to add to target listDescription
About the Role
We are hiring a Lead Engineer to head a small, high-ownership DevSecOps and SRE team of two to three engineers.The team owns the speed, reliability and security of our major external-facing applications β distributed microservices platforms running on managed Kubernetes in AWS.
The immediate focus is one significant system, and as the team matures and capacity allows, that scope is expected to extend to further business-critical applications.You are not a shared services function stretched thin across a sprawling portfolio.
You are the product-focused infrastructure owner for a small number of systems that genuinely matter, operating on Site Reliability Engineering principles.This is a genuine lead position.
We are looking for the seniority and judgement to set technical direction, to keep a small team focused and effective, and to stay composed when a production system is misbehaving and several people want an answer at once.You will work in close partnership with our central infrastructure team and our application engineering team.
Neither is a supervisor handing down instructions β technical direction is reached collaboratively, and your reasoning carries real weight in it.
Scope of Ownership
Application Environments
Full ownership of the Dev, UAT, Staging and Production environments for the applications in your teamβs scope.
Writing and maintaining the application-specific Infrastructure-as-Code covering their resources β Kubernetes clusters, databases, queues and caches.
Kubernetes Cluster Ownership
Cluster lifecycle management: creation, version upgrades, security patching, and scaling of both control plane and worker nodes.
Cluster cost optimisation through node autoscaling, Reserved and Spot instance strategy, and pod scheduling efficiency.
Internal cluster networking β CNI configuration, service mesh, and ingress/egress controllers, operating within the address space allocated by the central infrastructure team.
Continuous Integration & Delivery
Design, build and maintenance of end-to-end CI/CD pipelines for the applications in scope.
Implementation of advanced zero-downtime deployment strategies β blue/green and canary releases β across the microservices estate.
Management of application configuration and feature flags in production.
Site Reliability Engineering
Implementation and tuning of the monitoring, logging and distributed tracing stack for the systems you own.
Definition, tracking and reporting of Service Level Objectives and Indicators.
Owning the post-incident process β root cause analysis and corrective actions that actually get closed out.
DevSecOps Integration
Integration of security scanning (SAST/DAST) directly into the delivery pipeline.
Vulnerability management across application dependencies and runtime environments, with a rapid patching cadence.
Handling of application secrets, keys and certificates.
Capacity & Performance
Planning and execution of regular performance and load testing in close partnership with the QA team.
Configuration and optimisation of application-level auto-scaling across Kubernetes and compute resources.
Essential Skills
Experience: 6+ years in cloud infrastructure, DevOps, DevSecOps or SRE, including meaningful production ownership.
Prior technical lead experience, or senior experience with demonstrated mentoring and direction-setting.Cloud: Deep hands-on AWS.
Comfortable across compute, managed relational databases, managed caching, object storage, IAM and networking primitives.Kubernetes: Production ownership of a managed Kubernetes service (EKS, AKS or GKE) β cluster upgrades, node group management, ingress and load balancer routing, namespace segregation, and workload resource definitions.Infrastructure as Code: Strong proficiency with a declarative IaC toolchain (Terraform or CloudFormation) β authoring reusable modules, managing remote state and locking, debugging provider behaviour, and enforcing plan review gates.CI/CD: Demonstrated end-to-end ownership of pipelines on an enterprise platform (Azure DevOps, GitLab CI, GitHub Actions or Jenkins), including progressive and zero-downtime release strategies.Observability & SRE practice: Building and tuning monitoring, logging and tracing stacks.
Practical experience defining SLIs and SLOs and using error budgets to inform release decisions.Incident handling: Able to lead calmly during a production incident and to run a blameless post-mortem that produces corrective actions people actually complete.Security: Integrating SAST/DAST and container image scanning into delivery.
Managed secret handling.
Least-privilege identity design, including workload identity federation (IAM roles for service accounts or equivalent).Cost governance: Practical FinOps β right-sizing, autoscaling policy, Spot and Reserved capacity strategy, and identifying structural waste.Linux: Strong native Linux and WSL2 proficiency, with terminal-first operational habits.Communication: Able to explain a technical trade-off to a non-specialist stakeholder without either oversimplifying it or hiding behind detail.
We have 92,539 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume β in under a minute we'll analyze all 92,539 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume