Lead DevOps
Pvgrp — Indonesia · Posted ~1 day ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
The Lead DevOps is responsible for leading the management of one of the group entities running in rural bank business, covering IT infrastructure, cloud platforms, application deployment, observability, system security, and overall service reliability.
This is both a leadership and hands-on role, with a primary focus on ensuring that the bank’s systems operate securely, reliably, and efficiently, while maintaining proper documentation, scalability, and robust recovery capabilities in the event of system disruptions or incidents.
The Mission:
Lead the planning, implementation, and management of IT infrastructure and application platforms.Ensure application deployment processes are automated, consistent, reliable, and recoverable in the event of failures.Establish and implement standards for CI/CD, GitOps, Infrastructure as Code (IaC), and configuration management.Ensure the availability, performance, security, reliability, and scalability of production services.Manage monitoring, logging, tracing, alerting, and operational dashboards to ensure effective system observability.Establish and maintain incident response, escalation, root cause analysis (RCA), and post-incident review processes.Manage backup, restoration, and replication processes, and support the implementation and testing of the Disaster Recovery Plan (DRP) and Disaster Recovery Center (DRC).Conduct capacity planning and optimize infrastructure performance, utilization, and costs.Establish and enforce infrastructure security standards, including access management, credentials, secrets, certificates, and audit trails.Ensure regular system patching, vulnerability management, security hardening, and infrastructure maintenance.Support audit requirements, information security initiatives, and compliance with company policies and applicable regulatory requirements.Develop and maintain architecture documentation, operational Standard Operating Procedures (SOPs), runbooks, and troubleshooting guidelines.Collaborate closely with Development, QA, Security, Product, and third-party teams throughout the development lifecycle through production deployment and operations.Lead, mentor, and develop DevOps team members to continuously enhance their technical capabilities and overall team performance.Evaluate and recommend technologies and architectural solutions that align with the company’s business requirements, operational needs, and scale.
General Requirements:
Bachelor’s degree in Information Technology, Computer Science, Informatics Engineering, or a related field.Minimum of 5 years of professional experience in DevOps, Infrastructure Engineering, Cloud Engineering, Site Reliability Engineering (SRE), or related areas.Proven experience as a Lead DevOps Engineer, Senior DevOps Engineer, or in an equivalent technical leadership role.Strong understanding of high availability, scalability, disaster recovery, business continuity, Recovery Time Objective (RTO), and Recovery Point Objective (RPO) concepts.Strong understanding of networking fundamentals and technologies, including DNS, load balancers, reverse proxies, firewalls, TLS, and network security.Solid understanding of CI/CD, GitOps, Infrastructure as Code (IaC), containerization, and configuration management practices.Solid understanding of monitoring, logging, tracing, alerting, and incident management practices.Strong troubleshooting, root cause analysis (RCA), and problem-solving skills.Strong leadership, mentoring, communication, and cross-functional coordination skills.Willingness and ability to respond to production incidents outside regular working hours in accordance with the company’s on-call procedures.
Technical Requirements:
Strong expertise in Kubernetes cluster administration, management, and troubleshooting.Strong expertise in Linux system administration, security hardening, patching, and troubleshooting.Hands-on experience implementing GitOps practices and managing deployments using Argo CD.Hands-on experience with HashiCorp Vault for managing secrets, credentials, certificates, and encryption keys.Strong experience in implementing and managing the LGTM observability stack, including: Loki, Grafana, Tempo, Mimir and/or PrometheusExperience using Thanos for high availability and long-term metrics storage.Hands-on experience managing cloud services on one or more of the following platforms: Alibaba Cloud, Google Cloud Platform (GCP), Amazon Web Services (AWS)Strong proficiency in Infrastructure as Code (IaC) using OpenTofu.Experience using Pulumi for infrastructure provisioning and management.Hands-on experience in PostgreSQL administration, monitoring, backup, replication, and performance tuning.Solid understanding of message broker implementation and management using RabbitMQ.Solid understanding of event streaming implementation and management using Apache Kafka.Solid understanding of Redis implementation and management, including monitoring, persistence, and high-availability configurations.
We have 68,970 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 68,970 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume