DevOps Engineer / GCP Kubernetes Engineer

Accordinnovations — Indonesia · Posted ~2 hours ago

Mid Onsite

Skills

Google Kubernetes Engine (GKE) Kubernetes Google Cloud Platform (GCP) Monitoring Incident response Troubleshooting Cloud Monitoring Cloud Logging Grafana Prometheus Log analysis Metrics and tracing Root cause analysis GKE GCP

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A hands-on DevOps and cloud platform role focused on operating Kubernetes environments, monitoring production systems, responding to incidents, troubleshooting infrastructure and application issues, and improving platform stability. You will work extensively with cloud services and modern observability tools while participating in root cause analysis and operational support.

Highlights

Opportunity to work with cloud-native infrastructure, Kubernetes platforms, observability tooling, incident response, and production reliability in a hands-on operational engineering environment.

Description

· Work location: Jakarta · Shifting: 8*5 and ready on call on midnight Over all experience- Min 4+ Years · Relevant Experience (in years) ? Minimum 2 years’ experience Responsibilities: Monitoring & Incident Handling • Perform continuous 24x7 monitoring of Google Kubernetes Engine (GKE) clusters, applications, nodes, and supporting GCP services using approved monitoring tools. • Respond to alerts, events, and system notifications according to defined operational procedures and escalation matrices. • Perform incident triage and troubleshoot Kubernetes-related issues including pod failures, node issues, deployment failures, storage problems, ingress issues, and network connectivity problems. • Investigate and resolve incidents affecting application availability, cluster performance, and platform stability. • Analyze logs, metrics, traces, and events using Cloud Monitoring, Cloud Logging, Grafana, Prometheus, and other observability tools. • Execute Root Cause Analysis (RCA) for major incidents and identify corrective and preventive actions. • Coordinate with application, cloud, network, security, and development teams during incident resolution activities. • Manage major incidents and participate in bridge calls to ensure timely restoration of services. • Ensure compliance with SLA, SLO, and operational KPIs. • Implement proactive measures to reduce recurring incidents and improve platform reliability. Operational Tasks • Administer and maintain Google Kubernetes Engine (GKE) environments across development, testing, staging, and production platforms. • Perform cluster provisioning, upgrades, patching, maintenance, and lifecycle management activities. • Manage Kubernetes resources including namespaces, deployments, services, ingress controllers, StatefulSets, DaemonSets, ConfigMaps, and Secrets. • Support CI/CD deployment activities and troubleshoot release-related issues. • Implement and maintain Infrastructure as Code (IaC) using Terraform, Helm Charts, and Kubernetes manifests. • Configure and manage autoscaling, node pools, resource quotas, and capacity planning activities. • Perform routine health checks on clusters, workloads, nodes, storage, and network components. • Support backup, restoration, disaster recovery testing, and business continuity activities. • Manage security configurations including RBAC, IAM roles, service accounts, workload identities, and policy enforcement. • Perform vulnerability remediation, platform hardening, and security compliance activities. • Optimize cluster performance, resource utilization, and operational costs within GCP environments. • Develop automation scripts and operational improvements to enhance platform efficiency and reduce manual intervention. • Support onboarding of new applications and workloads onto Kubernetes platforms. • Manage and troubleshoot GCP services integrated with Kubernetes environments including Load Balancers, Cloud Storage, VPC Networks, Cloud DNS, IAM, and Cloud NAT. Documentation & Reporting • Maintain operational runbooks, standard operating procedures (SOPs), technical documentation, and troubleshooting guides. • Document cluster configurations, architecture diagrams, deployment procedures, and platform changes. • Prepare incident reports, Root Cause Analysis (RCA) documents, and post-incident review reports. • Maintain shift handover reports, operational logs, and daily activity records. • Track and report platform availability, capacity utilization, performance metrics, and incident trends. • Ensure all operational activities, changes, and incidents are properly documented within ticketing systems. • Produce weekly and monthly operational reports for management review. • Maintain compliance evidence and support audit requirements. • Recommend service improvement initiatives based on operational findings, recurring issues, and trend analysis. • Facilitate knowledge sharing and maintain a centralized knowledge base for common operational procedures and resolutions.