Description
KARS Technologies is looking for an experienced Kubernetes / GKE Operations Engineer to support the operation, monitoring, administration, and continuous improvement of enterprise Kubernetes environments running on Google Cloud Platform (GCP).
The ideal candidate should have hands-on experience with Google Kubernetes Engine (GKE), Kubernetes administration, incident management, cloud infrastructure, observability, automation, and production operations.
Key Responsibilities:
Monitoring & Incident Handling
Perform continuous monitoring of Google Kubernetes Engine (GKE) clusters, applications, nodes, and supporting GCP services using approved monitoring tools.Respond to alerts, events, and system notifications according to established operational procedures and escalation matrices.Perform incident triage and troubleshoot Kubernetes-related issues, including pod failures, node issues, deployment failures, storage problems, ingress issues, and network connectivity problems.Investigate and resolve incidents affecting application availability, cluster performance, and platform stability.Analyze logs, metrics, traces, and events using Google Cloud Monitoring, Cloud Logging, Grafana, Prometheus, and other observability tools.Perform Root Cause Analysis (RCA) for major incidents and identify corrective and preventive actions.Coordinate with application, cloud, network, security, and development teams during incident resolution.Participate in major incident bridge calls and ensure timely service restoration.Ensure compliance with agreed SLA, SLO, and operational KPIs.Proactively identify recurring issues and recommend improvements to enhance platform reliability.
GKE & Cloud Operations
Administer and maintain GKE environments across development, testing, staging, and production platforms.Perform cluster provisioning, upgrades, patching, maintenance, and lifecycle management.Manage Kubernetes resources, including Namespaces, Deployments, Services, Ingress Controllers, StatefulSets, DaemonSets, ConfigMaps, and Secrets.Support CI/CD deployment activities and troubleshoot release-related issues.Implement and maintain Infrastructure as Code (IaC) using Terraform, Helm Charts, and Kubernetes manifests.Configure and manage autoscaling, node pools, resource quotas, and capacity planning.Perform routine health checks across clusters, workloads, nodes, storage, and network components.Support backup, restoration, disaster recovery testing, and business continuity activities.Manage security configurations, including RBAC, IAM roles, service accounts, workload identities, and policy enforcement.Perform vulnerability remediation, platform hardening, and security compliance activities.Optimize cluster performance, resource utilization, and operational costs within GCP environments.Develop automation scripts and operational improvements to reduce manual intervention.Support onboarding and migration of applications and workloads onto Kubernetes platforms.Manage and troubleshoot Kubernetes-integrated GCP services, including Load Balancers, Cloud Storage, VPC Networks, Cloud DNS, IAM, and Cloud NAT.
Documentation & Reporting
Maintain operational runbooks, SOPs, technical documentation, and troubleshooting guides.Document cluster configurations, architecture diagrams, deployment procedures, and platform changes.Prepare incident reports, RCA documents, and post-incident review reports.Maintain shift handover reports, operational logs, and daily activity records.Track platform availability, capacity utilization, performance metrics, and incident trends.Ensure operational activities, changes, and incidents are properly recorded in ticketing systems.Prepare weekly and monthly operational reports for management review.Maintain compliance evidence and support audit requirements.Recommend service improvement initiatives based on recurring incidents, operational findings, and trend analysis.Facilitate knowledge sharing and maintain a centralized knowledge base.
Required Qualifications
Minimum 2 years of relevant hands-on experience in Kubernetes, cloud infrastructure, DevOps, SRE, or platform operations.Strong hands-on experience administering Kubernetes and Google Kubernetes Engine (GKE).Good understanding of Linux administration, networking, storage, security, and cloud infrastructure.Experience with Terraform, Helm, CI/CD pipelines, Prometheus, Grafana, Google Cloud Monitoring, and Cloud Logging is highly preferred.Strong troubleshooting and incident management skills in production environments.Ability to work in a shift-based operational environment and provide midnight/on-call support when required.Good Bahasa Indonesia & English communication, documentation, and cross-functional collaboration skills.
Certification Requirements
Candidates should hold relevant professional certifications, including:
✓ Certified Kubernetes Administrator (CKA)
✓ Red Hat Certification – RHCSA, RHCE, or Red Hat Certified OpenShift Administrator (RHCOSA)
✓ Google Cloud Certification – Associate Cloud Engineer and/or Professional Cloud DevOps Engineer
Employment / Project Details
This is a 7-month project-based position, with an estimated engagement period from September 2026 through March 2027.
The selected engineer will be based in Jakarta and will work closely with cloud, infrastructure, application, security, and development teams to maintain the reliability, availability, security, and performance of mission-critical Kubernetes environments.