Description
· Work location: Jakarta
· Shifting: 8*5 and ready on call on midnight
Over all experience- Min 4+ Years
· Relevant Experience (in years) ? Minimum 2 years’ experience
Responsibilities:
Monitoring & Incident Handling
• Perform continuous 24x7 monitoring of Google Kubernetes Engine (GKE) clusters, applications, nodes, and supporting GCP services using approved monitoring tools.
• Respond to alerts, events, and system notifications according to defined operational procedures and escalation matrices.
• Perform incident triage and troubleshoot Kubernetes-related issues including pod failures, node issues, deployment failures, storage problems, ingress issues, and network connectivity problems.
• Investigate and resolve incidents affecting application availability, cluster performance, and platform stability.
• Analyze logs, metrics, traces, and events using Cloud Monitoring, Cloud Logging, Grafana, Prometheus, and other observability tools.
• Execute Root Cause Analysis (RCA) for major incidents and identify corrective and preventive actions.
• Coordinate with application, cloud, network, security, and development teams during incident resolution activities.
• Manage major incidents and participate in bridge calls to ensure timely restoration of services.
• Ensure compliance with SLA, SLO, and operational KPIs.
• Implement proactive measures to reduce recurring incidents and improve platform reliability.
Operational Tasks
• Administer and maintain Google Kubernetes Engine (GKE) environments across development, testing, staging, and production platforms.
• Perform cluster provisioning, upgrades, patching, maintenance, and lifecycle management activities.
• Manage Kubernetes resources including namespaces, deployments, services, ingress controllers, StatefulSets, DaemonSets, ConfigMaps, and Secrets.
• Support CI/CD deployment activities and troubleshoot release-related issues.
• Implement and maintain Infrastructure as Code (IaC) using Terraform, Helm Charts, and Kubernetes manifests.
• Configure and manage autoscaling, node pools, resource quotas, and capacity planning activities.
• Perform routine health checks on clusters, workloads, nodes, storage, and network components.
• Support backup, restoration, disaster recovery testing, and business continuity activities.
• Manage security configurations including RBAC, IAM roles, service accounts, workload identities, and policy enforcement.
• Perform vulnerability remediation, platform hardening, and security compliance activities.
• Optimize cluster performance, resource utilization, and operational costs within GCP environments.
• Develop automation scripts and operational improvements to enhance platform efficiency and reduce manual intervention.
• Support onboarding of new applications and workloads onto Kubernetes platforms.
• Manage and troubleshoot GCP services integrated with Kubernetes environments including Load Balancers, Cloud Storage, VPC Networks, Cloud DNS, IAM, and Cloud NAT.
Documentation & Reporting
• Maintain operational runbooks, standard operating procedures (SOPs), technical documentation, and troubleshooting guides.
• Document cluster configurations, architecture diagrams, deployment procedures, and platform changes.
• Prepare incident reports, Root Cause Analysis (RCA) documents, and post-incident review reports.
• Maintain shift handover reports, operational logs, and daily activity records.
• Track and report platform availability, capacity utilization, performance metrics, and incident trends.
• Ensure all operational activities, changes, and incidents are properly documented within ticketing systems.
• Produce weekly and monthly operational reports for management review.
• Maintain compliance evidence and support audit requirements.
• Recommend service improvement initiatives based on operational findings, recurring issues, and trend analysis.
• Facilitate knowledge sharing and maintain a centralized knowledge base for common operational procedures and resolutions.