Summary
✨ AI‑Generated
A senior cloud infrastructure engineer is sought to troubleshoot complex production networking issues across major cloud platforms, support Kubernetes environments, lead incident response and root-cause analysis, and improve reliability through automation and observability. Strong expertise in networking fundamentals and practical diagnostic techniques is essential.
Highlights
Senior-level cloud networking and reliability role focused on complex production troubleshooting, automation, observability, Kubernetes platforms, and continuous reliability improvements.
Description
We are looking for a Senior Cloud Network SRE / Cloud Platform Engineer with strong expertise in GCP/AWS networking, incident management, Kubernetes, and platform reliability engineering.
The ideal candidate should be capable of troubleshooting complex cloud networking issues, leading production incident resolution, and driving reliability improvements through automation and observability.
Key Responsibilities
Troubleshoot and resolve complex GCP/AWS network incidents in production environments.Perform deep-dive analysis using TCP Dumps, packet captures, and network diagnostics tools.Diagnose and resolve issues related to TCP/IP, DNS, Routing, Firewalls, VPNs, Load Balancers, Network Overlays, Network Segregation, and Peer-to-Peer communication.Support and manage Kubernetes platforms (GKE/EKS).Build and maintain monitoring, alerting, and observability solutions.Perform incident response, problem management, and Root Cause Analysis (RCA).Automate operational tasks using Python and Shell scripting.Manage Infrastructure as Code (Terraform) and CI/CD pipelines.Mandatory Skills
Strong hands-on experience in GCP and/or AWS Networking.Expertise in TCP/IP, DNS, Routing, Firewalls, VPNs, Load Balancing, and Packet Analysis.Experience with tcpdump, and network troubleshooting tools.Strong knowledge of Network Overlays, Network Segmentation, and Kubernetes Networking.Hands-on experience with Kubernetes (GKE/EKS).Experience with Dynatrace, or Cloud Monitoring tools.Python and Shell scripting expertise.Terraform and CI/CD experience.Strong Incident Management and RCA skills.Good to Have
Banking/Financial Services experience.Service Mesh (Istio) knowledge.GCP/AWS/CKA/Terraform Certifications.Understanding of SLI/SLO/SLA and SRE practices.Screening Priority
Cloud Networking & TroubleshootingKubernetes NetworkingIncident Management & RCAObservability & MonitoringPython AutomationTerraform & DevOpsGCP/AWS Platform Engineering