Senior Cloud Network Site Reliability Engineer

Pracyva Ltd — United Kingdom · Posted ~21 hours ago

Senior

Skills

GCP networking AWS networking Kubernetes Incident management TCP/IP DNS Routing Firewalls VPNs Load balancers Network troubleshooting Root cause analysis Python Observability Automation GCP AWS GKE EKS VPN Load Balancers

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior cloud infrastructure engineer is sought to troubleshoot complex production networking issues across major cloud platforms, support Kubernetes environments, lead incident response and root-cause analysis, and improve reliability through automation and observability. Strong expertise in networking fundamentals and practical diagnostic techniques is essential.

Highlights

Senior-level cloud networking and reliability role focused on complex production troubleshooting, automation, observability, Kubernetes platforms, and continuous reliability improvements.

Description

We are looking for a Senior Cloud Network SRE / Cloud Platform Engineer with strong expertise in GCP/AWS networking, incident management, Kubernetes, and platform reliability engineering. The ideal candidate should be capable of troubleshooting complex cloud networking issues, leading production incident resolution, and driving reliability improvements through automation and observability. Key Responsibilities Troubleshoot and resolve complex GCP/AWS network incidents in production environments.Perform deep-dive analysis using TCP Dumps, packet captures, and network diagnostics tools.Diagnose and resolve issues related to TCP/IP, DNS, Routing, Firewalls, VPNs, Load Balancers, Network Overlays, Network Segregation, and Peer-to-Peer communication.Support and manage Kubernetes platforms (GKE/EKS).Build and maintain monitoring, alerting, and observability solutions.Perform incident response, problem management, and Root Cause Analysis (RCA).Automate operational tasks using Python and Shell scripting.Manage Infrastructure as Code (Terraform) and CI/CD pipelines.Mandatory Skills Strong hands-on experience in GCP and/or AWS Networking.Expertise in TCP/IP, DNS, Routing, Firewalls, VPNs, Load Balancing, and Packet Analysis.Experience with tcpdump, and network troubleshooting tools.Strong knowledge of Network Overlays, Network Segmentation, and Kubernetes Networking.Hands-on experience with Kubernetes (GKE/EKS).Experience with Dynatrace, or Cloud Monitoring tools.Python and Shell scripting expertise.Terraform and CI/CD experience.Strong Incident Management and RCA skills.Good to Have Banking/Financial Services experience.Service Mesh (Istio) knowledge.GCP/AWS/CKA/Terraform Certifications.Understanding of SLI/SLO/SLA and SRE practices.Screening Priority Cloud Networking & TroubleshootingKubernetes NetworkingIncident Management & RCAObservability & MonitoringPython AutomationTerraform & DevOpsGCP/AWS Platform Engineering