Site Reliability Engineer – Kubernetes/Linux Platform (Remote – MENA)

Saturn Cloud — United Arab Emirates · Posted ~4 hours ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

This is a fully remote role open to candidates across the MENA region About Saturn CloudSaturn Cloud builds infrastructure for running AI, machine learning, and data workloads at scale. Our platform helps teams develop, deploy, and operate compute-intensive workloads across modern cloud and GPU infrastructure. We’re looking for a Site Reliability Engineer based in the MENA region to help operate and troubleshoot the production infrastructure supporting Saturn Cloud Token Factory. This is an infrastructure/SRE role rather than a traditional customer support position. You’ll be expected to independently diagnose production incidents across Kubernetes, Linux, networking, containers, storage, and cloud infrastructure, participate in incident response and SLA coverage, and work directly with customers and GPU infrastructure providers when necessary. What You’ll DoOwn and troubleshoot production incidents across GPU-backed Kubernetes environmentsDiagnose issues spanning Kubernetes, Linux, networking, containers, storage, and cloud infrastructureInvestigate Kubernetes scheduling failures involving resource requests and limits, affinity, taints and tolerations, PriorityClasses, quotas, and related configurationTroubleshoot container runtimes including containerd, Docker, and OCI runtimesOperate and debug Kubernetes networking, including Services, Ingress, DNS, NetworkPolicy, CNI, and load balancersDiagnose TCP/IP, DNS, and TLS issues in production environmentsTroubleshoot Kubernetes storage and CSI issues, including PVCs/PVs, block storage, shared filesystems, and mount failuresImprove monitoring, alerting, and production observability using tools such as Prometheus and GrafanaBuild diagnostic and operational automation using Bash and PythonWork directly with customer and infrastructure-provider teams during production incidentsDetermine whether failures originate within Saturn Cloud, Kubernetes, customer infrastructure, networking/storage, or underlying GPU infrastructureProduce clear technical evidence for escalation when an issue belongs outside Saturn Cloud’s infrastructure layer What We’re Looking ForCore SkillsDeep hands-on experience troubleshooting Kubernetes in productionStrong Linux systems administration and debugging skillsStrong knowledge of Kubernetes scheduling and resource managementExperience with containerd, Docker, or other OCI-compatible container runtimesExperience with Helm and Kubernetes deployment/configuration managementStrong understanding of Kubernetes networkingStrong TCP/IP, DNS, and TLS troubleshooting skillsExperience troubleshooting Kubernetes storage and CSIExperience with Prometheus, Grafana, or comparable observability toolingBash and Python experience for diagnostics and operational automationExperience with at least one major managed Kubernetes environment such as EKS, GKE, AKS, or OKE GPU KnowledgeYou don’t need to be the team’s deepest GPU specialist, but you should be comfortable operating GPU workloads and identifying when an incident has moved beyond the Kubernetes or container layer. Experience should include: NVIDIA GPU Operator and Kubernetes device pluginsGPU resource allocation and schedulingnvidia-smi and basic GPU health diagnosticsNVIDIA Container Toolkit/runtimeCUDA and driver compatibility conceptsIdentifying GPU, driver, or hardware failures that require deeper infrastructure escalation Strong PlusesCilium or CalicoTerraformKAI Scheduler, Grove, Volcano, or other batch/GPU schedulersvLLM, NVIDIA Dynamo, Triton, or other inference systemsMulti-tenant Kubernetes platformsExperience operating infrastructure against contractual availability SLAsExperience working directly with enterprise customers during production incidents What Success Looks LikeWhen an inference endpoint becomes unavailable, you can systematically trace the issue through the load balancer, ingress, Kubernetes Service, workload, scheduler, container runtime, and node. You can determine whether the root cause belongs to Saturn Cloud, Kubernetes, the customer’s network or storage environment, or the underlying GPU infrastructure, and provide useful evidence to the appropriate team. You’re comfortable being the primary incident owner, not simply collecting logs and passing the problem to someone else. Why Saturn CloudRemote-first culture with a high-trust, high-ownership environmentWork on production infrastructure at the center of modern AI and GPU computingSolve technically challenging reliability problems across Kubernetes, cloud, and GPU infrastructureHelp shape the operational practices of a rapidly evolving AI infrastructure platform Compensation & BenefitsCompetitive salaryFlexible PTOFully remote position