SENIOR ENGINEER – KUBERNETES ADMINISTRATION

Rstratsolutions — United Arab Emirates · Posted ~22 hours ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

WE ARE HIRING | SENIOR ENGINEER – KUBERNETES ADMINISTRATION Rstrat Technologies Rstrat Technologies invites applications from experienced Kubernetes and infrastructure professionals for a Senior Engineer role focused on designing, building, securing, and operating on-premises Kubernetes platforms for AI and machine learning workloads. The successful candidate will take ownership of the complete cluster lifecycle, from provisioning and deployment through upgrades, optimisation, and decommissioning. This role requires strong technical ownership and hands-on experience supporting reliable, secure production infrastructure in data-centre environments. KEY RESPONSIBILITIES Cluster Architecture and Deployment • Design and provision production-grade Kubernetes clusters across bare-metal and virtualised environments. • Implement highly available control planes, maintain etcd health, and configure node pools. • Support kubeadm-based deployments, upstream Kubernetes, OpenShift, or Rancher. • Establish secure cluster bootstrapping in connected and air-gapped environments, including private registries, image mirrors, and operating-system hardening. Cluster Operations and Reliability • Manage upgrades, patching, capacity planning, and node lifecycle automation. • Configure resource quotas, workload scheduling, affinities, taints, tolerations, and multi-tenancy isolation. • Implement autoscaling and resource optimisation. • Maintain backup, restore, and disaster recovery procedures with documented runbooks and regular testing. AI/ML and GPU Infrastructure • Enable and manage GPU access for AI training, inference, and related workloads. • Support model-serving platforms such as KServe or Seldon and orchestration tools such as Kubeflow or Argo Workflows. • Improve GPU utilisation, workload throughput, and inference latency. Networking, Storage, and Security • Configure CNI networking, ingress, load balancing, DNS, and network policies. • Manage CSI drivers, dynamic storage provisioning, and stateful workloads. • Implement RBAC, Pod Security, admission policies, image scanning and signing, and secrets management. • Integrate enterprise identity services, certificate management, and audit logging. Platform Automation and Observability • Standardise deployments using Helm, Kustomize, GitOps, and infrastructure as code. • Support on-premises CI/CD runners and artifact registries. • Implement monitoring, logging, and tracing using tools such as Prometheus, Grafana, ELK/EFK, Loki, and OpenTelemetry. • Define service-level indicators and objectives, lead incident response and root-cause analysis, and participate in an on-call rotation. EXPERIENCE AND QUALIFICATIONS • Bachelor’s degree in Computer Science, Engineering, or a related discipline, or equivalent practical experience. • 7–10 years of experience in Platform Engineering, SRE, DevOps, or Systems Administration. • At least 4 years of hands-on experience administering production Kubernetes environments. • Proven experience designing, building, and operating on-premises Kubernetes clusters independently of managed cloud services. • Practical experience with GPU infrastructure, air-gapped operations, private registries, and enterprise security integrations. • CKA certification is required, or must be obtained within six months. CKS or CKAD certifications are preferred. RELEVANT TECHNICAL SKILLS • Kubernetes: Control-plane operations, etcd, scheduling, admission controllers, namespaces, quotas, and PodDisruptionBudgets. • Linux and Containers: Linux administration, containerd/CRI-O, systemd, kernel and GPU drivers, and OS hardening. • Networking: Cilium/Calico, NGINX/HAProxy, MetalLB, CoreDNS, and ingress/egress controls. • Storage: CSI, PV/PVC, snapshots, backup and recovery, and platforms such as Ceph, Rook, Longhorn, NetApp, or NFS. • Security: OPA Gatekeeper/Kyverno, Vault/Sealed Secrets, cosign/Trivy, and OIDC/LDAP/Active Directory. • Automation: Helm/Kustomize, Argo CD/Flux, Terraform/Ansible, and Bash/Python scripting. We are looking for candidates who demonstrate clear documentation, effective incident communication, strong problem-solving skills, and the ability to collaborate with stakeholders and mentor team members.