Kubernetes / HPC Engineer

Epsilon Solutions Canada — Canada · Posted ~5 hours ago

Senior Full-time Remote

Skills

Kubernetes Docker Helm container orchestration HPC Slurm Volcano distributed computing NVIDIA GPU infrastructure Linux Python Bash CI/CD monitoring infrastructure automation cloud platforms NVIDIA GPUs Azure AKS AWS EKS GCP GKE

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Work remotely as an infrastructure engineer responsible for Kubernetes and high-performance computing environments. You will operate containerized workloads, GPU infrastructure, cloud platforms, CI/CD pipelines, monitoring, and automation, while supporting large-scale distributed computing. Experience with AI training, scientific workloads, or other demanding compute environments is valuable.

Highlights

Remote Canadian role combining Kubernetes, cloud infrastructure, HPC, GPU computing, AI/ML workloads, and large-scale distributed systems, with exposure to advanced compute environments.

Description

Role name: Kubernetes (K8S) / HPC Engineer Work Location: CANADA (Remote) Key Skills: Hands-on experience with Kubernetes (K8S), Docker, Helm, and container orchestration.Strong knowledge of HPC clusters, workload schedulers (Slurm/Volcano), and distributed computing environments.Experience with GPU infrastructure (NVIDIA), AI/ML workloads, and performance optimization.Expertise in Linux, scripting (Python/Bash), CI/CD, monitoring, and infrastructure automation.Experience with cloud platforms such as Azure AKS, AWS EKS, or GCP GKE. Preferred Qualifications: Familiarity with high-speed networking (RDMA/InfiniBand), parallel storage systems, and large-scale compute environments.Experience supporting engineering simulations, semiconductor workloads, AI model training, or scientific computing platforms.Experience: 5+ years in Kubernetes/Cloud Infrastructure with at least 2+ years in HPC or large-scale compute platforms.