Senior Kubernetes Platform Engineer - GPU & AI Infrastructure

Teamgtn — Unknown · Posted ~1 day ago

Senior Full-time Hybrid Visa History ✓

Skills

Kubernetes GPU orchestration Workload scheduling Infrastructure automation Performance engineering Reliability engineering Multi-tenant infrastructure AI/ML infrastructure High-performance computing GPU infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Engineer and scale a next-generation Kubernetes platform designed for GPU-intensive AI, machine learning, large language model, and high-performance computing workloads. You will own areas including cluster architecture, GPU orchestration, workload scheduling, automation, performance, reliability, and multi-tenant infrastructure. The preferred arrangement is hybrid, with full remote work considered for exceptional candidates.

Highlights

Build and scale next-generation GPU-accelerated infrastructure for AI, machine learning, large language models, and high-performance computing. The role offers significant technical ownership, hybrid flexibility, relocation support, performance bonuses, and fully company-paid benefits.

Description

Senior Kubernetes Platform Engineer – GPU & AI Infrastructure Location: Dallas, TX preferred Work Arrangement: Hybrid, 3 days onsite / 2 days remote Remote Flexibility: Full remote may be considered for the right candidate Relocation: Available Employment Type: Direct Hire Compensation: Competitive base salaryPerformance bonus100% company-paid benefitsOverview We are seeking a Senior Kubernetes Platform Engineer to help design, build, and scale a next-generation GPU-accelerated infrastructure platform supporting artificial intelligence, machine learning, large language models, and high-performance computing workloads. This is not a traditional Kubernetes administration role. The position is focused on engineering and evolving the underlying compute platform, with significant ownership across Kubernetes architecture, GPU orchestration, workload scheduling, automation, performance, reliability, and multi-tenant infrastructure. The preferred work arrangement is hybrid in Dallas, TX, with three days onsite and two days remote. However, full remote may be considered for a highly qualified candidate with strong Kubernetes, GPU infrastructure, and AI/HPC platform experience. The successful candidate will work closely with platform, infrastructure, HPC, and machine learning engineering teams to build highly scalable Kubernetes environments capable of supporting large GPU clusters and compute-intensive workloads. This is an opportunity for an engineer who enjoys working deep within Kubernetes, solving infrastructure problems at scale, and building the platforms that power modern AI workloads. Key Responsibilities Kubernetes Platform Engineering Design, build, operate, and scale large production Kubernetes environments supporting GPU-intensive workloads.Develop Kubernetes platforms for AI/ML training, inference, LLM workloads, and high-performance computing.Extend Kubernetes functionality through custom operators, controllers, CRDs, and platform automation.Improve cluster lifecycle management, provisioning, upgrades, configuration, and operational consistency.Develop reusable platform services and infrastructure capabilities for engineering and research teams.GPU Infrastructure & Scheduling Integrate and optimize NVIDIA GPU technologies within Kubernetes environments.Support technologies including NVIDIA GPU Operator, device plugins, DCGM, and GPU telemetry.Design GPU scheduling and allocation strategies including MIG, GPU sharing, workload placement, and resource isolation.Improve utilization and throughput across large GPU clusters.Evaluate and implement scheduling technologies such as Kubernetes scheduler extensions, Volcano, Slurm, or similar platforms.Partner with AI/ML and HPC engineering teams to optimize compute environments for training, inference, and scientific workloads.Performance & Scalability Identify and resolve performance bottlenecks across compute, networking, storage, Kubernetes, and GPU infrastructure.Optimize platforms for high-throughput, distributed, and latency-sensitive workloads.Improve cluster density, resource utilization, workload startup times, and overall infrastructure efficiency.Support performance testing, benchmarking, capacity planning, and platform scalability initiatives.Help design infrastructure capable of supporting rapidly growing GPU and compute environments.Platform Automation Build automation using Go, Python, or similar languages.Develop Kubernetes operators, controllers, APIs, and internal platform tooling.Implement infrastructure-as-code using technologies such as Terraform, Helm, and Kustomize.Build and improve GitOps workflows using Argo CD, Flux, or similar platforms.Automate cluster provisioning, configuration, policy enforcement, upgrades, and application deployment.Observability & Reliability Design monitoring, observability, and telemetry for Kubernetes and GPU infrastructure.Implement solutions using Prometheus, Grafana, DCGM Exporter, OpenTelemetry, and related technologies.Develop meaningful metrics and dashboards around GPU utilization, scheduling, cluster health, capacity, and workload performance.Participate in production readiness reviews, incident response, troubleshooting, and root cause analysis.Improve platform reliability through automation, monitoring, testing, and operational standards.Networking & Storage Support Kubernetes networking architectures for high-performance compute environments.Work with CNI technologies such as Multus, NVIDIA networking components, or similar solutions.Partner with infrastructure teams around high-performance networking, storage, and distributed systems.Troubleshoot complex issues spanning containers, networking, storage, GPUs, operating systems, and application workloads.Security & Multi-Tenancy Design secure multi-tenant Kubernetes environments.Implement namespace isolation, RBAC, resource quotas, network policies, and workload controls.Develop policy and governance frameworks using technologies such as OPA, Gatekeeper, or similar solutions.Help establish secure platform standards while maintaining developer and workload flexibility.Required Qualifications Strong experience designing and operating Kubernetes in large-scale production environments.Deep understanding of Kubernetes architecture and internals, including controllers, operators, CRDs, scheduling, RBAC, networking, and cluster lifecycle management.Hands-on experience supporting GPU-based compute environments.Experience with NVIDIA GPU technologies such as GPU Operator, device plugins, MIG, DCGM, or similar tools.Experience supporting AI/ML, LLM, HPC, scientific computing, or other compute-intensive workloads.Proficiency with Go, Python, or another programming language used for infrastructure automation.Experience developing Kubernetes operators, controllers, or custom automation.Strong background in infrastructure-as-code and automated platform deployment.Experience with Terraform, Helm, Kustomize, or similar technologies.Experience with GitOps platforms such as Argo CD or Flux.Strong Linux systems knowledge.Experience troubleshooting distributed infrastructure across compute, networking, storage, and application layers.Experience with Kubernetes monitoring and observability technologies.Preferred Qualifications Experience operating large NVIDIA GPU clusters.Experience with Slurm, Volcano, kube-scheduler extensions, or advanced workload scheduling.Experience supporting distributed machine learning or LLM training environments.Familiarity with PyTorch, TensorFlow, NCCL, CUDA, or other accelerated computing technologies.Experience with high-performance networking technologies such as InfiniBand, RDMA, RoCE, or related architectures.Experience with distributed or parallel storage platforms.Familiarity with bare-metal Kubernetes infrastructure.Experience building internal developer platforms or self-service infrastructure capabilities.Background working in HPC, AI infrastructure, cloud infrastructure, or large-scale distributed systems.Ideal Candidate The ideal candidate is a platform engineer who understands Kubernetes beyond deployment and administration. This individual should be comfortable working deep within Kubernetes architecture, developing operators and automation, troubleshooting complex distributed systems, and optimizing infrastructure for demanding GPU workloads. The strongest candidates will combine Kubernetes engineering expertise with experience across GPU infrastructure, AI/ML workloads, automation, distributed systems, and production platform reliability. Dallas-based candidates are preferred, but location will not outweigh exceptional technical fit. Full remote arrangements may be available for candidates with particularly strong experience in large-scale Kubernetes and GPU infrastructure. This is a high-impact opportunity to help build and scale the infrastructure powering next-generation AI and high-performance computing.