GPU Infrastructure Architect

Vbeyond Corporation — United States · Posted ~3 hours ago

Senior

Skills

GPU profiling Distributed training High-performance computing LLM infrastructure GenAI infrastructure Multi-GPU environments GPU benchmarking Performance optimization Slurm Cloud/platform infrastructure Production deployments GPU GenAI LLM Distributed computing

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking an experienced infrastructure architect to design and scale platforms for demanding LLM and GenAI workloads. The role focuses on multi-GPU environments, GPU profiling and benchmarking, distributed training, job scheduling with Slurm, and production-grade deployments. You will collaborate with data, software, ML, MLOps, and application engineering teams to build reliable AI infrastructure.

Highlights

Work on scalable infrastructure for advanced AI workloads, collaborate across engineering and data disciplines, and contribute to production-grade GenAI platforms with a strong focus on GPU performance and distributed computing.

Description

We are looking for a highly skilled Architect - Platform Engineer to design, optimize, and scale infrastructure for GenAI and LLM workloads. This role is ideal for someone with deep hands-on experience in GPU profiling, distributed training, and high-performance compute environments. You will be working with Architects from other specialties such as Data engineering, Software engineering, ML engineering to create platforms, solutions and applications that cater to latest trends You’ll play a key role in building out GenAI platform foundations, supporting production-grade deployments, and partnering closely with data science, MLOps, and application teams to bring cutting-edge AI solutions to life. Key Responsibilities: Design and implement scalable infrastructure for LLM and GenAI workloads across multi-GPU environmentsPerform GPU profiling, benchmarking, and performance optimization for distributed training workloadsManage and schedule compute-intensive jobs using Slurm-based clusters and OpenShift/Kubernetes environmentsEnable and optimize the NVIDIA GPU stack (CUDA, cuDNN, NCCL, Triton, RAPIDS, etc.)Collaborate with cross-functional teams to deploy models in research and production environmentsBuild and support GenAI pipelines (fine-tuning, RAG, multi-modal inferencing, LLMOps)Develop reusable infrastructure templates using tools like Terraform and HelmContribute to internal innovation (PoCs, workshops) and support client-facing delivery engagementsDevelop and deliver automation software required for building & improving the functionality, reliability, availability, and manageability of applications and cloud platformsChampion and drive the adoption of Infrastructure as Code (IaC) practices and mindsetDesign, architect, and build self-service, self-healing, synthetic monitoring and alerting platform and toolsAutomate the development and test automation processes through CI/CD pipeline (Git, Jenkins, SonarQube, Artifactory, Docker containers)Build container hosting-platform using KubernetesIntroduce new cloud technologies, tools; processes to keep innovating in the commerce area to drive greater business value.Lead the technical discussion regarding architecture designing and troubleshooting with the clients and provide solutions proactively as requiredBasic Qualifications: Strong experience with Slurm and distributed training environmentsHands-on expertise with Red Hat OpenShift and/or KubernetesDeep knowledge of the NVIDIA GPU ecosystem (CUDA, cuDNN, NCCL, Nsight, Triton/TensorRT)Strong foundation in Linux systems, performance tuning, and multi-GPU optimizationExperience deploying GenAI workloads (LLM fine-tuning, RAG pipelines, multi-modal systems)Familiarity with Infrastructure-as-Code tools (Terraform, Ansible)Experience with cloud GPU environments (GCP, Azure, AWS, OCI) and/or on-prem GPU clustersServe as a mentor or guide for senior resources / team leads.Lead the technical discussion regarding architecture designOther Qualifications (OQs): Experience with NVIDIA NIMs, DGX systems, or GPU-accelerated containersKnowledge of LLMOps frameworks and MLOps integrationFamiliarity with vector databases and retrieval systems for RAG architecturesComfortable working in client-facing environments and collaborating with AI solution teamsHealthcare Domain Experience (Nice to Have): Experience working with FHIR R4, HL7 v2, or SMART on FHIRIntegration with EHR systems (e.g., Epic)Understanding of HIPAA compliance and healthcare data privacyExposure to clinical workflows, CDS Hooks, or patient-facing applicationsExperience building clinical decision support systems or healthcare interoperability solutions