GCP Platform Engineer

Quantiphi — Canada · Posted ~3 hours ago

Senior

Skills

Google Cloud Platform Kubernetes GPU computing Distributed training MLOps GCP CUDA cuDNN NCCL OpenShift

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking a Platform Engineer to design scalable cloud infrastructure for generative AI workloads, distributed computing, and production ML systems.

Highlights

Build advanced AI infrastructure supporting large-scale machine learning workloads and collaborate with AI engineering teams.

Description

Role Overview: We are looking for a highly skilled Senior Platform Engineer to design, optimize, and scale infrastructure for GenAI and LLM workloads. This role is ideal for someone with deep hands-on experience in GPU profiling, distributed training, and high-performance compute environments. You’ll play a key role in building out GenAI platform foundations, supporting production-grade deployments, and partnering closely with data science, MLOps, and application teams to bring cutting-edge AI solutions to life. Key Responsibilities: Design and implement scalable infrastructure for LLM and GenAI workloads across multi-GPU environmentsPerform GPU profiling, benchmarking, and performance optimization for distributed training workloadsManage and schedule compute-intensive jobs using Slurm-based clusters and OpenShift/Kubernetes environmentsEnable and optimize the NVIDIA GPU stack (CUDA, cuDNN, NCCL, Triton, RAPIDS, etc.)Collaborate with cross-functional teams to deploy models in research and production environmentsBuild and support GenAI pipelines (fine-tuning, RAG, multi-modal inferencing, LLMOps)Develop reusable infrastructure templates using tools like Terraform and HelmContribute to internal innovation (PoCs, workshops) and support client-facing delivery engagements Basic Qualifications: Strong experience with Slurm and distributed training environmentsHands-on expertise with Red Hat OpenShift and/or KubernetesDeep knowledge of the NVIDIA GPU ecosystem (CUDA, cuDNN, NCCL, Nsight, Triton/TensorRT)Strong foundation in Linux systems, performance tuning, and multi-GPU optimizationExperience deploying GenAI workloads (LLM fine-tuning, RAG pipelines, multi-modal systems)Familiarity with Infrastructure-as-Code tools (Terraform, Ansible)Experience with cloud GPU environments (GCP, Azure, AWS, OCI) and/or on-prem GPU clusters Other Qualifications (OQs): Experience with NVIDIA NIMs, DGX systems, or GPU-accelerated containersKnowledge of LLMOps frameworks and MLOps integrationFamiliarity with vector databases and retrieval systems for RAG architecturesComfortable working in client-facing environments and collaborating with AI solution teams Healthcare Domain Experience (Nice to Have): Experience working with FHIR R4, HL7 v2, or SMART on FHIRIntegration with EHR systems (e.g., Epic)Understanding of HIPAA compliance and healthcare data privacyExposure to clinical workflows, CDS Hooks, or patient-facing applicationsExperience building clinical decision support systems or healthcare interoperability solutions What’s in it for YOU at Quantiphi: Make an impact at one of the world’s fastest-growing AI-first digital engineering companies.Upskill and discover your potential as you solve complex challenges in cutting-edge areas of technology alongside passionate, talented colleagues.Work where innovation happens - work with disruptive innovators in a research-focused organization with 60+ patents filed across various disciplines.Stay ahead of the curve, immerse yourself in breakthrough AI, ML, data, and cloud technologies and gain exposure working with Fortune 500 companies.