AI Infrastructure Engineer

Workerbeeai — Canada · Posted ~22 hours ago

Skills

GPU infrastructure Model serving AI infrastructure Infrastructure engineering GPU Machine learning infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Work as an AI Infrastructure Engineer helping organizations solve demanding infrastructure challenges for modern machine-learning workloads. The role focuses on hands-on expertise with GPU platforms and model-serving infrastructure, with opportunities to contribute through project-based, contract, or permanent engagements.

Highlights

Flexible opportunity supporting project-based, contract, or permanent career paths while solving challenging AI infrastructure problems involving GPU platforms and model-serving systems.

Description

C2C not available | No third-party suppliers Workerbee employees will not respond to direct communication attempts via email, phone, social media, or LinkedIn regarding status within the talent network or customer needs. Any such inquiries will not receive a response. About Workerbee Workerbee connects workers with employers through trusted introductions. By joining Workerbee you can be matched for project-based, contract, or permanent opportunities with leading organizations. Over time, Workerbee helps you: Keep a living record of what you have actually accomplishedSee how your experience carries across roles and pathsExplore options without pressure to applyMove through change with clarity instead of urgencyWe are connected to those who hire talent and are in need of AI Infrastructure Engineers. In this work, you will help employers solve complex problems and bring real value through hands-on expertise in GPU platforms and model serving infrastructure. Where Your Expertise Makes an Impact Build and operate GPU clusters and model serving platforms on Kubernetes using vLLM, TensorRT-LLM, Triton, or Ray ServeTune throughput and latency through batching strategy, quantization, KV cache management, and model parallelism across acceleratorsManage capacity, scheduling, and the mix of reserved and spot GPU supply so training and inference workloads share hardware without contentionTrack cost per token and per request, then bring it down through right-sizing, autoscaling policy, and workload placementHarden the platform with rollout automation, canary deploys, regional failover, and monitoring against latency and error budgets What Stands Out Among Top Talent 5+ years in platform, SRE, or infrastructure engineering with production Kubernetes at scaleDirect experience serving models on GPUs, including CUDA fundamentals, memory profiling, and utilization tuningFluency with Terraform, Helm, Argo CD, and observability stacks such as Prometheus and GrafanaWorking knowledge of PyTorch, ONNX, and quantized model formats including GGUF, AWQ, and FP8Cost discipline and the ability to explain infrastructure tradeoffs to both engineering and finance stakeholders Why Join Workerbee Earlier visibility into opportunitiesBetter-fit introductionsAccess to meaningful workLess application noiseA network that improves over time Workerbee Terms of Service