Summary
✨ AI‑Generated
Lead the architecture and evolution of large-scale infrastructure used for advanced AI model training. You will own multi-GPU and cross-cloud platforms, infrastructure as code, CI/CD for infrastructure and ML workflows, distributed workload optimization, and operational reliability while partnering closely with ML specialists and mentoring engineers.
Highlights
High-impact senior-to-staff role focused on large-scale AI infrastructure, multi-GPU computing, distributed systems, architecture, reliability, and performance, with strong technical ownership and mentoring responsibilities.
Description
We are looking for a Senior to Staff Infrastructure Engineer to lead the design and evolution of large-scale, multi-GPU compute infrastructure used to train next-generation robotics and AI models.
This role sits at the intersection of DevOps, MLOps, and distributed systems — owning architecture, reliability, and performance at scale in a fast-moving, cutting-edge environment.
Responsibilities:
Lead architecture and long-term technical direction of multi-GPU, cross-cloud training platformsBuild and evolve infrastructure-as-code for provisioning, orchestration, and lifecycle managementArchitect and improve CI/CD systems for infrastructure and ML training workflowsOptimize distributed training workloads — scheduling, resource utilization, observabilityPartner with ML engineers and researchers to enable efficient experimentation and productionizationMentor engineers and drive operational excellence across the orgDocument architecture, systems, and key technical decisions
Requirements:
Production-grade Kubernetes experience (CKA preferred)Hands-on Terraform (infrastructure-as-code)Kubernetes packaging and release management via HelmAWS cloud operations experienceCI/CD pipeline experience including self-hosted runners (GitHub Actions)Prometheus/Grafana monitoring and alertingLinux administration, containerization, scripting (Python & Bash)Availability for on-call rotation
Nice to Have:
GPU-accelerated Kubernetes clusters (NVIDIA)Cluster autoscaling (Karpenter)Workflow orchestration (Prefect)Gang scheduling, fair-share resource allocationHigh-performance storage (FSx for Lustre, EFS)
About us:
Grid Dynamics (NASDAQ: GDYN) is a leading provider of technology consulting, platform and product engineering, AI, and advanced analytics services.
Fusing technical vision with business acumen, we solve the most pressing technical challenges and enable positive business outcomes for enterprise companies undergoing business transformation.
A key differentiator for Grid Dynamics is our 8 years of experience and leadership in enterprise AI, supported by profound expertise and ongoing investment in data, analytics, cloud & DevOps, application modernization and customer experience.
Founded in 2006, Grid Dynamics is headquartered in Silicon Valley with offices across the Americas, Europe, and India.