Site Reliability Engineer - AI/GPU Infrastructure

Orangetalents — United States · Posted ~7 hours ago

Mid Full-time

Skills

SRE DevOps Platform engineering Infrastructure engineering Kubernetes GPU infrastructure Linux Networking Storage Distributed systems Automation Observability Incident response Root-cause analysis Capacity planning NVIDIA GPUs CNI CSI Load balancing Monitoring Logging

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An SRE is sought to operate and optimize production GPU and Kubernetes infrastructure supporting demanding AI workloads. The role spans reliability, scalability, automation, observability, resource allocation, networking, storage, security, incident response, and root-cause analysis in a distributed technical environment.

Highlights

Work with modern GPU clusters, Kubernetes, distributed systems, and AI workloads while shaping reliability, automation, observability, capacity planning, and incident response practices in an international environment.

Description

About the Opportunity Join a fast-growing AI infrastructure company building GPU cloud and production-grade training and inference platforms. You will work with modern NVIDIA GPU clusters, Kubernetes, distributed systems, and AI workloads while helping shape the reliability and automation practices of a growing international platform. Key Responsibilities Operate and optimize production Kubernetes and GPU clusters.Improve platform reliability, scalability, automation, and observability.Troubleshoot issues across Linux, networking, storage, Kubernetes, and distributed services.Manage resource allocation, utilization, quotas, and capacity planning.Optimize CNI, CSI, load balancing, monitoring, logging, and cluster security.Participate in on-call support, incident response, and root-cause analysis.Collaborate with technical teams and stakeholders across different regions. Requirements 3+ years of experience in SRE, DevOps, platform, or infrastructure engineering.Strong hands-on experience with production Kubernetes environments.Solid Linux administration and system troubleshooting skills.Familiarity with Kubernetes networking, storage, monitoring, and logging.Proficiency in Python, Go, Shell, or another scripting language.Strong problem-solving skills and ownership mindset. Nice to Have Experience with AI, machine learning, GPU, or HPC clusters.Familiarity with NVIDIA GPU infrastructure and distributed workloads.Experience with Terraform, Ansible, Helm, GitOps, or CI/CD.Experience with Prometheus, Grafana, ELK, or OpenTelemetry.Mandarin communication skills.