Site Reliability Engineer - AI/GPU Infrastructure
Orangetalents — United States · Posted ~7 hours ago
Skills
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listSummary ✨ AI‑Generated
An SRE is sought to operate and optimize production GPU and Kubernetes infrastructure supporting demanding AI workloads. The role spans reliability, scalability, automation, observability, resource allocation, networking, storage, security, incident response, and root-cause analysis in a distributed technical environment.
Highlights
Work with modern GPU clusters, Kubernetes, distributed systems, and AI workloads while shaping reliability, automation, observability, capacity planning, and incident response practices in an international environment.
Description
We have 122,795 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 122,795 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume