Senior Infrastructure Engineer
Insight Global โ United States ยท Posted ~1 day ago
๐ Log in to save this job, tailor your resume & track your apply process โ 7 days free, no card needed.
Log in to add to target listDescription
THIS ROLE IS LOCATED IN A FEW DIFFERENT LOCATIONS:
New York, Texas, California
Total Compensation: $140k-220k + Company Equity
*Relocation Package is provided if needed*
The below is a starting point.
We always make space for exceptional people, so if you do not fit every requirement exactly, we still encourage you to apply.
Required Qualifications
Experience bringing up server or GPU fleets at scale, hundreds of nodes or more, and taking them all the way to production.
Strong Linux background with hands-on experience using out-of-band management technologies such as BMC, IPMI, and Redfish.
Experience automating hardware workflows using Python or Go rather than relying on manual processes.
Hands-on data center experience including racking, cabling, hardware installation, troubleshooting, and component replacement.
Ability to methodically troubleshoot issues across hardware, firmware, and software layers, isolating root causes before implementing fixes.
Willingness and ability to travel during deployment and turn-up activities.
Nice to Have Skills & Experience
Kubernetes-based bare metal provisioning.
Accelerator platform bring-up and validation (NVIDIA, AMD, or custom hardware).
Burn-in and stress-testing framework design.
DCIM and inventory management tooling experience.
Job Description
Our client is a leading AI infrastructure company building and operating large-scale compute environments that power next-generation AI workloads.
Their teams design, deploy, and operate high-performance data center infrastructure at massive scale, with a focus on speed, reliability, and operational excellence.
They are seeking a Compute Deployment Engineer to own server and accelerator cluster bring-up from facility readiness through production deployment.
This role is ideal for someone who thrives in fast-paced environments, takes ownership of large-scale deployments, and enjoys solving complex infrastructure challenges.
Responsibilities:
Own compute turn-up from facility availability to ready-for-service: the stretch after the network hands off and before customers run workloads.
Qualify racks at scale: establish firmware baselines, configure BMC and BIOS, run burn-in, and validate at node and cluster level across hundreds of racks per site on GPU and custom accelerator platforms.
Drive qualification through the base-management Kubernetes platform and provisioning stack (discovery, imaging, firmware updates, shared services), burning down qualification queues with tooling rather than manual processes.
Triage hardware failures found in qualification: isolate to component, drive RMA and vendor escalation, and feed failure patterns back into qualification gates.
Run turn-up remotely by default, with on-site deployments of roughly a week per data hall as new halls reach facility availability, plus occasional overlapping-site engagements.
Partner with network deployment, ICT, data center operations, and hardware teams during turn-up windows, and support incident response on newly deployed capacity.
Ability to travel 20-30% of the time to data centers and lab environments as needed.
We have 70,648 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume โ in under a minute we'll analyze all 70,648 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume