Senior Infrastructure Engineer

Insight Global โ€” United States ยท Posted ~1 day ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

THIS ROLE IS LOCATED IN A FEW DIFFERENT LOCATIONS: New York, Texas, California Total Compensation: $140k-220k + Company Equity *Relocation Package is provided if needed* The below is a starting point. We always make space for exceptional people, so if you do not fit every requirement exactly, we still encourage you to apply. Required Qualifications Experience bringing up server or GPU fleets at scale, hundreds of nodes or more, and taking them all the way to production. Strong Linux background with hands-on experience using out-of-band management technologies such as BMC, IPMI, and Redfish. Experience automating hardware workflows using Python or Go rather than relying on manual processes. Hands-on data center experience including racking, cabling, hardware installation, troubleshooting, and component replacement. Ability to methodically troubleshoot issues across hardware, firmware, and software layers, isolating root causes before implementing fixes. Willingness and ability to travel during deployment and turn-up activities. Nice to Have Skills & Experience Kubernetes-based bare metal provisioning. Accelerator platform bring-up and validation (NVIDIA, AMD, or custom hardware). Burn-in and stress-testing framework design. DCIM and inventory management tooling experience. Job Description Our client is a leading AI infrastructure company building and operating large-scale compute environments that power next-generation AI workloads. Their teams design, deploy, and operate high-performance data center infrastructure at massive scale, with a focus on speed, reliability, and operational excellence. They are seeking a Compute Deployment Engineer to own server and accelerator cluster bring-up from facility readiness through production deployment. This role is ideal for someone who thrives in fast-paced environments, takes ownership of large-scale deployments, and enjoys solving complex infrastructure challenges. Responsibilities: Own compute turn-up from facility availability to ready-for-service: the stretch after the network hands off and before customers run workloads. Qualify racks at scale: establish firmware baselines, configure BMC and BIOS, run burn-in, and validate at node and cluster level across hundreds of racks per site on GPU and custom accelerator platforms. Drive qualification through the base-management Kubernetes platform and provisioning stack (discovery, imaging, firmware updates, shared services), burning down qualification queues with tooling rather than manual processes. Triage hardware failures found in qualification: isolate to component, drive RMA and vendor escalation, and feed failure patterns back into qualification gates. Run turn-up remotely by default, with on-site deployments of roughly a week per data hall as new halls reach facility availability, plus occasional overlapping-site engagements. Partner with network deployment, ICT, data center operations, and hardware teams during turn-up windows, and support incident response on newly deployed capacity. Ability to travel 20-30% of the time to data centers and lab environments as needed.