Summary
✨ AI‑Generated
An advanced computing organization is seeking an HPC platform engineer to build and operate infrastructure supporting large-scale workloads across compute, storage, and networking environments.
Highlights
High-impact infrastructure role with competitive compensation, benefits, and work on advanced computing systems.
Description
Job Title: HPC Platform Engineer
Industry: High Performance Computing / AI Infrastructure
Location (city, state): Dallas, TX
Assignment Type: Direct hire
Pay: $180,000–$260,000 base salary, plus a potential $50,000–$100,000 bonus
Work Schedule: Hybrid; three days in the Dallas office and two days remote.
The team manager determines the in-office schedule.
Benefits: This position is eligible for medical, dental, vision, and 401(k).
Additional benefits include 25 days of PTO, an HSA contribution, a gym membership, and lunch on office days.
About The Company:
Our client develops advanced computing and cloud infrastructure for demanding AI, research, and simulation workloads.
The organization is investing in the systems and engineering teams needed to expand its computing capacity.
Job Description:
We are seeking an HPC Platform Engineer to build, operate, and improve the infrastructure behind large-scale computing workloads.
This person will work across compute, storage, networking, and platform software, taking ownership from design and deployment through ongoing support and capacity planning.
The role also provides an opportunity to guide technical standards and mentor other engineers.
Key Responsibilities:
Design and implement HPC infrastructure across servers, storage, networking, and related data center systems.Find performance bottlenecks and improve system throughput, reliability, and resource use.Partner with engineering, operations, and research teams to install, configure, test, and support new systems.Use monitoring data to troubleshoot complex issues and plan for future capacity needs.Assess new technologies and vendor solutions, and recommend improvements to the platform.Document architectures, configurations, and operating procedures while helping junior engineers develop their skills.
Qualifications:
Five or more years of HPC engineering or closely related infrastructure experience.Working knowledge of parallel computing, distributed storage, workload scheduling, high-speed networking, and GPU-based systems.Experience with relevant technologies such as Slurm or PBS Pro; Lustre or GPFS; InfiniBand; OpenStack; and Docker or Kubernetes.Ability to automate infrastructure work with Python and tools such as Ansible, Puppet, or Chef.Experience using monitoring tools to diagnose and improve platform performance.Strong troubleshooting, documentation, and communication skills.Bachelor’s degree or equivalent practical experience.
Experience with ZFS, NiFi, or computational fluid dynamics workloads is a plus.
Additional Details:
Relocation assistance may be tailored to the candidate.
TN visa candidates may be considered.
The interview process is expected to include an HR conversation, a hiring manager meeting, a technical discussion focused on prior experience and concepts, and an onsite visit.