Senior Software Engineer, Fleet Automation and GPU Infrastructure

Teamgtn — United States · Posted ~5 hours ago

Senior Full-time Hybrid Visa History ✓

Skills

software development automation Linux distributed systems GPU infrastructure GPU Automation Distributed systems APIs

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior software engineering role focused on automation platforms for managing large-scale compute environments, combining backend development with infrastructure engineering and hardware interaction.

Highlights

Opportunity to build large-scale automation platforms for advanced compute infrastructure with relocation support available.

Description

Senior Software Engineer Location: Dallas, TX Work Model: Hybrid, 3 Days Onsite Employment Type: Direct Hire Relocation: Available for qualified non-local candidates Compensation: Competitive Base Salary + Performance Bonus Benefits: 100% Company-Paid Benefits Overview Our client is seeking a Senior Software Engineer to build the software, automation, and internal platforms used to provision, configure, monitor, and manage large-scale GPU and CPU compute infrastructure. This is a software-first engineering role operating at the intersection of backend development, infrastructure automation, Linux systems, distributed systems, and high-performance compute. The ideal candidate enjoys building production-quality services and APIs while also understanding how software interacts with physical hardware, networking, storage, and GPU infrastructure. Key Responsibilities Design and build automation platforms for provisioning, configuration, validation, and lifecycle management of GPU and CPU compute nodes.Develop backend services and APIs for hardware deployment, imaging, remediation, and decommissioning.Build reliable software using Go, C#, TypeScript, or similar backend languages.Design data models and persistent state for infrastructure automation workflows.Develop and maintain CI/CD pipelines for infrastructure and configuration changes.Automate hardware validation and testing across large compute environments.Build monitoring, observability, dashboards, and alerting using Prometheus, Grafana, Alertmanager, ELK, or similar tools.Partner with Infrastructure, Network, Operations, and Research teams to automate operational workflows.Participate in incident response, root-cause analysis, and reliability improvement efforts.Identify systemic infrastructure issues and develop software solutions that improve scalability and reliability.Required Qualifications 5+ years of software engineering experience building backend services, infrastructure platforms, or automation tooling.Strong development experience with Go, C#, TypeScript, or another modern backend language.Experience designing APIs, backend services, and distributed or stateful systems.Strong experience with relational and/or NoSQL databases.Strong Linux knowledge, including networking, storage, process management, and system troubleshooting.Experience with Ubuntu and/or RHEL environments.Experience building and maintaining CI/CD pipelines.Hands-on experience with production monitoring and observability platforms.Strong troubleshooting and problem-solving skills across software and infrastructure environments.Preferred Experience GPU, HPC, AI/ML, or large-scale compute infrastructure.NVIDIA technologies such as DCGM, nvidia-smi, or NVIDIA Container Toolkit.Bare-metal provisioning and hardware lifecycle automation.Kafka or other event-driven architectures.Experience working with infrastructure, SRE, network, or platform engineering teams.Bachelor's degree in Computer Science, Software Engineering, or equivalent practical experience.Ideal Candidate The ideal candidate is a software engineer first with strong backend development skills and enough infrastructure depth to build automation for complex, large-scale compute environments. The strongest profiles will combine software engineering, Linux systems, distributed infrastructure, automation, and observability, with GPU or HPC experience considered a strong plus.