Network Software Engineer - AI Data Centers

Intelletec Energy — United States · Posted ~3 hours ago

Skills

Network software engineering Network diagnostics Telemetry Monitoring and alerting Automation Distributed systems Network infrastructure Network automation Monitoring Alerting Switches Routers NICs

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A network software engineering role focused on operating rapidly expanding AI data-center infrastructure at scale. You will build diagnostic and telemetry tooling across network and host connectivity, automate hardware recovery from detection through restoration, and establish monitoring, alerting, and health-management systems that make large infrastructure fleets easier to operate.

Highlights

Build software systems for large-scale AI data-center infrastructure, including diagnostics, telemetry, automated hardware recovery, monitoring, and operational visibility. The role offers substantial ownership over systems that improve reliability and reduce manual operational work across a rapidly expanding infrastructure fleet.

Description

Our client is a fast-growing hyperscale Data Center Operator building infrastructure to support top AI labs and enterprises. DataCenter Network Production Engineer What You'll Help Build Make network incidents diagnosable at scale. Create tooling that connects telemetry and diagnostics across switches, routers, NICs, and host-to-network paths. Enable engineers to run commands remotely across the infrastructure and quickly visualize failures, dependencies, and likely causes.Create an automated hardware recovery workflow. Build the systems that take an issue from initial detection through troubleshooting, ticket creation, hardware replacement, and restoration of service. The goal is to standardize recovery across data center fabrics, edge infrastructure, and server connectivity without relying on manual coordination.Develop the operational visibility layer for a rapidly expanding fleet. Establish the monitoring, alerting, and health-management infrastructure needed to operate multiple large-scale facilities today and a dramatically larger footprint tomorrow. You'll be building core capabilities that don't yet exist. What You'll Own Develop software in Go to eliminate repetitive troubleshooting work. This includes automated connectivity checks, fleet-wide remote execution, diagnostic workflows, and tools that make complex network failures easier to understand.Build and improve the complete equipment-recovery lifecycle, connecting automated fault detection with service tickets, replacement workflows, component tracking, and final restoration.Operate and evolve the monitoring and alerting systems that provide engineers with an accurate, real-time view of infrastructure health across every facility.Bring new facilities and network equipment into production by developing and executing validation procedures that establish reliability before systems begin carrying live workloads. Salary & Benefits: $225-275k base salary + meaningful equity.Great health insuranceRelocation Assistance3 Days/week Hybrid schedule