Summary
A reliability-focused engineer is sought to investigate large-scale infrastructure incidents, analyze diagnostics, perform root cause analysis, and improve operational readiness across complex computing platforms.
Highlights
Opportunity to solve complex infrastructure issues, improve reliability, and work across hardware, software, and operations teams.
Description
Job Title: Fleet Debugging & Incident Management Engineer
Key areas to look for include:
Server platform validation/SDET/FW development etc (BIOS, BMC, DPU,GPU, storage, networking, server bring-up)
Fleet diagnostics, root cause analysis, and issue triage
Telemetry, monitoring, dashboarding, and log analysis
Python/PowerShell automation and operational tooling
Azure DevOps, incident management, and release validation
Experience supporting server cloud infrastructure, datacenter, or large-scale deployment environments
Job Summary
We are seeking a proactive engineer to support large-scale fleet operations, incident management, and platform debugging activities.
The role requires ownership of production issues, cross-functional coordination, root cause investigations, and executive-level reporting to ensure fleet health and operational readiness.
Key Responsibilities
· Own ICM triage, incident tracking, escalation, and closure.
· Perform fleet debugging across hardware, firmware, BMC, BIOS, and platform software layers.
· Analyze logs, telemetry, crash dumps, and diagnostics to identify failure patterns and root causes.
· Coordinate with platform, validation, firmware, release, and deployment teams to drive issue resolution.
· Lead operational reviews and provide status, risk, and incident trend reporting to stakeholders.
· Work closely with offshore engineering teams, providing technical guidance and execution oversight.
· Support release readiness assessments and fleet health monitoring activities.
· Identify opportunities for automation and operational process improvements.
Required Skills
· Strong communication, stakeholder management, and problem-solving skills.
· Strong experience in hardware and firmware debugging.
· Hands-on knowledge of server platforms, BMC, BIOS, Linux, and low-level system architecture.
· Experience with incident management, production support, or fleet operations.
· Ability to analyze logs, telemetry, and system diagnostics to troubleshoot complex issues.
· Proficiency in Python, Shell, or PowerShell scripting.
Preferred Skills
· Experience with DPU, GPU, accelerator, or cloud-scale server platforms.
· Exposure to Azure infrastructure, data center operations, or reliability engineering.
· Familiarity with ICM, monitoring systems.
· Experience leading investigations across geographically distributed teams.
Key Focus Areas: Fleet Debugging, ICM Management, Root Cause Analysis, Platform Reliability, Operational Reporting, Cross-Team Coordination, and Incident Resolution.
--
Thanks
Prashant Bansal
Raas infotek corporation
262 Chapman road, Suite 105A, Newark, DE-19702
Phone: 302-565-0188 Ext: 144,
Email: Prashant.bansal@raasinfotek.com
Website: raasinfotek.com