Fleet Debugging and Incident Management Engineer

Raasinfotek — Canada · Posted ~2 days ago

Senior

Skills

Hardware debugging Firmware debugging BIOS BMC Linux Incident management Root cause analysis Python PowerShell Telemetry analysis Azure DevOps Telemetry

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary

A reliability-focused engineer is sought to investigate large-scale infrastructure incidents, analyze diagnostics, perform root cause analysis, and improve operational readiness across complex computing platforms.

Highlights

Opportunity to solve complex infrastructure issues, improve reliability, and work across hardware, software, and operations teams.

Description

Job Title: Fleet Debugging & Incident Management Engineer Key areas to look for include: Server platform validation/SDET/FW development etc (BIOS, BMC, DPU,GPU, storage, networking, server bring-up) Fleet diagnostics, root cause analysis, and issue triage Telemetry, monitoring, dashboarding, and log analysis Python/PowerShell automation and operational tooling Azure DevOps, incident management, and release validation Experience supporting server cloud infrastructure, datacenter, or large-scale deployment environments Job Summary We are seeking a proactive engineer to support large-scale fleet operations, incident management, and platform debugging activities. The role requires ownership of production issues, cross-functional coordination, root cause investigations, and executive-level reporting to ensure fleet health and operational readiness. Key Responsibilities · Own ICM triage, incident tracking, escalation, and closure. · Perform fleet debugging across hardware, firmware, BMC, BIOS, and platform software layers. · Analyze logs, telemetry, crash dumps, and diagnostics to identify failure patterns and root causes. · Coordinate with platform, validation, firmware, release, and deployment teams to drive issue resolution. · Lead operational reviews and provide status, risk, and incident trend reporting to stakeholders. · Work closely with offshore engineering teams, providing technical guidance and execution oversight. · Support release readiness assessments and fleet health monitoring activities. · Identify opportunities for automation and operational process improvements. Required Skills · Strong communication, stakeholder management, and problem-solving skills. · Strong experience in hardware and firmware debugging. · Hands-on knowledge of server platforms, BMC, BIOS, Linux, and low-level system architecture. · Experience with incident management, production support, or fleet operations. · Ability to analyze logs, telemetry, and system diagnostics to troubleshoot complex issues. · Proficiency in Python, Shell, or PowerShell scripting. Preferred Skills · Experience with DPU, GPU, accelerator, or cloud-scale server platforms. · Exposure to Azure infrastructure, data center operations, or reliability engineering. · Familiarity with ICM, monitoring systems. · Experience leading investigations across geographically distributed teams. Key Focus Areas: Fleet Debugging, ICM Management, Root Cause Analysis, Platform Reliability, Operational Reporting, Cross-Team Coordination, and Incident Resolution. -- Thanks Prashant Bansal Raas infotek corporation 262 Chapman road, Suite 105A, Newark, DE-19702 Phone: 302-565-0188 Ext: 144, Email: Prashant.bansal@raasinfotek.com Website: raasinfotek.com