DevOps Engineer

Imcsgroupofficial — United States · Posted ~1 day ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

Role-DevOps Engineer Location-Philadelphia, PA 19103 Role Summary The Analyst IV is the senior-most technical expert in Reliability Operations. This role provides strategic leadership in major incident response, enterprise problem management, and change governance, while influencing reliability architecture and cross‑org operational excellence. Key Responsibilities Incident Management Act as senior Incident Commander for the most complex, cross‑platform P1/P0 outages.Provide executive-level communication and decision-making support.Guide engineering teams through crisis diagnostics and service recovery.Own and improve enterprise-wide incident command processes.Problem Management Lead systemic, long-term problem investigations across multiple services.Drive root cause elimination strategies at engineering and architectural levels.Influence roadmaps to incorporate reliability enhancements.Provide executive reviews of major problem investigations.Change Management Serve as reliability advisor for high‑risk changes across the organization.Lead or oversee major enterprise-level change events (e.g., data center cutovers, platform migrations).Shape governance and guardrails for operationally safe change delivery.Mentor teams on designing resilient, low-risk change mechanisms.Reliability Strategy & Leadership Champion reliability best practices across engineering, SRE, and operations.Influence architecture, monitoring frameworks, observability standards, and automation investments.Mentor all analyst tiers and serve as a technical authority for Ops.Collaborate with engineering leadership on long-term reliability goals.Lead multi-team efforts to improve system availability, performance, and scalability.Own and improve incident management processes, tooling, and standards.Architect improvements to monitoring, alerting, and reliability frameworksContribute to Automation & AI Ideas for improving efficiency & reduce MTTM /MTTRTechnical Leadership Serve as a technical authority on service reliability and operational best practices.Review system architectures for reliability risks.Guide automation, observability, and self-healing strategies across the org.Coaching & Organizational Influence Mentor all analyst tiers and contribute to career development pathways.Partner with SRE Managers, Engineering Ops, and Architecture teams on roadmap planning.Lead training on incident handling, tooling, and operational maturity. Required Qualifications 5+ years in SRE, Incident Management, or high‑level Operations Engineering roles.Deep experience managing complex distributed systems and cloud infrastructure.Proven leadership during high-severity, high-impact outages.Ability to influence architecture and reliability decisions at scale. Preferred Qualifications Strong automation and scripting skills.Experience in designing reliable, scalable architectures.Certifications in cloud, ITIL, SRE, or DevOps disciplines.