DevOps Engineer
Imcsgroupofficial — United States · Posted ~1 day ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
Role-DevOps Engineer
Location-Philadelphia, PA 19103
Role Summary
The Analyst IV is the senior-most technical expert in Reliability Operations.
This role provides strategic leadership in major incident response, enterprise problem management, and change governance, while influencing reliability architecture and cross‑org operational excellence.
Key Responsibilities
Incident Management
Act as senior Incident Commander for the most complex, cross‑platform P1/P0 outages.Provide executive-level communication and decision-making support.Guide engineering teams through crisis diagnostics and service recovery.Own and improve enterprise-wide incident command processes.Problem Management
Lead systemic, long-term problem investigations across multiple services.Drive root cause elimination strategies at engineering and architectural levels.Influence roadmaps to incorporate reliability enhancements.Provide executive reviews of major problem investigations.Change Management
Serve as reliability advisor for high‑risk changes across the organization.Lead or oversee major enterprise-level change events (e.g., data center cutovers, platform migrations).Shape governance and guardrails for operationally safe change delivery.Mentor teams on designing resilient, low-risk change mechanisms.Reliability Strategy & Leadership
Champion reliability best practices across engineering, SRE, and operations.Influence architecture, monitoring frameworks, observability standards, and automation investments.Mentor all analyst tiers and serve as a technical authority for Ops.Collaborate with engineering leadership on long-term reliability goals.Lead multi-team efforts to improve system availability, performance, and scalability.Own and improve incident management processes, tooling, and standards.Architect improvements to monitoring, alerting, and reliability frameworksContribute to Automation & AI Ideas for improving efficiency & reduce MTTM /MTTRTechnical Leadership
Serve as a technical authority on service reliability and operational best practices.Review system architectures for reliability risks.Guide automation, observability, and self-healing strategies across the org.Coaching & Organizational Influence
Mentor all analyst tiers and contribute to career development pathways.Partner with SRE Managers, Engineering Ops, and Architecture teams on roadmap planning.Lead training on incident handling, tooling, and operational maturity.
Required Qualifications
5+ years in SRE, Incident Management, or high‑level Operations Engineering roles.Deep experience managing complex distributed systems and cloud infrastructure.Proven leadership during high-severity, high-impact outages.Ability to influence architecture and reliability decisions at scale.
Preferred Qualifications
Strong automation and scripting skills.Experience in designing reliable, scalable architectures.Certifications in cloud, ITIL, SRE, or DevOps disciplines.
We have 66,600 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 66,600 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume