Engineer – Incident & Problem Management
Malomatia — Qatar · Posted ~3 days ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
Job Description
To support effective operation, coordination, control, and continual improvement of Incident Management, Major Incident Management, and Reactive and Proactive Problem Management across in-scope services.
The role helps ensure that incidents are accurately recorded, prioritized, escalated, communicated, restored, validated, and closed within agreed targets, and that recurring or high-impact issues are investigated and reduced through structured root-cause analysis, known-error management, and corrective action.
Working under the guidance of the Practice Manager or Senior Consultant, the role contributes to service reliability, customer experience, operational resilience, auditability, and continual learning.
Responsibilities
RESPONSIBILITIES
Support the day-to-day administration and consistent execution of Incident Management, Major Incident Management, and Problem Management procedures, workflows, templates, escalation models, and controls.
Review incident records for accurate logging, categorization, impact, urgency, priority, assignment, timestamps, service and configuration-item links, diagnostic information, resolution details, user validation, and closure quality.
Monitor high-priority, aging, breached, reopened, customer-sensitive, and supplier-dependent incidents; follow up with assigned teams and escalate risks or delays through approved paths.
Support major-incident declaration, bridge-call coordination, participant activation, timeline and action logging, restoration tracking, stakeholder communications, and formal stand-down under the direction of the Major Incident Manager or Practice Lead.
Coordinate service desk, technical, cloud, infrastructure, application, cybersecurity, supplier, customer, and business teams to support rapid and safe restoration, ensuring workarounds, emergency changes, failover, and recovery actions are authorized and recorded.
Collect logs, monitoring information, decisions, impact details, timelines, and supporting evidence for major-incident reports, post-incident reviews, customer reporting, and audits; document lessons learned and track actions.
Identify and register problem candidates from major and recurring incidents, trend and monitoring data, failed changes, post-release issues, supplier performance, availability or capacity concerns, and known operational risks.
Support structured root-cause analysis using approved techniques, document causes and contributing factors, and track corrective and preventive actions to verified closure.
Maintain problem backlogs, known errors, workarounds, temporary fixes, permanent-resolution actions, owners, target dates, risks, dependencies, and links among incident, problem, change, knowledge, supplier, service, and configuration-item records.
Ensure verified workarounds and known errors are published to service desk and support teams, remain current, and are used to improve first-line resolution.
Produce operational reports and dashboards covering incident trends, SLA performance, response and restoration times, backlog age, reopen rate, major and repeat incidents, problem status, known-error age, and action closure.
Prepare meeting packs, minutes, decision records, and action logs for major-incident reviews, problem reviews, backlog meetings, supplier discussions, and service-improvement forums.
Support monitoring, event-management, automation, runbook, knowledge, shift-left, swarming, and self-healing improvements that reduce detection, diagnosis, restoration, and recurrence times.
Maintain procedures, priority models, communication templates, escalation lists, major-incident checklists, root-cause templates, known-error content, and knowledge articles under document control.
Assist with record-quality reviews, process compliance checks, supplier follow-up, customer or audit evidence preparation, and closure of identified gaps or corrective actions.
Identify recurring process weaknesses, data-quality issues, bottlenecks, supplier delays, and automation opportunities, and contribute to practical continual-improvement initiatives.
Required Education, Knowledge, And Certifications
Bachelor’s degree in information technology, computer science, software engineering, engineering, information systems, or a related discipline.
ITIL 4 Foundation certification is required; relevant training in Incident Management, Major Incident Management, Problem Management, Monitor, Support and Fulfil, or Create, Deliver and Support is strongly preferred.
Training in Site Reliability Engineering, SIAM, COBIT, ISO/IEC 20000, Lean Six Sigma, risk management, business continuity, project management, or information security is advantageous.
Good working knowledge of IT service management, service operation and transition, service reliability, operational resilience, customer experience, service levels, supplier obligations, and continual improvement.
Required Hard Skills
Working knowledge of Incident, Major Incident, and Problem Management lifecycles, roles, priority models, escalation paths, communications, governance forums, and record requirements.
Ability to review incident and problem records, identify missing or inconsistent data, monitor SLA and aging status, maintain traceability, and follow approved escalation and closure procedures.
Practical capability in major-incident coordination, bridge-call administration, timeline and action logging, stakeholder updates, restoration tracking, and post-incident documentation.
Ability to support root-cause analysis, distinguish symptoms from causes, document contributing factors, and track corrective and preventive actions to verified closure.
Working knowledge of enterprise ITSM platforms, monitoring and observability tools, knowledge bases, dashboards, workflow automation, CMDB, service mapping, and audit trails.
Understanding of infrastructure, cloud, networks, telecommunications, cybersecurity, databases, middleware, applications, and end-user computing sufficient to coordinate multi-team support.
Ability to prepare operational dashboards, trend reports, incident summaries, root-cause documents, meeting minutes, management presentations, and auditable supporting evidence.
Working knowledge of analytical techniques such as Pareto analysis, 5 Whys, fishbone analysis, chronology analysis, causal mapping, and basic data analysis.
Qualifications
Required Soft Skills:
Clear verbal and written communication, with the ability to translate technical events into understandable business impact, status, risk, decisions, and actions.
Good facilitation, coordination, follow-up, and stakeholder-management skills across technical teams, customers, service managers, and suppliers.
Strong planning, prioritization, organizational discipline, and ability to manage multiple incidents, problems, actions, and deadlines.
Analytical thinking, focused questioning, attention to detail, and ability to work with incomplete or rapidly changing information.
Customer-focused collaboration, expectation setting, empathetic communication, prompt escalation, and disciplined action follow-through.
Behavioral
Remains calm, structured, responsive, and credible during major incidents, service degradation, customer pressure, or competing priorities.
Demonstrates ownership, accountability, reliability, integrity, confidentiality, and consistent adherence to approved procedures.
Balances rapid restoration with customer impact, technical risk, security, compliance, resilience, and long-term prevention.
Shows curiosity, persistence, critical thinking, adaptability, and willingness to investigate recurring issues rather than accept temporary workarounds as permanent solutions.
Promotes collaboration, swarming, transparency, knowledge sharing, psychological safety, and a no-blame learning culture while maintaining accountability.
Experience
4–6 years of relevant experience in IT service operations, Incident Management, Major Incident Management, Problem Management, service desk, service reliability, or a comparable function.
At least 2 years of practical experience coordinating priority incidents, supporting major incidents, maintaining problem records, assisting root-cause analysis, or tracking corrective actions.
Experience using enterprise ITSM platforms, monitoring or observability tools, dashboards, knowledge bases, CMDB or service-mapping information, and workflow automation.
Experience working with cross-functional infrastructure, cloud, application, network, cybersecurity, service desk, customer, and supplier teams in operational or 24x7 environments.
Experience supporting operational reporting, customer communications, audits, post-incident reviews, problem backlogs, known-error management, or continual-improvement activities is advantageous.
We have 66,600 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 66,600 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume