The Seven Stages of Problem Management — A Complete Guide to the Process
Overview of the Problem Management Lifecycle
The Problem Management lifecycle ensures that problems are systematically identified, investigated, and resolved. Each stage has specific activities, inputs, and outputs.
Stage 1: Problem Detection
Problems can be detected through multiple channels:
- Major incidents: When a major incident occurs, a problem should always be raised
- Recurring incidents: Patterns of similar incidents indicate an underlying problem
- Monitoring events: Alerts that suggest an underlying issue
- Trend analysis: Patterns identified through data analysis
- Vendor reports: Suppliers reporting known issues
- Technical support staff: Internal experts identifying issues
Key question: Is this a one-time incident or a pattern indicating a problem?
Stage 2: Problem Logging
Once detected, the problem must be properly logged:
Core fields :
|
Field |
Description |
|
Subject |
Title or short summary |
|
Description |
Detailed description including actual behavior, expected behavior, and steps to reproduce |
|
Resources |
Devices on which the problem is identified |
|
Category |
Category to which the problem is mapped |
|
Sub Category |
Subcategory under the category |
|
Priority |
How soon the problem needs to be fixed |
Additional fields:
- Requested By: User who requested the problem
- Assignee Group: User group that manages the problem
- Assign to: Specific user
- Application: Applications where the problem is detected
- Root Cause: Factors on resolution of which incidents can be prevented
- Work Around: Temporary method for achieving the task
Stage 3: Problem Categorization
Categorization helps with reporting, routing, and analysis:
- Product categorization: The IT asset or system affected
- Operational categorization: The action or function required
- Component/s: Segments of IT infrastructure related to the problem
Good categorization ensures consistency and enables meaningful reporting.
Stage 4: Problem Prioritization
Priority is based on :
- Impact: The effect of the problem, usually in regards to service level agreements
- Urgency: The time available before the business feels the problem's impact
- Frequency: How often related incidents occur
The goal is to prioritize problems that have the greatest business impact.
Stage 5: Problem Investigation and Diagnosis
This is the root cause analysis phase, where the team determines the underlying cause of the problem:
- Root cause analysis: Using techniques like Five Whys, Ishikawa diagrams, and Pareto analysis
- Investigation reason: The trigger for prompting an investigation (e.g., recurring incidents, non-routine incidents)
- Cross-functional collaboration: Involving subject matter experts
Key question: What is the underlying cause that, if fixed, would prevent recurrence?
Stage 6: Creating a Known Error Record
When the root cause is identified and a workaround is developed, the problem becomes a "known error" :
- Known Error documentation: Symptoms of related incidents, root cause, and workarounds
- Knowledge Base integration: Known errors are added to the Knowledge Base
- Incident linking: All related incidents are linked to the known error
This is where problem management creates lasting value—by ensuring that when the same issue occurs again, it can be resolved faster.
Stage 7: Problem Resolution and Closure
The final stage involves implementing a permanent fix:
- Propose a change: The service desk team proposes a change to the infrastructure to resolve the problem
- Implement the fix: Through change management
- Verify resolution: Confirm the problem is resolved
- Close the problem: After verification and documentation
Testing the fix: Ensure it doesn't introduce new problems.
Stage 8: Major Problem Review
Team members should carry out in-depth reviews of major problems :
- Lessons learned: What worked, what didn't
- Preventive measures: How to prevent similar problems
- Process improvements: How to improve the problem management process itself
- Action items: Follow-up activities
Visualizing the Lifecycle
The problem management workflow should complement these ITIL-recommended activities :
- Problem investigation
- Identification of workarounds
- Recording of known errors
Start with the default workflow and adapt it to your specific needs over time.
Conclusion
The Problem Management lifecycle provides a systematic approach to identifying, investigating, and resolving problems. By following each stage, organizations can ensure that recurring incidents are eliminated and service quality improves over time.
Action Items for Your Organization
- Map your current problem management process against the lifecycle
- Identify gaps in your process
- Document the workflow in your ITSM platform
- Train your team on each stage of the lifecycle
- Measure completion time for each stage