Workarounds, known errors and problem control

The three phases of problem management and where the known error sits between the last two, what a workaround is, and how an organisation is expected to behave when it has more problems than it can analyse.

Lesson 4 of 12 in objective 7. Seven ITIL practices in detail, part of ITIL 4 Foundation.

The three phases of problem management, and the one that keeps going. A chain of 4 steps whose last one repeats on its own: Problem identification (from incident trends, and proactively), then Problem control (analyse it, work out the cause, find a workaround), then A known error exists (analysed and not resolved. The boundary marker), then Error control (ongoing management of the known error, no longer investigation). Then an arrow from Error control round to itself: Re-examining the known error's impact, the cost of the incidents it keeps causing and whether a permanent fix is now worthwhile is error control running again, not the work going back to analysis. Problem identification from incident trends, and proactively Problem control analyse it, work out the cause, find a workaround A known error exists analysed and not resolved. The boundary marker Error control ongoing management of the known error, no longer investigation Re-examining the known error's impact, the cost of the incidents it keeps causing and whether a permanent fix is now worthwhile is error control running again, not the work going back to analysis
The three phases of problem management, and the one that keeps going.

Workarounds

A workaround reduces or removes the impact of an incident or a problem for which no full resolution is yet available. It is a solution — something you do — rather than a record, and it can be permanent in practice: a workaround that is never replaced by a fix is still a workaround.

Workarounds sit inside the purpose of problem management rather than being incidental to it, and they are documented against the problem or known error so the service desk can apply them again without rediscovering them.

Control, error control, and the line between them

Problem control is the analysis phase: understanding the problem, working out the cause, and finding a workaround where a permanent fix is not available. Its output is understanding.

Error control is what happens afterwards. Once a problem has been analysed and not resolved, it is a known error, and error control manages that known error over time — periodically re-examining its impact, the cost of the incidents it keeps causing, and whether a permanent fix has now become worthwhile. A team doing that periodic re-examination is in error control, not in problem control, and the question usually describes the re-examination without naming it.

The reason the phases are worth separating is that the known error is the boundary marker between them. Analysis finishes, the known error exists, and the work changes character from investigation to ongoing management.

When there are more problems than capacity

The practice does not pretend every problem gets solved. Problems are prioritised for analysis according to the RISK they pose, and it is explicitly accepted that not all of them will be resolved — some are left with a workaround and revisited, and some are simply not worth the cost of fixing.

That is worth remembering because the wrong options here are the responsible-sounding ones: analyse everything, resolve every problem before accepting new ones, hire more people. The examinable position is risk-based prioritisation plus acceptance that resolution is not universal.

The stated position on capacity, next to the responsible-sounding wrong answers. The practice — Which problems get analysed: The riskiest first; What happens to the rest: A workaround and a revisit, or left alone; Is every problem resolved?: No, and that is the stated position. The wrong options — Which problems get analysed: All of them; What happens to the rest: Hire until nothing waits; Is every problem resolved?: Yes, before new ones are accepted The practice The wrong options Which problems get analysed The riskiest first All of them What happens to the rest A workaround and a revisit, or left alone Hire until nothing waits Is every problem resolved? No, and that is the stated position Yes, before new ones are accepted
The stated position on capacity, next to the responsible-sounding wrong answers.

Where the incident practice hands over

The recurring scenario: an application fails for the third time this month, the service desk restarts it, users are working again in ten minutes, and nobody knows why it keeps failing. Every one of those failures is an incident, and incident management has done its job correctly each time — service was restored quickly.

What has not happened is any handling of the repeated cause, and that belongs to problem management. Both practices are involved, they run as separate records with separate lifecycles, and neither replaces the other — an incident is not held open until its cause is understood, and a problem does not close the incidents it explains.

One fault that keeps returning, and the records each practice opens for it. One fault, failing three times contains Incident management (restores service each time, correctly), Three incident records (one per failure, not held open for the cause), Problem management (owns the repeated cause), One problem record (separate lifecycle, closes no incident). One fault, failing three times Incident management restores service each time, correctly Three incident records one per failure, not held open for the cause Problem management owns the repeated cause One problem record separate lifecycle, closes no incident
One fault that keeps returning, and the records each practice opens for it.

Worth carrying in

Workaround
Reduces or removes impact while no full resolution exists. A solution, not a record.
Problem control
Analysis: understand the problem, find the cause, find a workaround.
Error control
Manage the known error over time; revisit whether a permanent fix is now worth it.
Risk-based prioritisation
Problems are analysed in risk order, and not all will be resolved.
Separate lifecycles
Incident records and problem records coexist. Neither replaces the other.

What the exam does with this

Objective
7. Seven ITIL practices in detail
Share of the exam
47.5% (the whole objective)
Questions in this lesson
4
Signed for by a person
0

Partly checked. None of the 4 questions here has been read against the cited source by a person. 4 questions have been checked against their cited clause by an automated pass — which is not the same thing, and is not a signature.

Only questions a person has signed for are used in mock exams here. That is the whole difference between the two kinds of checking above.

How these questions are written — where each question comes from, what the verification ledger records, and what happens when one is found wrong.

Drill this lesson

A lesson is one sitting: the trainer draws a short run from these questions alone and spaces the ones you get wrong.

Practise Workarounds, known errors and problem control

Questions in this lesson

Practise Workarounds, known errors and problem control

The rest of objective 7