Incident management, priority and major incidents

Incident management in detail: what restoring service quickly actually means, how the order of work is decided when several incidents are open at once, and why major incidents get a procedure of their own.

Lesson 1 of 12 in objective 7. Seven ITIL practices in detail, part of ITIL 4 Foundation.

Why major incidents get a separate procedure at all. Ordinary incident — Timescale: The agreed target for its priority; Who works it: The team the category routes it to; How it is run: The normal procedure. Major incident — Timescale: Shorter, with greater urgency; Who works it: Often a dedicated temporary team; How it is run: A separate, pre-agreed procedure Ordinary incident Major incident Timescale The agreed target for its priority Shorter, with greater urgency Who works it The team the category routes it to Often a dedicated temporary team How it is run The normal procedure A separate, pre-agreed procedure
Why major incidents get a separate procedure at all.

What the practice is trying to achieve

Incident management exists to keep the damage an incident does as small as it can be, by getting normal operation back quickly. Speed of RESTORATION is the objective, and understanding the cause is not part of it. A workaround that gets users working again, with the cause still unknown, is a completely successful incident resolution, and any option that makes root cause a condition of closing an incident belongs to a different practice.

The definition to keep alongside it: an incident is an interruption nobody planned, or a fall in the quality of a service. Degradation counts, which is why "the service is much slower than usual" is an incident and not a request.

Deciding what to work on first

When several incidents are open, the order is decided by an agreed classification scheme so that the ones with the highest business impact are worked first. The two words carrying the weight are AGREED and IMPACT. Not first-come-first-served, not whichever user shouts loudest, not the one that looks technically most interesting — a scheme settled in advance, applied consistently.

That is also why the categorisation applied when an incident is logged matters beyond tidiness: it is what feeds the prioritisation and, separately, what routes the incident to the right team.

One categorisation is read twice: it routes the incident, and it feeds the priority. Left column, What feeds the decision; right column, The decision made. The category of the incident points at The team it is routed to (The category names the right team). The category of the incident and Business impact, weighed by the agreed scheme both point at Its place in the order (Never arrival order, and never whoever shouts loudest). What feeds the decision The decision made The category of the incident The team it is routed to The category names the right team The category of the incident Business impact, weighed by the agreed scheme Its place in the order Never arrival order, and never whoever shouts loudest
One categorisation is read twice: it routes the incident, and it feeds the priority.

Swarming, and major incidents

Swarming is the technique where several specialists from different teams work an incident together at the same time, and once the right person to carry on has become obvious, the others step away. It is worth knowing by name and worth contrasting with tiered escalation, which passes an incident from one level to the next in sequence. Swarming brings the expertise to the incident; escalation moves the incident to the expertise.

Major incidents get a separate procedure because they need shorter timescales and greater urgency than the normal path can deliver, and they are often handled by a dedicated temporary team assembled for the purpose. Note the reason the exam wants: it is about urgency and timescale, not about the incident being technically harder, and the separate procedure is agreed in advance rather than invented during the outage.

Both get the expert onto the incident. They move opposite things. Swarming — What moves: The people, to the incident; How many at once: Several, from different teams; The incident is: Worked by all of them, then handed to one. Tiered escalation — What moves: The incident, to the people; How many at once: One level at a time; The incident is: Passed up in sequence Swarming Tiered escalation What moves The people, to the incident The incident, to the people How many at once Several, from different teams One level at a time The incident is Worked by all of them, then handed to one Passed up in sequence
Both get the expert onto the incident. They move opposite things.

Worth carrying in

Incident management
Damage kept small by getting normal operation back fast.
Incident
An interruption nobody planned, or a fall in service quality.
Prioritisation
By an agreed classification scheme, highest business impact first.
Swarming
Several specialists work it together, then all but one step away.
Major incident
Shorter timescales, greater urgency, often a dedicated temporary team.

What the exam does with this

Objective
7. Seven ITIL practices in detail
Share of the exam
47.5% (the whole objective)
Questions in this lesson
5
Signed for by a person
0

Partly checked. None of the 5 questions here has been read against the cited source by a person. 5 questions have been checked against their cited clause by an automated pass — which is not the same thing, and is not a signature.

Only questions a person has signed for are used in mock exams here. That is the whole difference between the two kinds of checking above.

How these questions are written — where each question comes from, what the verification ledger records, and what happens when one is found wrong.

Drill this lesson

A lesson is one sitting: the trainer draws a short run from these questions alone and spaces the ones you get wrong.

Practise Incident management, priority and major incidents

Questions in this lesson

Practise Incident management, priority and major incidents

The rest of objective 7