Major incident management is one of the parts of service operations where mistakes have visible consequences. A poorly run major incident becomes a visible failure that affects customers, damages stakeholder confidence, and produces extended cleanup work. A well-run major incident gets resolved cleanly with minimum collateral damage.
The textbook major incident process gets you most of the way to running these well. The remaining 20% — the practical execution decisions that aren't captured in the process diagrams — is where most major incident management actually fails or succeeds. This piece is about that 20%.
The textbook process — quickly
The standard major incident process involves declaration based on severity criteria, activation of the incident response team, dedicated communication channels for the duration of the incident, structured update cadences, documented timeline of actions taken, formal closure with stakeholder confirmation, and post-incident review with documented findings.
If you're not doing these basics, fix that first. The rest of this piece won't help you if your foundational process isn't in place.
The basics are necessary but not sufficient. The execution of these steps is where major incidents actually succeed or fail. The execution requires specific judgment that the textbook doesn't always cover well.
Declare faster than feels comfortable
The single most common failure pattern in major incident management is delayed declaration. The team identifies that something is wrong but doesn't declare a major incident because they're not sure it's "really" major or because they hope it'll resolve quickly without the formal process.
By the time they actually declare, significant time has passed without coordinated response. Stakeholders are already aware of the problem through other channels. The incident has compounded because the formal coordination didn't kick in early enough.
The right discipline is to declare faster than feels comfortable. If you have any reasonable suspicion the incident might meet major criteria, declare. The cost of declaring and then de-escalating if the incident proves not to be major is much lower than the cost of delayed declaration.
Establishing this discipline requires explicit support from leadership. Service desk staff and operations engineers will not declare aggressively if they perceive that wrong declarations get them in trouble. Leadership needs to make explicit that early declaration is the desired behavior even when some declarations are subsequently de-escalated.
Get the right people in the room fast
The incident response team needs to include the people who can actually fix the problem. This is sometimes obvious and sometimes not.
The "obvious" inclusions — service desk lead, operations on-call, application owner — are usually present in incident response. The non-obvious inclusions are often missing. The people who actually wrote the code that's misbehaving. The people who built the integration that's failing. The vendor support contact who has been on the receiving end of similar issues before.
The team should expand quickly when initial responders identify that they don't have the right capability in the room. Hesitating to escalate to additional people because of organizational politeness is a common pattern that extends incident duration significantly.
The right discipline is to escalate broadly early and then narrow as the situation clarifies. It is much faster to have extra people on the bridge for the first hour and release them as you identify the actual problem space than to spend two hours getting the right people involved sequentially.
Communicate with stakeholders proactively
Stakeholder communication during major incidents is one of the most consistent areas of failure. The technical response team focuses on resolution and lets stakeholder communication slide. Stakeholders find out about the incident through other channels and become frustrated by lack of formal communication.
The right pattern is to designate a specific person responsible for stakeholder communication who is not the technical lead. This person handles the update cadence, drafts the messages, manages the stakeholder list, and keeps stakeholders informed proactively rather than reactively.
The communication should be honest about what is known and what isn't. Stakeholders generally accept honest acknowledgment of uncertainty better than vague reassurance that turns out to be wrong. "We know X is failing, we don't yet know why, here's what we're investigating, next update at [time]" is much better than "we're working on it" or premature claims that the issue is resolved.
The cadence should be predictable. Stakeholders should know when to expect the next update so they're not refreshing email constantly waiting. Even if the next update is "no significant change since last update, still investigating," the predictable cadence reduces stakeholder anxiety.
Maintain the timeline in real time
The incident timeline — what happened when, what actions were taken, what observations were made — needs to be maintained during the incident, not reconstructed afterward.
The reason isn't primarily for the post-incident review (though that matters too). It's that the timeline informs the in-flight decision-making. Patterns become visible when the timeline is captured systematically that aren't visible when it's all in people's heads. Decisions about what to try next get better when the team can see what has and hasn't been tried.
The timeline maintenance should be assigned to a specific person, often the same person handling stakeholder communication. The timeline should be on a shared surface that all responders can see. The discipline of capturing actions and observations as they happen pays back during the incident itself, not just afterward.
Make decisions explicitly
One of the failure patterns in major incident management is implicit decision-making that nobody is accountable for. The team starts going down a particular response path because someone suggested it, but no one has explicitly decided that this is the path. When the path doesn't work, no one is accountable for re-evaluating because no one explicitly chose it.
Better practice is to make decisions explicitly with named owners. "We're going to try X, owned by Y, expected outcome by Z, escalate to alternative approach if we don't see progress by then." This frames the decision so the team can evaluate it after the timebox and explicitly choose to continue, modify, or abandon.
This explicit decision-making feels formal during a stressful incident but it's actually faster than implicit decision-making because it prevents the team from spinning on approaches that aren't working without explicit reconsideration.
Know when to ask for help
Some major incidents exceed the response team's capability. The team needs to recognize this and escalate to additional resources rather than continuing to flounder with insufficient capability.
The escalation paths include vendor support, professional services contacts, internal experts not currently engaged, and in serious cases external incident response capabilities. Knowing what these paths are before you need them is part of major incident preparedness. Activating them when needed is a judgment call that requires recognizing that the current approach isn't producing progress.
The discipline failure here is that teams sometimes feel that escalating outside the immediate response group is admitting failure. It's not. Major incidents that exceed the response team's capability are a normal occurrence. Recognizing this and escalating appropriately is professional behavior, not failure.
Close cleanly
The incident resolution moment is not the end of the major incident process. The closure activities matter and are often handled poorly.
Closure should involve confirmation from stakeholders that the issue is actually resolved from their perspective. Internal team confirmation is necessary but not sufficient. Sometimes the technical fix doesn't fully address the customer experience and the closure based on technical resolution misses this.
Closure should include explicit handover of any monitoring or follow-up activity. If something needs to be watched for the next 24 hours to confirm stability, someone owns that watching. If a permanent fix is pending and the incident was resolved with a workaround, the path to permanent fix is owned and tracked.
Closure should include scheduling the post-incident review and identifying the participants. The review should happen within a few days while the details are still fresh. Delaying the review until "things calm down" usually means it doesn't happen with the rigor it warrants.
Run the post-incident review well
The post-incident review is where the organization actually learns from major incidents. Done well, it produces specific actionable improvements. Done poorly, it produces blame distribution and defensiveness with no actionable outcomes.
The review should be blameless in posture but should still produce honest analysis of what happened. The blameless framing isn't about avoiding hard truths; it's about creating conditions where hard truths can be discussed without people becoming defensive about their roles in the incident.
The review should produce specific actions with owners and timelines. "We should improve monitoring" is not an action. "Add specific monitoring for X by Y date, owned by Z, validated through specific test by date W" is an action. Actions without specificity rarely happen.
The review findings should be shared more broadly than the immediate response team. Other teams in the organization can learn from the patterns that the incident exposed even if they weren't directly involved in the response.
What this looks like over time
Organizations that run major incidents well tend to have specific characteristics. They have practiced the process enough that the basics are automatic. They have leadership culture that supports early declaration and broad escalation. They have communication discipline that keeps stakeholders informed proactively. They have post-incident learning culture that actually changes behavior across incidents.
These characteristics aren't accidents. They develop through deliberate practice and through accumulating experience across multiple incidents. Organizations that handle each incident as a one-off rarely develop the muscle memory that makes major incident management consistently effective.
The investment in major incident capability pays back across the years. The cost of one badly-run major incident — in customer impact, stakeholder confidence, internal frustration, and cleanup work — typically exceeds the cost of years of investment in major incident process and training. The math favors investment, even though the immediate returns aren't always visible.
The 20% that determines whether major incidents go cleanly or badly is mostly about the practical execution decisions described above. The textbook process gets you to 80%. The 20% is the work that actually matters in the moments when major incidents happen.