Incident response

Post-Incident Review and Root Cause Analysis Guide

A practical post-incident review and root cause analysis guide for documenting what happened, why it happened, which controls failed, and how to reduce recurrence.

Part of the topicIncident response workflows

Bring intake, triage, ownership, decisions, and post-incident review together.

By Aneo B.V.Published August 31, 2026Editorial standards
Post-incident reviewRoot cause analysisSecurity incident reviewIncident response improvementCorrective actions
Direct answerPractical stepsHuman review

A post-incident review should help the organisation learn from an incident and improve the way it prevents, detects, responds to, and recovers from similar events. It should not become a blame exercise or a long narrative that nobody turns into action.

Direct answer

Run a post-incident review by confirming the incident scope, reconstructing a fact-based timeline, identifying contributing causes and control gaps, reviewing response decisions, assigning corrective actions, and checking that improvements are completed and tested.

Root cause analysis is part of the review, but it is not the only part. A useful review also examines detection, triage, ownership, communication, containment, evidence, recovery, and the decisions made while facts were incomplete.

This guide is general operational guidance, not legal, regulatory, or forensic advice. Preserve evidence and involve qualified specialists when the incident requires it.

Source and scope: NIST SP 800-61 Rev. 3 connects incident lessons with ongoing risk management and improvement. The review steps, example, and corrective-action fields below are Aneo’s suggested workflow, not a required NIST report format.

When to hold a post-incident review

Use a review after:

  • A critical or high-severity security incident
  • A confirmed compromise, data exposure, or service disruption
  • An incident that required leadership or customer escalation
  • A near miss that exposed an important control weakness
  • A repeated incident with a known pattern
  • A response that was unusually slow, confusing, or difficult to coordinate

Not every low-impact event needs a large meeting. A short review note can be enough when the learning is clear and the risk is limited.

Hold the review while records, decisions, and participant memory are still available, but after the immediate response has stabilised. The incident record should remain the source for facts and timestamps.

Set the review objective

Start by writing what the review should answer. Good objectives include:

  • How did the incident enter or develop in the environment?
  • Why did existing controls not prevent or detect it earlier?
  • What helped the team contain and recover?
  • Which decisions were delayed by missing context or unclear ownership?
  • Which changes will reduce likelihood, impact, or response time?

Avoid starting with “who made the mistake?” That question narrows the review too early and can hide the process, design, or governance conditions that allowed the mistake to matter.

Step 1: Confirm the incident facts and scope

Before discussing causes, agree on what is known. Record:

  • Detection source and first known time
  • Affected assets, accounts, services, data, and locations
  • Confirmed and suspected impact
  • Containment and recovery status
  • Relevant customer, supplier, or business dependencies
  • Open questions and confidence level

Separate facts, assumptions, and unknowns. If the scope is still being investigated, say so. A review is stronger when uncertainty remains visible than when the team creates false precision.

Step 2: Reconstruct the timeline

Use a chronological record of observable events, decisions, actions, and communications. Include the time zone and source when possible.

A useful timeline may include:

Timeline item What to record
Signal Alert, report, log event, or user observation
Validation What confirmed or challenged the initial signal
Decision Severity, ownership, escalation, or risk decision
Action Containment, enrichment, recovery, or communication
Evidence Source, timestamp, affected asset, and relevant details
Outcome What changed after the action

Do not rewrite an uncertain early assumption as if it was known at the time. Recording how understanding changed is valuable for improving triage and escalation.

Step 3: Analyse contributing causes

Security incidents rarely have one simple cause. Examine several layers:

Trigger

What event started the incident or allowed the activity to become visible?

Enabling condition

What made the trigger possible, such as excessive access, a missing update, a weak configuration, or a supplier dependency?

Detection gap

Why was the activity not detected, enriched, routed, or understood earlier?

Response gap

Why did containment, ownership, communication, or recovery take longer than expected?

Systemic condition

What process, resource, design, priority, or governance condition allowed the other factors to persist?

Use “why” questions carefully. The goal is not to keep asking why until a person is named. The goal is to find changes that are specific enough to implement and verify.

Step 4: Review controls and evidence

Connect the incident to relevant controls, policies, processes, and evidence. Ask:

  • Which control was intended to reduce this risk?
  • Was the control selected and in scope?
  • Was it implemented as defined?
  • Was it operating at the time?
  • Was evidence available and current?
  • Did the response reveal a missing control or an execution gap?

Do not label a control as failed without checking the facts. A control may have worked as designed while the risk was outside its scope, or the policy may have been correct while implementation or monitoring was incomplete.

Step 5: Review the response itself

Assess the response across the full workflow:

  • Intake and initial ticket quality
  • Triage and severity classification
  • Ownership and handoffs
  • Asset, user, and business context
  • Evidence preservation
  • Containment decisions
  • Internal and external communications
  • Recovery and validation
  • Timeline and decision logging
  • Closure criteria and follow-up

Look for friction that can be removed before the next incident. Examples include missing contact details, unclear escalation rules, duplicate tickets, unavailable logs, or approvals that require a person who is not reachable during an incident.

Step 6: Create corrective actions that can be verified

Every action should have:

  • A clear outcome
  • One accountable owner
  • A priority and target date
  • A linked risk, control, or review finding
  • A verification method
  • A follow-up date

Prefer specific actions such as “enforce MFA for the remaining administrator group and record the configuration review” over “improve access security.”

Classify actions where useful:

  • Prevent: reduce the chance of recurrence
  • Detect: improve signal quality or time to detection
  • Respond: improve triage, ownership, containment, or communication
  • Recover: improve restoration and validation
  • Govern: improve policy, training, supplier, or review discipline

Step 7: Close the learning loop

An action is not complete because a ticket was created. Confirm that the change was implemented, tested, reviewed, and connected to the relevant policy, control map, or evidence source.

Update the risk register when the incident changes the risk assessment. Update policies and procedures when the expected behaviour changes. Update training or runbooks when the response exposed a knowledge gap.

Share the learning at the right level. A useful review can improve security without disclosing unnecessary incident details to people who do not need them.

Example: post-incident review outcome

An administrator account was used from an unexpected location. MFA prevented access to one service but was not enforced on a legacy administrative path. Logging existed, but the alert did not include the asset owner, so triage was delayed.

The review identifies three contributing causes: incomplete MFA coverage, missing ownership enrichment, and an escalation path that was unclear outside business hours.

Corrective actions are assigned to close the legacy access path, enrich alerts with asset ownership, and update the on-call escalation record. Each action has an owner, target date, and verification evidence.

Common review mistakes

Blaming individuals instead of improving conditions

Accountability matters, but a review should also examine process, technology, workload, training, and governance.

Stopping at the first technical cause

“A stolen password was used” may describe the trigger. It does not explain why access, detection, or response controls did not reduce the impact.

Creating actions nobody can verify

Vague actions produce vague progress. Define what done looks like.

Ignoring successful response decisions

The review should preserve practices that helped the team, not only document failures.

Treating closure as the end of learning

Follow-up actions, evidence, and risk decisions are part of the review outcome.

Where IncidentAI fits

IncidentAI helps teams maintain structured incident records, timelines, notes, audit logs, summaries, and RCA drafts. It can organise the review starting point, while responders remain responsible for validating facts, approving conclusions, assigning actions, and making material decisions.

For the first response stage, see the security incident response first-hour checklist. For deeper RCA context, see root cause analysis for SMBs.

Frequently asked questions

What is a post-incident review?

A post-incident review is a structured assessment of what happened, how the organisation responded, what worked, what did not, and which changes should be made.

What is the difference between a post-incident review and root cause analysis?

Root cause analysis focuses on why the incident happened and progressed. A post-incident review is broader and also evaluates detection, response, decisions, communication, recovery, and improvement actions.

Should post-incident reviews assign blame?

The review should focus on facts, accountability, and improving the system. Deliberate misconduct or policy violations may require a separate process, but blame should not replace operational learning.

How long should a post-incident review take?

The effort should match severity and complexity. A critical incident may need a detailed review with several owners; a low-impact event may need a short structured record and one or two actions.

When is a corrective action complete?

It is complete when the change is implemented, verified, supported by evidence, and reflected in the relevant policy, control, process, or risk record.