A post-incident review should help the organisation learn from an incident and improve the way it prevents, detects, responds to, and recovers from similar events. It should not become a blame exercise or a long narrative that nobody turns into action.
Direct answer
Run a post-incident review by confirming the incident scope, reconstructing a fact-based timeline, identifying contributing causes and control gaps, reviewing response decisions, assigning corrective actions, and checking that improvements are completed and tested.
Root cause analysis is part of the review, but it is not the only part. A useful review also examines detection, triage, ownership, communication, containment, evidence, recovery, and the decisions made while facts were incomplete.
This guide is general operational guidance, not legal, regulatory, or forensic advice. Preserve evidence and involve qualified specialists when the incident requires it.
Source and scope: NIST SP 800-61 Rev. 3 connects incident lessons with ongoing risk management and improvement. The review steps, example, and corrective-action fields below are Aneo’s suggested workflow, not a required NIST report format.
When to hold a post-incident review
Use a review after:
- A critical or high-severity security incident
- A confirmed compromise, data exposure, or service disruption
- An incident that required leadership or customer escalation
- A near miss that exposed an important control weakness
- A repeated incident with a known pattern
- A response that was unusually slow, confusing, or difficult to coordinate
Not every low-impact event needs a large meeting. A short review note can be enough when the learning is clear and the risk is limited.
Hold the review while records, decisions, and participant memory are still available, but after the immediate response has stabilised. The incident record should remain the source for facts and timestamps.
Set the review objective
Start by writing what the review should answer. Good objectives include:
- How did the incident enter or develop in the environment?
- Why did existing controls not prevent or detect it earlier?
- What helped the team contain and recover?
- Which decisions were delayed by missing context or unclear ownership?
- Which changes will reduce likelihood, impact, or response time?
Avoid starting with “who made the mistake?” That question narrows the review too early and can hide the process, design, or governance conditions that allowed the mistake to matter.
Step 1: Confirm the incident facts and scope
Before discussing causes, agree on what is known. Record:
- Detection source and first known time
- Affected assets, accounts, services, data, and locations
- Confirmed and suspected impact
- Containment and recovery status
- Relevant customer, supplier, or business dependencies
- Open questions and confidence level
Separate facts, assumptions, and unknowns. If the scope is still being investigated, say so. A review is stronger when uncertainty remains visible than when the team creates false precision.
Step 2: Reconstruct the timeline
Use a chronological record of observable events, decisions, actions, and communications. Include the time zone and source when possible.
A useful timeline may include:
| Timeline item | What to record |
|---|---|
| Signal | Alert, report, log event, or user observation |
| Validation | What confirmed or challenged the initial signal |
| Decision | Severity, ownership, escalation, or risk decision |
| Action | Containment, enrichment, recovery, or communication |
| Evidence | Source, timestamp, affected asset, and relevant details |
| Outcome | What changed after the action |
Do not rewrite an uncertain early assumption as if it was known at the time. Recording how understanding changed is valuable for improving triage and escalation.
Step 3: Analyse contributing causes
Security incidents rarely have one simple cause. Examine several layers:
Trigger
What event started the incident or allowed the activity to become visible?
Enabling condition
What made the trigger possible, such as excessive access, a missing update, a weak configuration, or a supplier dependency?
Detection gap
Why was the activity not detected, enriched, routed, or understood earlier?
Response gap
Why did containment, ownership, communication, or recovery take longer than expected?
Systemic condition
What process, resource, design, priority, or governance condition allowed the other factors to persist?
Use “why” questions carefully. The goal is not to keep asking why until a person is named. The goal is to find changes that are specific enough to implement and verify.
Step 4: Review controls and evidence
Connect the incident to relevant controls, policies, processes, and evidence. Ask:
- Which control was intended to reduce this risk?
- Was the control selected and in scope?
- Was it implemented as defined?
- Was it operating at the time?
- Was evidence available and current?
- Did the response reveal a missing control or an execution gap?
Do not label a control as failed without checking the facts. A control may have worked as designed while the risk was outside its scope, or the policy may have been correct while implementation or monitoring was incomplete.
Step 5: Review the response itself
Assess the response across the full workflow:
- Intake and initial ticket quality
- Triage and severity classification
- Ownership and handoffs
- Asset, user, and business context
- Evidence preservation
- Containment decisions
- Internal and external communications
- Recovery and validation
- Timeline and decision logging
- Closure criteria and follow-up
Look for friction that can be removed before the next incident. Examples include missing contact details, unclear escalation rules, duplicate tickets, unavailable logs, or approvals that require a person who is not reachable during an incident.
Step 6: Create corrective actions that can be verified
Every action should have:
- A clear outcome
- One accountable owner
- A priority and target date
- A linked risk, control, or review finding
- A verification method
- A follow-up date
Prefer specific actions such as “enforce MFA for the remaining administrator group and record the configuration review” over “improve access security.”
Classify actions where useful:
- Prevent: reduce the chance of recurrence
- Detect: improve signal quality or time to detection
- Respond: improve triage, ownership, containment, or communication
- Recover: improve restoration and validation
- Govern: improve policy, training, supplier, or review discipline
Step 7: Close the learning loop
An action is not complete because a ticket was created. Confirm that the change was implemented, tested, reviewed, and connected to the relevant policy, control map, or evidence source.
Update the risk register when the incident changes the risk assessment. Update policies and procedures when the expected behaviour changes. Update training or runbooks when the response exposed a knowledge gap.
Share the learning at the right level. A useful review can improve security without disclosing unnecessary incident details to people who do not need them.
Example: post-incident review outcome
An administrator account was used from an unexpected location. MFA prevented access to one service but was not enforced on a legacy administrative path. Logging existed, but the alert did not include the asset owner, so triage was delayed.
The review identifies three contributing causes: incomplete MFA coverage, missing ownership enrichment, and an escalation path that was unclear outside business hours.
Corrective actions are assigned to close the legacy access path, enrich alerts with asset ownership, and update the on-call escalation record. Each action has an owner, target date, and verification evidence.
Common review mistakes
Blaming individuals instead of improving conditions
Accountability matters, but a review should also examine process, technology, workload, training, and governance.
Stopping at the first technical cause
“A stolen password was used” may describe the trigger. It does not explain why access, detection, or response controls did not reduce the impact.
Creating actions nobody can verify
Vague actions produce vague progress. Define what done looks like.
Ignoring successful response decisions
The review should preserve practices that helped the team, not only document failures.
Treating closure as the end of learning
Follow-up actions, evidence, and risk decisions are part of the review outcome.
Where IncidentAI fits
IncidentAI helps teams maintain structured incident records, timelines, notes, audit logs, summaries, and RCA drafts. It can organise the review starting point, while responders remain responsible for validating facts, approving conclusions, assigning actions, and making material decisions.
For the first response stage, see the security incident response first-hour checklist. For deeper RCA context, see root cause analysis for SMBs.
Frequently asked questions
What is a post-incident review?
A post-incident review is a structured assessment of what happened, how the organisation responded, what worked, what did not, and which changes should be made.
What is the difference between a post-incident review and root cause analysis?
Root cause analysis focuses on why the incident happened and progressed. A post-incident review is broader and also evaluates detection, response, decisions, communication, recovery, and improvement actions.
Should post-incident reviews assign blame?
The review should focus on facts, accountability, and improving the system. Deliberate misconduct or policy violations may require a separate process, but blame should not replace operational learning.
How long should a post-incident review take?
The effort should match severity and complexity. A critical incident may need a detailed review with several owners; a low-impact event may need a short structured record and one or two actions.
When is a corrective action complete?
It is complete when the change is implemented, verified, supported by evidence, and reflected in the relevant policy, control, process, or risk record.
