NIS2 Tabletop Exercise: 3 Scenarios, a Scoring Rubric, and the Article 21(2)(f) Evidence Pack
Article 21(2)(f) of the NIS2 Directive requires “policies and procedures to assess the effectiveness of cybersecurity risk-management measures.” It does not use the words “tabletop exercise” anywhere in its text. What it requires is evidence that your risk-management measures actually work — and a written incident response plan that has never been read aloud in a room full of the people meant to execute it is not evidence of anything.
That is the gap a tabletop exercise closes. This guide covers what the law and its implementing rules actually require (not what generic advice claims they require), three calibrated scenarios you can run this quarter, a scoring rubric that turns a two-hour discussion into an auditable score, and the corrective action plan and evidence pack format a competent authority actually wants to see afterward.
What Article 21(2)(f) Actually Requires (and What It Doesn’t)
The full text of Article 21(2)(f) is eleven words: “policies and procedures to assess the effectiveness of cybersecurity risk-management measures.” [1] It sits inside Article 21(1), which binds every essential and important entity to “appropriate and proportionate” measures scaled to the entity’s size, exposure, and the state of the art — not a fixed checklist. [1]
Read narrowly, (f) is a governance requirement: have a policy, have a procedure, run it. Read for what an auditor will actually test, it means you need to be able to show when you last assessed effectiveness, how, and what changed as a result. A tabletop exercise is the cheapest, fastest way to generate that evidence for the plans that matter most — incident response, business continuity, and crisis management — because it forces decisions out of a document and into a transcript.
Get the NIS2 Article 21 Compliance Checklist
90+ assessment items mapped to CIR 2024/2690 — instant PDF, no payment.
Compliance officers running their first exercise consistently find the same thing: the incident response plan names a role — “IT Security Lead authorizes containment” — and nobody in the room is completely sure who currently holds that role, or whether they have the authority to take a production system offline without a VP’s sign-off. That gap is invisible in a document review. It is the first thing a tabletop exercise surfaces, and it is exactly the kind of finding (f) exists to catch. For the full ten-measure picture Article 21(2) sets out, see our complete Article 21 breakdown.
Does CIR 2024/2690 Apply to You?
Not every NIS2 entity is bound by a specific testing regulation — only some are. Commission Implementing Regulation (EU) 2024/2690 lays down the technical detail behind Article 21(2), but its scope is limited to nine defined categories: DNS service providers, TLD name registries, cloud computing service providers, data centre service providers, content delivery network providers, managed service providers, managed security service providers, providers of online marketplaces, online search engines and social networking platforms, and trust service providers. [2]
| Your entity type | Bound directly by CIR 2024/2690? | What actually applies |
|---|---|---|
| DNS/TLD, cloud, data centre, CDN, MSP/MSSP, marketplaces, search engines, social platforms, trust services | Yes | CIR Annex points 3.5.5, 4.1.4, 4.2.6, 4.3.4, and 7 apply directly [2] |
| Manufacturing, energy, health, water, transport, banking, financial market infrastructure, public administration, and every other essential/important sector | No | Article 21(2)(f) applies directly; the CIR text is a useful benchmark, not your legal basis [1] |
If you’re in the second row, that’s most readers of this article — and it matters, because generic advice that quotes CIR section numbers at you as if they were universal law is simply wrong for your entity. See our CIR 2024/2690 overview for the full technical breakdown of what it does and doesn’t cover.
For the nine CIR-scoped categories, the most direct citation for a tabletop exercise is Annex point 3.5.5, which states plainly: “The relevant entities shall test at planned intervals their incident response procedures.” [2] Two adjacent points extend the same duty to the plans a broader exercise often touches: point 4.1.4 requires the business continuity and disaster recovery plan to be “tested, reviewed and, where appropriate, updated at planned intervals and following significant incidents or significant changes to operations or risks,” and point 4.3.4 uses near-identical language for the crisis management plan. [2] None of the three name a specific frequency. “Planned intervals” is your own risk-based decision to make and document — not a number the regulation hands you.
This guide covers the discussion-based tabletop format specifically — walking a room through a scenario without touching live systems. If you need the live, hands-on technical simulation format instead (restoring an actual backup, failing over an actual system), the site’s companion Incident Simulation guide covers that build separately; running both formats on a rotating cycle is common practice, not an either/or choice.
Tabletop Exercise vs. Functional Test vs. Full-Scale Exercise
“Testing” is not one activity. Auditors and the CIR text both distinguish between discussion-based and operational tests, and conflating them is the second most common gap after missing evidence entirely.
| Test type | What it actually tests | Typical effort | Evidences |
|---|---|---|---|
| Tabletop exercise | Roles, decisions, communication — discussion only, no live systems touched | Low–Medium (half day) | CIR 3.5.5 incident response procedure testing, 4.3.4 crisis management plan, Art. 21(2)(f) [1][2] |
| Functional / backup-restore test | Actual recovery of systems or backups in an isolated environment | Medium–High | CIR 4.2.6 backup and redundancy testing [2] |
| Full-scale exercise | Live activation of business continuity/disaster recovery procedures, often cross-functional | High | CIR 4.1.4 business continuity plan testing [2] |
No source mandates a fixed ratio of the three. As a practical heuristic — not a legal minimum — a defensible programme runs at least one tabletop annually, with functional or full-scale tests on a longer cycle or triggered by a significant incident or major operational change. Business continuity plan testing specifically is covered in more depth in our Article 21(2)(c) business continuity guide.
Building the Exercise: 3 Calibrated Scenarios
A scenario your team has never plausibly faced doesn’t surface real decision gaps — it surfaces how good people are at improvising fiction. Calibrate to your actual threat profile. These three are chosen to stress-test three distinct decision types, not because any one is statistically more common than another: a contain-or-pay call, an upstream-versus-downstream liability judgment, and a threshold decision made under ambiguity. Each is built with a trigger, an early decision point, and the Article 21(2) measures it exercises.
| Scenario | Trigger | Key decision point | Measures exercised |
|---|---|---|---|
| Ransomware | Helpdesk reports file-server encryption and a ransom note | Pay/don’t-pay position, and when the 24-hour Article 23 early warning clock starts | (b) incident handling, (c) business continuity, (f) effectiveness assessment |
| Supply chain compromise | A critical ICT supplier discloses a breach touching a system you share with them | Whether you have a reportable incident yourself, or are merely downstream-affected | (d) supply chain security, (f) effectiveness of vendor oversight |
| DDoS | Customer-facing service degrades, then becomes unreachable | The point at which degraded service crosses the “significant incident” threshold, versus a routine ops issue | (a) risk analysis, (b) incident handling |
Each scenario should carry two or three “injects” — new information dropped mid-exercise to test whether the room adapts, rather than just executing a rehearsed script. In the ransomware scenario, a strong inject is a journalist calling to confirm the breach twenty minutes in: it tests whether the communications role knows not to comment without Legal, and whether anyone remembers that role exists. This inject-based structure follows the exercise-design thinking behind the ENISA Cybersecurity Exercise Methodology, published February 2026, which frames exercise design as a distinct phase from execution and evaluation. [3]
Running the Exercise: A 90-Minute Structure
| Segment | Duration | What happens |
|---|---|---|
| Scenario briefing | 10 min | Facilitator reads the trigger; no solutions discussed yet |
| Initial response discussion | 25 min | Room works through first actions, roles, and escalation |
| First inject + reassessment | 20 min | New information changes the picture; decisions get revisited |
| Second inject + escalation decision | 20 min | Room decides whether/when the incident crosses a reporting threshold |
| Debrief (hot-wash) | 15 min | What worked, what didn’t, captured verbatim — this feeds the scoring rubric |
The facilitator should not play a response role. Their job is timekeeping, capturing decisions as they’re actually made (not as participants later remember making them), and resisting the urge to solve the scenario for the room. Germany’s national competent authority, BSI, frames the purpose of this kind of exercise plainly: checking information-exchange mechanisms and communication paths during an extraordinary incident, alongside situational awareness and recovery procedures. [4] That’s a narrower, more testable goal than “practice cybersecurity,” and it’s a useful yardstick for whether your exercise actually tested anything.
The Scoring Rubric: Turning a Discussion Into a Number
A facilitator’s private impression that the exercise “went fine” is not evidence of anything. A rubric is what turns ninety minutes of discussion into a defensible score — and into a corrective action plan with teeth.
| Dimension | 1 — Fail | 3 — Partial | 5 — Pass |
|---|---|---|---|
| Decision speed | No decision reached in the segment | Decision reached late, after prompting | Decision reached within the segment, unprompted |
| Role clarity | No one knew who was authorized to decide | Right role identified, wrong or absent person | Correct, present person made the call |
| Escalation accuracy | Significant-incident threshold missed entirely | Threshold discussed but decision delayed | Threshold correctly identified and acted on |
| Cross-team communication | Relevant team not contacted at all | Contacted late or through the wrong channel | Contacted promptly through the defined channel |
| Recovery prioritization | No prioritization discussed | Prioritized informally, no stated rationale | Prioritized against a documented criterion (e.g. RTO) |
Score each dimension per scenario, not once for the whole exercise — a team can nail role clarity on ransomware and fail it on supply chain, because different roles own each. Anything scoring 1 or 2 becomes a mandatory line in the corrective action plan; scores of 3 are optional but worth logging as a trend if they recur across exercises.
From Findings to Fix: The Corrective Action Plan
CIR points 4.1.4 and 4.3.4 both require plans to be “reviewed and, where appropriate, updated” following testing. [2] That’s the part most exercises skip: the finding gets discussed in the debrief and then evaporates. The exercise only counts as effectiveness assessment under Article 21(2)(f) if its findings feed back into the actual plan documents.
| Gap identified | Owner | Target date | Evidence of closure |
|---|---|---|---|
| e.g. Containment authority undefined for after-hours incidents | Named individual, not a title | Specific date | Updated policy section + re-test result |
An undated, unowned list of “lessons learned” bullet points is the single most common reason an auditor rejects an exercise as insufficient evidence — it shows the gap was found, not that it was closed.
The Auditor Evidence Pack: What a Competent Authority Wants to See
An exercise that happened but left no retained record is functionally invisible to an auditor. Point 4.1.4 requires testing at planned intervals with updates “where appropriate” — a verbal account that “we ran one” demonstrates neither the testing nor the update. [2] Build the pack as you go, not retroactively:
- Scenario document, dated, with the trigger and injects used
- Attendance list with roles (not just names)
- Completed scoring sheet per scenario
- Corrective action register with current status
- Sign-off from someone with governance authority over the plan
- The previous exercise’s corrective actions, shown closed — this is what proves a continuous cycle rather than a one-off event
That last point is the one generic checklists miss. A single exercise with a clean scoring sheet proves you can run an exercise. A second exercise that references the first one’s corrective actions as resolved proves the effectiveness-assessment loop Article 21(2)(f) actually asks for. Our audit preparation guide covers how competent authorities typically structure a compliance review beyond this specific evidence pack.
Who Owns This? Role Responsibility Table
| Role | Exercise responsibility |
|---|---|
| CISO / IT Security Manager | Designs technical scenario detail; plays the technical response role |
| Compliance Officer | Facilitates or observes; owns the evidence pack and corrective action register |
| Board / C-suite sponsor | Reviews the summary and signs off the corrective action plan — the governance evidence Article 20 expects |
| Legal, Communications, Operations leads | Play their actual functional role in the scenario, not observe from the side |
FAQ
How often does NIS2 actually require a tabletop exercise?
No provision in the Directive or the CIR sets a specific number. Article 21(2)(f) requires effectiveness assessment; CIR point 3.5.5 requires entities to “test at planned intervals their incident response procedures,” and points 4.1.4 and 4.3.4 require the same for business continuity and crisis management plans. [1][2] The interval itself is your own risk-based decision to make and document, not a fixed legal cadence.
We’re not one of the nine CIR-scoped entity types — do we still need to test?
Yes. Article 21(2)(f) binds every essential and important entity directly, regardless of sector. [1] The CIR simply adds sector-specific technical detail for the nine categories it covers. [2]
Does a tabletop exercise satisfy Article 23 incident notification testing too?
A well-designed scenario should include the decision point of when a simulated event crosses the significant-incident threshold, but the exercise doesn’t replace having a working notification procedure — it tests whether people can execute the procedure that must already exist.
Can one exercise cover more than one Article 21(2) measure?
Yes, and it usually should. A single ransomware scenario typically touches incident handling, business continuity, and effectiveness assessment simultaneously — a well-designed exercise is often the most efficient way to generate audit evidence across several measures in one sitting.
This article provides general information only and does not constitute legal or regulatory advice. Requirements may vary by jurisdiction and organisation type. Consult a qualified legal professional or compliance specialist for advice specific to your situation.
For a complete step-by-step walkthrough, see how to facilitate the exercise: the four-week countdown, the facilitator and note-taker split, and the evidence record.
Sources
- [1] NIS 2 Directive, Article 21 — Cybersecurity risk-management measures, nis-2-directive.com (verbatim mirror of Directive (EU) 2022/2555)
- [2] Commission Implementing Regulation (EU) 2024/2690 of 17 October 2024, EUR-Lex
- [3] The ENISA Cybersecurity Exercise Methodology, ENISA (February 2026)
- [4] Übungen — IT crisis exercises, BSI (national competent authority, Germany)
Get the NIS2 Article 21 Compliance Checklist
90+ assessment items mapped to CIR 2024/2690 — instant PDF, no payment.
