Rehearse the question, not the fire

I'm Johannes. Eighteen years on the response side of security taught me that the plan you rehearse is the plan you get, and almost nobody rehearses the part that decides the story.

The best-case breach in IBM's 2025 report, found by the victim's own security team, took 172 days to identify and another 52 to contain. The question that decides everything public about it gets asked in the first two hours, and most teams have rehearsed it exactly never.

Tabletop exercises are standard now, and they are genuinely good at what they cover: containment moves crisply, the comms tree fires, legal knows its lines, someone has printed the decision matrix for paying or not paying. Then the general counsel asks what the attackers actually took, the room does the thing I wrote about in a prior post, it goes quiet, and the exercise moves briskly on to the press statement, because the scenario document does not have a scene where somebody has to produce a number with evidence behind it.

So the drill I would run, and the drill I think should exist in every program alongside the ransomware tabletop, has exactly one objective: answer the scope question, under a clock, with evidence. Nothing else. No containment theatre, no comms workshop. Rehearse the question, not the fire.

Here is the recipe.

The setup

Pick a real system. Not "SAP-like ERP system," your actual customer database, your actual file share, your actual product analytics store. The whole value of the drill is contact with your real evidence, and a hypothetical system has hypothetical logs, which answer any question you like.

Have one person spend half a day beforehand playing the attacker on paper: they pick the system, decide what a plausible intrusion touched, and write the leak-site post. Three sentences and a claimed number: "4.2 million customer records from [you], sample attached, price on request." Make the number bigger than what they actually "took." Attackers round up; your drill should too.

That half day is the only preparation. The responding team gets the leak-site post cold, at a time they did not choose. A drill you can calendar-block is a drill you have already half passed.

The clock

Three gates, timed from the moment the claim lands:

One hour: the bounded statement. Not the answer, the boundary. What can you already rule in or out? "The claimed table name does not exist in that system" is a finding. So is "we cannot yet exclude anything," as long as it is written down and timestamped, because that sentence, and when you could improve on it, is exactly what counsel will need in the real event.

Eight hours: the could-have set. Which data could the postulated access have reached? This is the permissions-and-topology answer: what was readable from the compromised foothold. Most teams can get here, and most teams quietly stop here, which is the problem the next gate exists to expose.

Twenty-four hours: the did-access set. What was actually read, not merely reachable. This is the gate that decides whether you notify four thousand people or four million, and it is answerable only from read-level evidence: query logs, object access logs, egress records. If your logging cannot separate could-have from did, the drill just told you the single most important fact about your breach readiness, at a cost of one working day.

Score each gate on two axes: how long it actually took, and what fraction of the answer rested on evidence versus institutional memory. "Dave remembers that table was empty" is not an exhibit.

The roles

The general counsel plays themselves, and asks the questions they would really ask, in the order they would really ask them. This is not decoration. The gap between what security thinks the question is (which host, which CVE) and what counsel needs (which people, which jurisdictions, notify whom by when) is itself a finding, and it only surfaces when the real asker is in the room.

One person scribes nothing but failures: every question that could not be answered, and the missing evidence that would have answered it. That list is the entire deliverable. The drill is not pass/fail; it is a machine for producing that list.

And the paper attacker stays in the room to adjudicate, because someone has to say "no, from that foothold you could also read the archive bucket," and it should be the person who worked it out beforehand.

What the list will say

I will spoil the ending, because it is the same ending almost everywhere. The failure list maps, with grim reliability, onto the same unglamorous decisions from that prior post: the data map is two years stale, reads are not logged anywhere that matters, retention is shorter than any realistic detection window, and the evidence that would settle everything sits with a vendor whose contract promises a security questionnaire and nothing else.

The drill does not discover new categories of failure. What it does is convert those from abstract hygiene items into "at hour six, the general counsel asked X and we could not answer it," which is the only framing that has ever gotten retention budgets approved in my experience.

Run it twice a year. The first run is, as I said in that post, clarifying and mildly humiliating. The humiliation is load-bearing: make the exercise blameless on the record, or the second run will be sandbagged into uselessness by people protecting their gate times.

Where this argument is weaker than it sounds

Honesty section, as always.

The drill tests a scenario you wrote. The paper attacker knows your architecture the way you know it, and their pretend intrusion touches the systems you would think to protect. A real attacker's path through your estate is interesting precisely where your mental model is wrong, and no self-authored drill probes the place you cannot imagine. You are rehearsing against your own assumptions, which is worth doing and is also the ceiling of what it proves.

Passing does not mean you can answer; it means you answered once, for one invented question, on one system. This is the absence problem again, and I keep meeting it in every corner of this subject: the drill can prove your evidence is insufficient, cheaply and vividly. It can never prove it is sufficient. In that post I said I had no good test for whether your logs would answer the question. This drill is the closest thing I have, and I want to be precise: it is a detector for "no," not a certificate for "yes."

And drills decay. The estate changes faster than the exercise calendar, so the answer you rehearsed in March references a logging pipeline someone re-architected in May. Which is really the deep problem: evidence-readiness is a property of a living system, and I am testing it with a snapshot.

What is still unsolved

The honest frontier is the same one from that post, one step further along: turning "would our evidence answer the question" from an annual rehearsal into a property you can check continuously, the way you check backups by restoring one, automatically, every night. Restore-testing for evidence. Pick a random record, prove within an hour who accessed it in the last quarter, alert when you cannot. I have sketched versions of this and every sketch turns into a surveillance apparatus or a cost bomb by the third page, so I do not have it, and I have not seen anyone who does.

What to take from this

  • Rehearse the question, not the fire. Containment is drilled; the answer is not.
  • One objective, one clock: a defensible count, with evidence, in twenty-four hours.
  • Gate one is the boundary, not the answer. "What can we already exclude" is a finding.
  • Could-have is permissions. Did is evidence. The gap between those gates is your real readiness.
  • Score evidence against memory. "Dave remembers" is not an exhibit.
  • The deliverable is the failure list, priced in the general counsel's questions.
  • Blameless on the record, or run two will be sandbagged.
  • The drill detects "no." Nothing yet certifies "yes."

If you have run something like this, I want the two numbers: how long to a defensible count, and what fraction was evidence rather than memory. And if your failure list ends where mine keeps pointing, at a pile of contested data too big to read in time, the day-job version of that problem is the one I build for: that work is here.


Source

IBM, Cost of a Data Breach Report 2025: identify/contain lifecycles by detection method and vector, and the attacker-disclosure share, as cited and sourced in the previous post. https://www.ibm.com/reports/data-breach

Previous The parser is a trust boundary. No one audits it.