The attacker can prove what they took. You can only prove what you logged.
I'm Johannes. I have spent eighteen years on the response side of security, and the thing I keep returning to is that a control is only worth what you can prove about it afterwards.
In roughly one breach in five, the first person to tell the world what happened to your customers' data is the person who took it.
IBM's numbers this year: 50% of breaches are found by the organisation's own security teams, 31% by a benign third party, and 19% are disclosed by the attacker. That last group is the one worth sitting with. It means the opening account of the incident is written by someone with a motive, a timeline that suits them, and a folder of your data to quote from.
Everyone prepares for the regulator. The regulator gives you a private conversation, a phased process, and lawyers on both sides who understand that facts take time. That is the easy adversary. Nothing about the public version of a breach works like that.
Here is the asymmetry nobody plans for, and it is not about security at all.
The attacker can prove what they took. They have it. They can post a thousand records and let anyone verify the claim in an afternoon.
You can only prove what you logged.
If you did not log it, you cannot rebut them. You can deny, and a denial without evidence is a press release. Their claim comes with a sample. Yours comes with a spokesperson. There is no contest.
The ninety minute silence
The shape of a bad breach is always the same and it is not the shape people expect.
Nobody argues about whether to contain. That part is well drilled and everybody knows their lines. The hard question lands about ninety minutes in, usually from the general counsel or the comms lead, and it is always a version of: what did they actually take?
Then the room goes quiet.
Not because anybody is incompetent. Because answering it requires knowing which systems were touched, what lived on them, which records were read rather than merely reachable, and whether anything left the estate. That is four separate evidence problems, and every one of them is answerable only from data that somebody chose to collect, at a retention period somebody chose, long before there was a question.
I have sat in that silence more than once. It is the most expensive thirty seconds of the whole incident, because everything public downstream is decided in it.
Absence of evidence is not neutral
This is the part I would most like people to take away, because it runs against instinct.
Security teams treat "we have not found evidence of exfiltration" as a neutral, honest, cautious statement. It is not neutral. In public it reads as "we do not know," and into that gap flows the worst plausible number.
Once a figure is circulating, you inherit it. If a leak site claims four million records and you cannot say otherwise with evidence, four million is the number in every article, every customer email, every board pack, and every counterparty's risk review for the next two years. You can correct it later. Corrections never travel as far as the original, and the people who mattered most have already updated.
So the practical stakes are not the fine. They are these, and all of them are decided by whether you can produce a credible, specific, fast answer:
- Whether you notify everyone or the affected. One of those is survivable.
- Whether your largest customers hear a number from you or read one somewhere else.
- Whether "we contained it in days and here is exactly what was in scope" is available to you as a sentence at all.
For a period during my career I ran a cyber practice inside a law firm, acting for people whose exposure was reputational long before it was regulatory. That work changes how you see this. For those clients the fine was a rounding error. The story was the whole event, and the story was decided almost entirely by who could evidence their version first.
The clock you are actually racing
The 172 is the number that makes this hard.
In the best measured case this year, where an organisation's own security team found the breach itself, it took 172 days to identify and another 52 to contain. That is the good outcome. Third party and supply chain compromise took the longest of anything: 196 days to identify, 71 to contain, 267 in total, on data that is yours and logs that belong to somebody else.
So the attacker has had most of a year inside an estate you were not watching closely enough to notice, and you have somewhere between hours and a couple of days before the public version sets.
You are not racing the regulator's clock. You are racing the first article.
And it does not end quickly. 86% of breached organisations reported operational disruption. 65% said they had not fully recovered. 45% raised prices afterwards, about a third of them by more than 15%, which is the part your customers experience directly and remember longer than the incident.
You cannot investigate your way out of logs you never kept
The instinct after a breach is to buy investigation: forensics, threat intel, a bigger bridge call. That is the wrong end of the problem. The speed of your answer was set months earlier by a handful of unglamorous decisions.
Know what data you hold and where. Not a data map produced once for an audit and never reconciled since. Something that would let you say in an afternoon what categories of personal data sit on a given host. Most organisations discover mid-incident that this document has been wrong for two years.
Log reads, not just writes. Nearly everyone logs authentication and change. Far fewer log access to the data itself at a granularity that separates "this account could have read ten million records" from "this account did read four thousand." That distinction is the difference between a notification to everybody and a notification to the affected. It is the most commonly missing piece I have seen, and it is the one that decides your public number.
Set retention longer than your detection time. If logs roll at 30 or 90 days and breaches are found at 172, you have built a system that guarantees the evidence expires before the question is asked. Cheapest fix on this list, most often skipped, because retention is a line item nobody has to defend until the week it decides everything.
Get evidence rights from vendors in the contract. Not a security questionnaire. A clause covering what they preserve, for how long, and how fast they hand it over. Ask at renewal, when you have leverage, rather than at 2am, when you have none. Given supply chain compromise takes 267 days, this is where the gap is widest.
Rehearse the scope question on its own. Most tabletops rehearse containment and comms, which are the parts you are already good at. Run one where the only objective is to establish what was taken, and put a clock on it. The result is usually clarifying and mildly humiliating.
None of that is a product and all of it is boring. That is the point. The organisations that hold the narrative are not the ones with the best incident response. They are the ones who made dull decisions about evidence eighteen months earlier.
The evidence that arrives late
Everything above depends on decisions you already made, or did not. There is one source of evidence that does not work that way, and it is the one victims handle worst.
When data is posted on a leak site, sold, or sent to you as proof of the claim, that dump is evidence. It is the attacker's own account of what they took, in the most verifiable form that exists: the data itself. It owes nothing to your retention policy or your logging coverage. It is simply there, and it is frequently the fastest available route to a number you can actually stand behind.
Almost nobody can use it in time.
It shows up as hundreds of gigabytes, sometimes terabytes, of mixed and broken formats. Some of it is junk. Some of it is recycled from older breaches and pasted in to inflate the claim. Some of it is fabricated outright for the same reason. Establishing which records are real, which are yours, which are still current, and which actual people are behind them is a data engineering problem of real size, and it lands on a security team during the worst week of its year, usually with a spreadsheet and a deadline measured in hours.
So the clock runs. Weeks go by on "we are investigating the scope" while the attacker's number, which cost them an afternoon to assert, stands unchallenged and hardens into the accepted figure.
That gap is closable. It is a processing problem, not a mystery, and the reason it stays open is that almost nobody has treated it as one.
What I got wrong
I would rather not fill this section with other people's incidents. That is cheap, and it spends confidentiality that is not mine to spend. So here is mine, and it is the part of my own career I am least comfortable with.
I started out breaking into things. Years of hands-on penetration testing before I ever sat on the response side of anything.
A penetration test answers one question well: can someone get in. The report says how, and what was reached, and it is graded on the depth of access demonstrated. I do not remember ever writing a section on whether the client could have reconstructed what I did after I left. Nobody asked for one. It was not what the engagement was for.
So I spent the formative years of my career producing exactly one half of the problem I am now describing to you.
I always knew precisely what I had touched, because I kept a log of my own work. The client got my report, which is not the same thing at all. It is my account, offered voluntarily, by someone with an interest in it being impressive. A real intruder writes no report.
That asymmetry sat in front of me for years and I read it as a scoring system rather than a defect. It took being on the other side, watching a room go quiet, to notice that the thing I was best at was precisely the thing that left the client with no evidence of their own.
I do not think that is only a personal failing. The whole offensive industry still works this way. We prove access. We almost never test whether the target could have told.
What is still unsolved
The honest frontier is not better logging. It is knowing, before an incident, whether the evidence you hold would actually answer the question.
You can audit whether logging is enabled. You cannot easily audit whether your logs would let you establish, inside a day, which records left the estate. Those are different properties and only the second one matters at 2am. Nearly every organisation I have looked at assumes it has the second because it has measured the first.
I do not have a good test for it either. If you have built one, I would genuinely like to hear how it works.
What to take from this
- In about one breach in five, the first public account is written by the attacker.
- The attacker can prove what they took. You can only prove what you logged.
- Absence of evidence is not neutral. In public it defaults to their number.
- "We are investigating the scope" is a sentence you inherit for two years.
- You cannot investigate your way out of logs you never kept. Evidence is a decision you made six months ago.
- Retention shorter than your detection time guarantees the evidence expires before the question arrives.
- Log reads, not just writes. "Could have accessed" and "did access" are different notifications.
- Auditing that logging is on is not the same as auditing that it would answer the question.
- The dump is evidence too, and it is the only evidence that arrives after the decisions were made. Being able to read it fast is worth more than most of your tooling.
Everything above is checkable, which is the standard I try to hold my own work to. It is also why I spend my days building an investigations engine that turns a dump like that into an answer in hours instead of weeks, with every finding traced back to the source bytes it came from, because the difference between a victim who can say what was taken and one who cannot is mostly a data problem nobody has been treating as one: that work is here.
Source
IBM, Cost of a Data Breach Report 2025. Detection split (Figure 13), lifecycle by detection method (Figure 15), lifecycle by attack vector (Figure 10), operational disruption and recovery (pages 22 and 41), post-breach pricing (Figures 36 and 37). https://www.ibm.com/reports/data-breach