The standard I cited does not exist

I'm Johannes. I have spent eighteen years on the response side of security, and I build tooling that turns evidence into something you can defend in front of a court. This one is about a source I got wrong, and how it happened.

For twenty-six days this summer, my own architecture decision record rested on a federal standard that does not exist.

The sentence was the first premise of the first point in the document. "DOJ's standard for forensic algorithms is the calibrated instrument (version + calibration state + audit logs preserved)." It read like something a careful person had looked up. It shaped a decision about how strongly my product is allowed to claim its results can be reproduced. The design review whose whole job was to attack that decision read the sentence and listed it as support.

There is an irony here that I did not enjoy. The calibrated-instrument idea is about traceability: every measurement can be traced to an instrument, and every instrument to its version and its calibration record. The claim that forensic tools must be traceable could not itself be traced to anything.

A real standard has an address. This one did not have an address. This post is about how a sentence like that gets written, why careful people keep reading it back as evidence, and what I now do differently. The honest version includes the part where my correction needed correcting.

The spinach was never a decimal point

The canonical version of this story is about spinach, and it is better than the one you know.

You probably know that spinach's reputation for iron came from a nineteenth-century scientist who slipped a decimal point and made it ten times richer than it is. It is a favourite example of a scientific error spreading unchecked.

The sociologist Ole Bjørn Rekdal went looking for the decimal point in 2014 and found that no one had ever documented it. What he did find was a paper trail you can date. In 1972 the nutritionist Arnold Bender suggested that spinach's fame "may well have grown from a misplaced decimal point". In 1977 it "appears to have been based on" one. In 1981 a piece in the British Medical Journal stated it as fact. By 1982, in a textbook, it had a year: "traced back to a mistake in the transcription of analytical results in 1870". No step in that chain added a source. The hedge just wore off.

The best detail is this. In 1995 a paper titled "The dissemination of false data through inadequate citation" repeated the decimal-point story as fact, citing the 1981 article. When the author of that article was later asked where he had got it, he said he could not remember, and guessed Reader's Digest.

The myth about a careless error in science was itself an error, spread by exactly the carelessness it was used to warn against. Hold on to that, because it is the shape of this post.

Two problems wearing one coat

There is a well-studied mechanism here, and the vocabulary is already written.

Steven Greenberg, a neurologist, mapped every citation in one biomedical claim network in 2009: 242 papers and 675 citations. In that network, the papers supporting the belief got 94% of the citations to primary data, and the handful of critical ones got 6%. He named the failure modes. Citation transmutation: "the conversion of hypothesis into fact through the act of citation alone". Dead end citation: supporting a claim "with citation to papers that do not contain content addressing the claim".

That is the spinach chain, and it is my ADR.

It is not rare. Christopher Baethge and Hannah Jergas pooled 46 studies of quotation accuracy in medical journals in 2025, covering 32,000 quotations. 16.9% were wrong, about half of them badly, with no significant improvement over time. Earlier, Mikhail Simkin and Vwani Roychowdhury modelled how misprints propagate in reference lists and estimated that only about one citer in five had read the paper they were citing. That is a model estimate, not a survey, but it is the right order of magnitude to be worried about.

So a phantom standard is really two problems wearing one coat. Someone writes a plausible sentence and attaches the name of an institution. Then everyone downstream checks whether the sentence sounds right, not whether the institution said it. The first problem is the error. The second is what makes it permanent.

Where the sentence actually came from

I cannot tell you where my sentence came from. What survives of the research session that produced it is its conclusions, not its reading list. That is its own finding, and I come back to it below.

What I can tell you is where the claim lives in public. A verification pass this week ran about twenty search phrasings across preprint servers, law firm blogs, vendor sites and the usual long-form platforms. It found exactly one source that asserts it: a seminar paper by Sahibpreet Singh and Lalita Devi, "Reliability and Admissibility of AI-Generated Forensic Evidence in Criminal Trials", posted to arXiv in December 2025 and to OSF in April 2026. On page 50: "The U.S. Department of Justice has advised that forensic algorithms be treated as instruments whose calibration, software version, and audit logs must be preserved and documented."

That is close to my ADR's wording. It is not proof it was my source. I have no way to establish that, and I am not claiming it.

Their footnote for that sentence is not a DOJ document. It is a 2021 report by the Government Accountability Office, a different body, on forensic algorithms. I searched that report's full text. It does not contain "calibrat", "instrument", or "audit log". Its only version language is a sensible line about testing: that "clearly identifying the software versions used in testing could also improve public confidence". GAO's 2020 and 2024 reports in the same series contain none of the three terms either.

I want to be fair to the paper, because the idea underneath the sentence is right. Forensic practice does treat equipment, software included, as something with a version and a calibration history. The sentence gets the substance roughly right and the authority wrong. In law, authority is not a detail. It is the claim.

What DOJ actually says

Then I checked DOJ itself, which is what the ADR should have done in July.

DOJ's own Code of Professional Responsibility for the Practice of Forensic Science, issued with the Attorney General's memorandum of 6 September 2016, contains no calibration language at all. Its closest requirement is number 9: make and retain records "in sufficient detail to allow meaningful review and assessment by an independent professional proficient in the discipline". That is a good standard. It is about records a peer can review, not about treating an algorithm as an instrument.

DOJ's December 2024 final report on artificial intelligence and criminal justice uses "calibration" once, about bias metrics for risk assessment models. Its version and logging language sits in the facial recognition chapter: vendors should report evaluation results "for the system version that is procured", agencies should "log FRT uses, enabling auditing". Close in spirit, never framed as an instrument, and never a rule for forensic algorithms in general. None of DOJ's 35 published Uniform Language for Testimony and Reports documents mention calibration.

And then the part that turned this from an embarrassing footnote into a post. In January 2021 DOJ published a formal response to the 2016 PCAST report on forensic science. PCAST had argued that feature-comparison methods "belong squarely to the discipline of metrology". DOJ disagreed. It quoted the international metrology vocabulary's definition of measurement, which presupposes "a calibrated measuring system", and concluded that pattern-comparison methods "compare the features/characteristics and overall patterns of a questioned sample to a known source; they do not measure them."

So the phantom does not just lack a source. On the record, DOJ has argued close to the opposite of what the phantom has it saying. A litigator who cited "DOJ's calibrated-instrument standard" in an expert report could be answered with DOJ's own published statement.

Where the real standard lives

The model my ADR wanted does exist. It is not American and it is not a DOJ document.

It is ISO/IEC 17025:2017, the accreditation standard forensic laboratories are assessed against. Clause 6.4.13 requires laboratories to keep records for equipment that can influence results, including the equipment's identity with its software and firmware version, and its calibration dates, calibration results, adjustments, acceptance criteria and next calibration due date. I am paraphrasing it deliberately. It is a paid standard, and I will not quote text I have not read in a licensed copy. That rule came from getting this wrong once already, which I explain below.

In England and Wales the same model has statutory force. The Forensic Science Regulator's Code of Practice, version 2, in force since 2 October 2025, requires records to capture "the instrument configuration, e.g. software version or, if relevant, firmware" (19.2.2(b)(iii)), and to be written so that "another practitioner in the same field, and in the absence of the original practitioner, can follow the nature of the work undertaken" (19.2.3). Under section 4 of the Forensic Science Regulator Act 2021, "the code is admissible in evidence in criminal and civil proceedings in England and Wales", and a court may take a failure to follow it into account. The Code has no mention of artificial intelligence or machine learning anywhere in it.

Look at the pattern. The real standards are dated, numbered, published and locatable down to the clause. You can put your finger on the sentence. The phantom has every word of a standard and none of the address.

The reason I could refute the phantom at all is that the real ones have an address. DOJ's forensic instruments are a finite, published set. You can read all of them and search them. Proving that something is not there is only possible against a list of what should be. I wrote about that last month in a post about files that are missing. This is the same problem in a different setting.

"The substance was right, so who cares?"

It is fair to object that I am making a fuss about a footnote. The requirement I cited, version plus calibration plus preserved records, is a real requirement. I just gave it the wrong home. Nothing about the decision changed when I fixed it, and the ADR's amendment says exactly that.

The objection is right about the engineering and wrong about the evidence. In a technical document, a wrong citation for a right idea is a small error. In anything that reaches a court, the citation is the argument. Evidence law does not ask whether your method sounds right. Under the 2023 amendment to Federal Rule of Evidence 702, the proponent has to show the court "that it is more likely than not" that the opinion "reflects a reliable application" of reliable methods. You meet that burden with things the judge can check. A standard that cannot be found is not something the judge can check, and one the named institution has argued against is worse than nothing.

There is a recent case that shows how courts treat this. In Matter of Weber (Surrogate's Court, Saratoga County, October 2024), an expert had used Microsoft Copilot to cross-check a damages figure and could not explain how it worked or what he had asked it. The court ran the question itself and got $949,070.97, then $948,209.63, then "a little more than $951,000.00". It held that, at least in that court's practice, "counsel has an affirmative duty to disclose the use of artificial intelligence". The expert was not caught fabricating. He was caught unable to say where his answer came from. A phantom citation is the same failure, one level up.

"It is one seminar paper"

The stronger objection is about scale. The pipeline note I wrote a month ago, planning this post, said the phrase "is repeated across the industry" and "appears in vendor blogs and preprints". The verification pass found one paper, posted in two places.

That was a second phantom, and I wrote it myself, in the note planning a post about phantoms.

The objection wins that point, and I have changed the title to match: this was once "the standard everyone cites". The mechanism does not need an industry. Greenberg's amplification finding is that a small number of confident papers, cited by others that add no data, can make a belief look settled. One confident source is enough when the reader is an agent that summarises the first plausible answer and moves on. During this week's verification, the search tool's own AI-written summary stated "The U.S. Department of Justice has advised..." as fact, twice, next to links to GAO and DOJ documents that do not say it. I cannot show that is how the sentence reached my ADR. It is certainly a way it could.

"Humans did this for decades before AI"

They did. The spinach chain ran from 1972 to 1995 without a language model anywhere near it, and the 16.9% quotation error rate comes from a literature written almost entirely by humans. I am not arguing that AI invented phantom citations.

I am arguing that it changes the numbers. In 2023, William Walters and Esther Isabelle Wilder checked 636 citations generated by ChatGPT and found that 55% of GPT-3.5's and 18% of GPT-4's were fabricated outright. That was older models, and current ones are better. But the unit cost of a plausible, authoritative-sounding sentence has dropped to almost nothing, and the cost of checking it at the source has not dropped at all. Damien Charlotin's database of court decisions dealing with AI-hallucinated content stood at 2,095 cases on 28 September 2026. That is a floor, not a rate, but a growing one.

My own failure had a specific shape. An agent wrote the research. An agent-assisted document turned it into a premise. An agent-run design review, whose job was to attack that document, read the premise and called it support. Nobody in that loop was lying. Every step checked whether the sentence was plausible, and it was.

"So have the agent verify it"

This is the objection I most want to be true, and it is the one I have the most evidence against.

The claim was caught on 29 July by a research pass that did exactly that: several agents, told to read primary sources and label every finding as confirmed, reported, or unverified. It worked. Then I corrected the ADR, and the correction needed two further review rounds before it was right.

The first version of the correction said the ISO clause had been "verified against a primary-source PDF". It had not. ISO/IEC 17025 is paywalled, and what had been read was an unofficial copy. The claim of verification was itself unverifiable. The same correction quoted a Third Circuit opinion but cut the quote before its conditional, "had Defendant not pleaded guilty", and truncated the ISO clause mid-item. Both were caught by review, not by the agent that wrote them.

So yes, have agents verify. Just do not treat the word "verified" as a fact about the source. It is another claim, made by whoever wrote it, and it needs a locator like any other.

Where my own work stops

Here is what I cannot claim.

I cannot tell you where the sentence came from. The session that wrote it is gone, and with it the one record that would have answered the question. For someone who builds tooling on the principle that every conclusion should trace back to the bytes it came from, that is an uncomfortable thing to admit. My design documents did not meet the standard I hold my product to. The research notes kept conclusions and dropped sources, which is exactly what makes a claim uncheckable.

I cannot tell you this was the only one. It was caught because one pass happened to check that premise at its source. A later spot check of fourteen outside claims in my own notes found three that were fabricated and seven that meant something other than what I had recorded. That is one sample from one person's notes, not a rate for anyone else's. But it is not zero.

And I cannot fully verify the one standard I now point to. I have paraphrased ISO/IEC 17025 by clause number because I have not read a licensed copy. The correct move is to buy one. Until then, "clause 6.4.13 requires this" is as strong as I can honestly say it, and I am telling you so.

The unsolved problem

The fix I use now is small and mechanical. A claim attributed to an institution carries the document, the date, and the pinpoint, or it is written as my inference. A research note keeps its reading list, not just its conclusions. "Verified" has to name what was read and whether it was the real thing.

None of that solves the underlying problem. Checking a citation at its source costs minutes to hours, and writing one now costs almost nothing. Proving that something does not exist is only possible where the real sources are a closed, published set, as DOJ's are. Most of what we cite is not like that. I know how to show that a sentence traces back to a document. I do not know how to cheaply show that a sentence written by an agent traces back to nothing, before it becomes the premise of a decision. If you have built that check and it holds up, I would like to know how.

What to take from this

  • A real standard has an address. If you cannot name the document, date and clause, you have a rumour.
  • In law the citation is the claim. Right idea, wrong authority is not a small error.
  • Hedges wear off in transit. "May well have" becomes "was" in about three hops.
  • A dead end citation looks exactly like a real one until someone opens it.
  • "Verified" is a claim too. It needs a locator.
  • Keep the reading list, not just the conclusions. A research note without sources cannot be checked later.
  • You can only prove absence against a closed list of what should be there.

If you run legal or forensic research through agents and have a check for attributed claims that actually catches these, I want to hear how it works. Every method I have seen, mine included, checks whether a sentence is plausible, and the phantoms are always plausible.

Building evidence records that trace every conclusion back to its source, so claims can be checked rather than trusted, is the problem I work on during the day: that work is here.


Sources

  • The spinach decimal-point chain (Bender 1972, 1977, 1982; Hamblin 1981; Larsson 1995; the Reader's Digest reply): Ole Bjørn Rekdal, "Academic urban legends", Social Studies of Science 44(4):638-654, 2014. https://doi.org/10.1177/0306312714535679
  • Citation transmutation, dead end citation, the 94% / 6% split in a 242-paper network: Steven A. Greenberg, "How citation distortions create unfounded authority: analysis of a citation network", BMJ 339:b2680, 2009. https://doi.org/10.1136/bmj.b2680
  • 16.9% of 32,000 medical quotations incorrect, about half major, no improvement over time: Christopher Baethge and Hannah Jergas, "Systematic review and meta-analysis of quotation inaccuracy in medicine", Research Integrity and Peer Review 10:13, 2025. https://doi.org/10.1186/s41073-025-00173-z
  • About one citer in five reads the original (a model estimate from misprint propagation): M. V. Simkin and V. P. Roychowdhury, "Read before you cite!", Complex Systems 14:269-274, 2003. https://arxiv.org/abs/cond-mat/0212043
  • 55% (GPT-3.5) and 18% (GPT-4) of generated citations fabricated: William H. Walters and Esther Isabelle Wilder, "Fabrication and errors in the bibliographic citations generated by ChatGPT", Scientific Reports 13:14045, 2023. https://doi.org/10.1038/s41598-023-41032-5
  • 2,095 court decisions involving AI-hallucinated content, as of 28 September 2026: Damien Charlotin, AI Hallucination Cases database. https://www.damiencharlotin.com/hallucinations/
  • The only public source found asserting the claim (p. 50, footnote 43): Sahibpreet Singh and Lalita Devi, "Reliability and Admissibility of AI-Generated Forensic Evidence in Criminal Trials", arXiv:2601.06048, 2025. https://arxiv.org/abs/2601.06048
  • The report that footnote cites, with its software-version sentence: US Government Accountability Office, "Forensic Technology: Algorithms Strengthen Forensic Analysis, but Several Factors Can Affect Outcomes", GAO-21-435SP, 2021. https://www.gao.gov/assets/gao-21-435sp.pdf
  • Requirement 9 on records for independent review: US Department of Justice, Code of Professional Responsibility for the Practice of Forensic Science, with the Attorney General's memorandum of 6 September 2016. https://www.justice.gov/opa/file/891366/download
  • The facial recognition version and logging language: US Department of Justice, Artificial Intelligence and Criminal Justice, Final Report, December 2024. https://www.justice.gov/olp/media/1381796/dl
  • DOJ's rejection of the metrology framing (pp. 3-4): US Department of Justice, Statement on the PCAST Report, January 2021. https://www.justice.gov/olp/page/file/1352496/dl
  • The metrology claim DOJ was answering: President's Council of Advisors on Science and Technology, "Forensic Science in Criminal Courts: Ensuring Scientific Validity of Feature-Comparison Methods", 2016. https://obamawhitehouse.archives.gov/sites/default/files/microsites/ostp/PCAST/pcast_forensic_science_report_final.pdf
  • Equipment records including software and firmware version and calibration history (clause 6.4.13, paraphrased): ISO/IEC 17025:2017, General requirements for the competence of testing and calibration laboratories. https://www.iso.org/standard/66912.html
  • Software version and firmware records (19.2.2(b)(iii)) and records another practitioner can follow (19.2.3): Forensic Science Regulator, Code of Practice, version 2, in force 2 October 2025. https://www.gov.uk/government/publications/forensic-science-regulator-code-of-practice
  • The Code is admissible in evidence in England and Wales: Forensic Science Regulator Act 2021, section 4(2). https://www.legislation.gov.uk/ukpga/2021/14/section/4
  • "More likely than not" and "reflects a reliable application": Federal Rule of Evidence 702, as amended 1 December 2023. https://www.law.cornell.edu/rules/fre/rule_702
  • Copilot variations and the duty to disclose AI use: Matter of Weber, 2024 NY Slip Op 24258, 85 Misc 3d 727 (Sur Ct, Saratoga County, 10 October 2024). https://nycourts.gov/reporter/3dseries/2024/2024_24258.htm
Previous Rehearse the question, not the fire