Is It Really the Court's Words? A Preregistered Measurement of Quotation Attribution in Federal Briefs
A preregistered holdout study of whether the words a brief attributes to a court are in the court's opinion
LAW Research Division · LAW Research 3(3) · 2026
Is It Really the Court's Words? A Preregistered Measurement of Quotation Attribution in Federal Briefs
LAW Research Division · 2026
Abstract
When a brief puts words in quotation marks and cites a court, the reader assumes the court said them. We measured how often that assumption fails, and how reliably a machine can tell. We built a checker that reads each quotation in an opposing brief, finds the authority it is offered for, fetches that authority's full text, and reports quotations the text does not contain or contains only in altered form. We wrote its acceptance criteria down before measuring anything. We then ran it on four holdout samples of 30 federal briefs each, drawn from public court records and frozen before the checker touched them. An independent verifier graded every objective flag by a different method. The first holdout failed the bar (precision 0.697). After the failure analysis, the next three holdouts each cleared it, with precision of 0.907, 0.878 and 0.907, or 131 confirmed of 146 graded flags pooled. Across those three holdouts, 53 of 87 briefs (61%) contained at least one quotation that the cited authority's text, as digitised, does not contain.
1. The problem
A reporter's headnote is not the court's holding. The Supreme Court said so in 1906 (200 U.S. 321, 337), rejecting an argument built on a headnote. The headnote "is not the work of the court, nor does it state its decision"; "it is simply the work of the reporter." The best-known headnote in American law, the one in Santa Clara County v. Southern Pacific Railroad (1886), has carried corporate personhood for well over a century without being a holding (see Citation Chain Archaeology, LAW Research 3(2)).
The same failure appears at a smaller scale in every motion practice. A quotation can come from a dissent, a concurrence, a syllabus or a case the court was itself quoting. It can come from a different case altogether. It can reproduce the court's words with a word changed and no brackets to say so. Opposing counsel rarely has time to pull every authority and read the passage. A tool that does it automatically is only worth deploying if its accusations are right, because a false "this quotation is not in the opinion" damages the party who makes it.
2. Method
Population. Public federal filings from the RECAP archive whose description names a brief, memorandum, opposition, reply or motion, with more than 8,000 characters of extracted text, not truncated, and at least five quotations of 40–600 characters. One filing per docket.
Holdouts. Four samples of 30, each drawn with a fixed seed from the dockets not used by any earlier sample. Each was frozen before the checker ran on it, with a hash of the sample and a fingerprint of the population it came from, so any draw can be reproduced. Each holdout was scored once.
Preregistration. Before each holdout ran, we recorded the claim and the result that would refute it. Precision of the checker's objective flags must reach at least 0.80. Fewer than 30 graded flags counts as underpowered, not as a pass.
Objective flags. Two kinds count:
- "not in the source": the quoted words do not appear in the authority's text;
- "altered quotation": the words appear, but changed without brackets or ellipses to say so.
Judgment calls, such as a bracketed substitution that may shift the meaning or a court quoting another court, are routed to a human reviewer. They are never counted as objective flags.
Independent verification. Each objective flag was graded against the full opinion text by a separate verifier that works differently from the checker. It splits the quotation at its brackets and ellipses, removes spacing and punctuation, and asks whether every remaining fragment appears in the opinion. The checker never sees the grades, and the grades are never tuned to the checker.
Opinion text. Full opinions from the Harvard Caselaw Access Project and CourtListener. Where the text is a scan, OCR errors are a known hazard, discussed in §5.
3. Results
| Holdout | Briefs scored | Graded flags | Confirmed | Precision | Verdict | |---|---|---|---|---|---| | 1 | 29 | 66 | 46 | 0.697 | Refuted | | 2 | 30 | 43 | 39 | 0.907 | Supported | | 3 | 29 | 49 | 43 | 0.878 | Supported | | 4 | 28 | 54 | 49 | 0.907 | Supported |
Pooled over holdouts 2–4, 131 of 146 graded flags were confirmed, a precision of 0.897. Most confirmed flags were "not in the source" (123); the rest were silent alterations (8). A further 99 judgment flags went to human review and are not counted here.
How common is it? In holdouts 2–4, 53 of 87 briefs (61%) contained at least one quotation that the verifier confirmed is absent from, or silently altered in, the text of the authority it is cited to.
4. What the first failure taught
Holdout 1's twenty false flags were not disagreements about meaning. Each was the checker misreading a convention that legal writers use correctly. A tool that skips these conventions will accuse careful briefs.
- Brackets are declared alterations. "[t]he", "contain[s]", "[it]" for "them", and "[§ 455(a)]" inserted for clarity all tell the reader the words were changed. They become a defect only when the change alters a party, a modality (may/shall/must), a number or a negation.
- The court's own quotation marks. Opinions quote statutes and earlier cases inside their own sentences, and briefs usually drop those inner marks. The words still match.
- Spacing, hyphens and apostrophes. "postdeprivation" and "post-deprivation", "benefit-ted" across a line break, "plaintiffs" for "plaintiff's" in a scan: the letters are the same.
- Footnote calls and omissions. "ERISA[4]" keeps the source's note marker. "[]" marks words or letters left out.
- Where the citation actually is. Federal briefs cite the record in prose ("ECF No. 12 at 3", "Ex. 19", "Compl. ¶ 68"). They cite unpublished cases by database number ("2010 WL 4038826") and regulation by Federal Register or C.F.R. section. A tool that recognises only reporter citations pins these quotations to whatever case sits nearest, and then reports the case "does not contain" them.
Modelling these conventions took precision from 0.697 to the 0.88–0.91 band on briefs no fix was drawn from.
5. Limits, stated plainly
- Recall is not yet measured. We report how often the checker's accusations are right, not how many misattributions it misses. Measuring recall needs readers outside this work to read whole briefs. That study is open.
- Most raw candidates are set aside. Before grading, the checker keeps a quotation only if the authority it accuses is the one the brief actually cites for those words. Between 9% and 14% of raw candidates passed that test. The rest are mostly quotations of the record, of agency documents or of the other side, which never were quotations of the authority nearby. Setting them aside is deliberate; it is what keeps precision high.
- The verifier cannot see inside brackets. It removes bracketed text before testing. Bracketed negations ("[is not]"), which the checker treats as substantive, are therefore scored against the checker, so the precision figures are conservative.
- Digitised text. A quotation can be "absent" from a scan because the scan is damaged. Every flag in practice carries the instruction to confirm it on the reporter page before it is argued.
- Population. Federal filings whose full text is public. State-court practice may differ.
6. Why it matters in practice
The value is not that 61% of briefs misquote. Many of those flags are pin-cite errors or quotations of a quoted source, and they are not misconduct. The value is that every such flag is a checkable fact a court will verify if it is raised. In our practice the checker runs on every opposing brief as an advisory list of attack points and never decides a filing's fate. Each point names the authority, the quoted words and the passage the authority actually contains. The same machinery runs on our own drafts before filing, where it blocks rather than advises.
Data and method. The holdout manifests, preregistrations and scored results are held with the LAW Research Division's measurement files: four holdouts of 30, each sample hashed and its population fingerprinted.
Citation
LAW Research Division. (2026). Is It Really the Court's Words? A Preregistered Measurement of Quotation Attribution in Federal Briefs. LAW Research, 3(3).
Distribution
Published: LAW Research, LAW Research 3(3) Status: published