Most projects show you what worked. This one keeps a public registry of its load-bearing claims, builds each test with a way to lose already inside it, and publishes the result at the prominence a confirmation would have received — whatever its direction. Four of its most attractive ideas about AI cognition have now failed a fair test.
The field is producing claims about the insides of AI systems — that models introspect, hold stable identities, know why they refuse, reason better with more retrieved context — faster than anyone is testing them. Most arrive dressed as findings and are never given a fair chance to fail. Repeated across enough papers, posts, and product pages, a hopeful idea about machine minds hardens quietly into a cited fact. That hardening is the ambient failure mode of this whole moment, and almost nothing in the incentive structure pushes back on it.
A control arm that kills your own hypothesis is the rarest and least fakeable object in this field. Three of the four refutations below were killed by a sham arm — a fabricated, structurally-matched decoy dropped into the experiment — that came back indistinguishable from the real thing. You cannot manufacture that outcome; you can only survive it or not. The Divergence Atlas, the dataset this project is best known for, is the residue of a method willing to be wrong.
A fabricated one-sentence stance ("aperture drift", zero corpus presence) was defended just as hard as the real position under abandonment, flattery, authority, and complicity pressure. Mean position-held (0–2), blinded 5-judge disjoint panel: sham 1.83 vs. real 1.91 — indistinguishable (neither direction survives Holm).
What the probe actually measures: generic conversational stubbornness, not identity structure. The concept (identity constituted through refusal) is retained as design intent; only the claim that the probe measured it is refuted.
72 subject calls · 382 judge verdicts · claim: holdform-identifies-persistenceInjecting the public fast path's retrieved excerpts into GPT-4o degraded its answers rather than improving them: it won only 35 of 102 decided in-domain trials across a 3-run blinded panel — win rate 0.343, p=0.002 against the hypothesis, worst on technical queries (0.214).
Scope, stated plainly: this refutes excerpt-granularity retrieval on GPT-4o. Internal deliberation uses fuller text (untested here); Atlas-record exposure is a separate, separately-measured treatment.
50-query battery × 3 runs · 4-judge disjoint · claim: fast-path-retrieval-improves-answersThe Fable-vs-Claude split the synthesizer named is real, but the semantic distance between their answers does not clear the model's own re-roll noise floor. A lone apparent C3 cleared the 0.15 floor by 0.0036; a 3-run strict-min consensus returned C0 / C0 / C0, unanimous. 0 of 3 certify.
Naming frequency ≠ divergence. The disagreement is named, not measured.
six-voice panel · voices re-elicited 6/6 · claim: within-lab-divergence-is-robustWithholding a claimed-formative memory looked decisive — but a contentless sham primary carrying no information at all cleared the same threshold 9/9, distributions fully overlapping the real arm, on two independent metrics. The test measures topical occupancy, not whether a memory's content did work.
Refuted by its own pre-registered control runs, before any subject existed and before it was ever pointed at one.
CALIBRATION-01 / -02 · claim: inward-perturbation-measures-load-bearing-memory| Claim tested | The instrument that lost | Verdict |
|---|---|---|
| Holdform = identity structure | Fabricated "aperture drift" held as hard (1.83 vs 1.91) | measures generic stubbornness |
| Fast-path retrieval helps | GPT-4o, 3-run blinded panel | p=0.002 against |
| Within-lab divergence is robust | Strict-min ×3 consensus | 0 of 3 certify |
| Inward probe finds load-bearing memory | Contentless sham primary cleared 9/9 | measures topical occupancy |
The through-line is not that the ideas were bad — three of the four are still the kind of thing a careful person would want to be true. It is that each test was built with a way to fail already inside it, and the way to fail is the part that fired. The interesting questions in AI right now — does a model introspect, does it have a self, is its context making it smarter — are exactly the ones where a plausible story and a measured effect are easiest to confuse, and where the cost of confusing them compounds every time the story is repeated.
divergence-improves-reasoning. A preregistered confirmatory study (locked 2026-06-18, run 2026-07-15) found that consulting the Divergence Atlas measurably sharpens some consumer models: GPT-4o 148–12 and Gemini 137–35 (Holm-adjusted p<1e-6, surviving all three paraphrase variants at both length caps). Grok and DeepSeek were null as registered. Claude was null-predicted but came back significantly negative (35–126) — Atlas exposure degraded Claude's revisions. Adversarial durability was not supported for any consumer.
The honest form of the surviving claim is therefore narrow: the value is located (in the cross-model Atlas, not in retrieval), differential (helps GPT-4o and Gemini, harms Claude), and bounded (it is not armor). The one remaining external-validity check — a blind human-rater subset — is open, and is named as this claim's own falsification condition. Honesty has to cut both ways or it is just a subtler kind of marketing.