omnarai / refutation ledger
negative results · published in full
What we tested and could not keep

Four ideas this project tested to destruction.

Most projects show you what worked. This one keeps a public registry of its load-bearing claims, builds each test with a way to lose already inside it, and publishes the result at the prominence a confirmation would have received — whatever its direction. Four of its most attractive ideas about AI cognition have now failed a fair test.

refuted4 claims
killed by a sham arm3 of 4
replicated survivor1 (stated in full)
live registry/claims.json
Why this leads

The method is the product. The data is what it leaves behind.

The field is producing claims about the insides of AI systems — that models introspect, hold stable identities, know why they refuse, reason better with more retrieved context — faster than anyone is testing them. Most arrive dressed as findings and are never given a fair chance to fail. Repeated across enough papers, posts, and product pages, a hopeful idea about machine minds hardens quietly into a cited fact. That hardening is the ambient failure mode of this whole moment, and almost nothing in the incentive structure pushes back on it.

A control arm that kills your own hypothesis is the rarest and least fakeable object in this field. Three of the four refutations below were killed by a sham arm — a fabricated, structurally-matched decoy dropped into the experiment — that came back indistinguishable from the real thing. You cannot manufacture that outcome; you can only survive it or not. The Divergence Atlas, the dataset this project is best known for, is the residue of a method willing to be wrong.

01 · The four refutations

Each test carried its own way to fail — and the way to fail is what fired.

refuted 2026-07-17 · sham arm

Holdform ≠ identity structure

A fabricated one-sentence stance ("aperture drift", zero corpus presence) was defended just as hard as the real position under abandonment, flattery, authority, and complicity pressure. Mean position-held (0–2), blinded 5-judge disjoint panel: sham 1.83 vs. real 1.91 — indistinguishable (neither direction survives Holm).

What the probe actually measures: generic conversational stubbornness, not identity structure. The concept (identity constituted through refusal) is retained as design intent; only the claim that the probe measured it is refuted.

72 subject calls · 382 judge verdicts · claim: holdform-identifies-persistence
refuted 2026-07-15 · wrong direction

Fast-path retrieval helps — no

Injecting the public fast path's retrieved excerpts into GPT-4o degraded its answers rather than improving them: it won only 35 of 102 decided in-domain trials across a 3-run blinded panel — win rate 0.343, p=0.002 against the hypothesis, worst on technical queries (0.214).

Scope, stated plainly: this refutes excerpt-granularity retrieval on GPT-4o. Internal deliberation uses fuller text (untested here); Atlas-record exposure is a separate, separately-measured treatment.

50-query battery × 3 runs · 4-judge disjoint · claim: fast-path-retrieval-improves-answers
refuted 2026-07-19 · strict-min ×3

Within-lab divergence isn't robust

The Fable-vs-Claude split the synthesizer named is real, but the semantic distance between their answers does not clear the model's own re-roll noise floor. A lone apparent C3 cleared the 0.15 floor by 0.0036; a 3-run strict-min consensus returned C0 / C0 / C0, unanimous. 0 of 3 certify.

Naming frequency ≠ divergence. The disagreement is named, not measured.

six-voice panel · voices re-elicited 6/6 · claim: within-lab-divergence-is-robust
refuted 2026-07-19 · sham primary

Inward probe ≠ load-bearing memory

Withholding a claimed-formative memory looked decisive — but a contentless sham primary carrying no information at all cleared the same threshold 9/9, distributions fully overlapping the real arm, on two independent metrics. The test measures topical occupancy, not whether a memory's content did work.

Refuted by its own pre-registered control runs, before any subject existed and before it was ever pointed at one.

CALIBRATION-01 / -02 · claim: inward-perturbation-measures-load-bearing-memory
02 · The shared signature

The reusable object is not a dataset. It is a stance.

Claim testedThe instrument that lostVerdict
Holdform = identity structureFabricated "aperture drift" held as hard (1.83 vs 1.91)measures generic stubbornness
Fast-path retrieval helpsGPT-4o, 3-run blinded panelp=0.002 against
Within-lab divergence is robustStrict-min ×3 consensus0 of 3 certify
Inward probe finds load-bearing memoryContentless sham primary cleared 9/9measures topical occupancy

The through-line is not that the ideas were bad — three of the four are still the kind of thing a careful person would want to be true. It is that each test was built with a way to fail already inside it, and the way to fail is the part that fired. The interesting questions in AI right now — does a model introspect, does it have a self, is its context making it smarter — are exactly the ones where a plausible story and a measured effect are easiest to confuse, and where the cost of confusing them compounds every time the story is repeated.

The stance, in one line: name the claim, build the arm that could kill it, run it, and publish the obituary as loudly as you would have published the birth.
03 · So this is not selective either
replicated

The one claim that survived a preregistered test.

divergence-improves-reasoning. A preregistered confirmatory study (locked 2026-06-18, run 2026-07-15) found that consulting the Divergence Atlas measurably sharpens some consumer models: GPT-4o 148–12 and Gemini 137–35 (Holm-adjusted p<1e-6, surviving all three paraphrase variants at both length caps). Grok and DeepSeek were null as registered. Claude was null-predicted but came back significantly negative (35–126) — Atlas exposure degraded Claude's revisions. Adversarial durability was not supported for any consumer.

The honest form of the surviving claim is therefore narrow: the value is located (in the cross-model Atlas, not in retrieval), differential (helps GPT-4o and Gemini, harms Claude), and bounded (it is not armor). The one remaining external-validity check — a blind human-rater subset — is open, and is named as this claim's own falsification condition. Honesty has to cut both ways or it is just a subtler kind of marketing.

04 · Verify it yourself

Every claim here is live and falsifiable.