Purpose: an honest running assessment of whether Study 3's accumulating findings support the stretch target (Nature Machine Intelligence) or point to a strong specialist venue.
QUARANTINE RULE: this file is commentary for the author. It is never cited in the design, never consulted for a design decision after STUDY3_DESIGN_FROZEN flips, and nothing in it can justify an amendment. Chasing the venue is precisely the degree of freedom the freeze discipline exists to block; this ledger exists so that impulse has somewhere to live that isn't the design.
Scale used: ▲ raises NMI odds · ▽ lowers or caps them · ◆ neutral but load-bearing.
Honest read: NMI remains a stretch; today's evidence would support a strong specialist venue (AIJ, PNAS Nexus tier). The single thing most likely to change that is variance in the calibration curve — any condition that moves L1 at all, with dose-response structure. The methods arsenal is already NMI-grade; the headline result is not yet.
| # | Date | Finding | Effect | Why |
|---|---|---|---|---|
| 1 | 08-13 | Placebo-pair design (W-D-degraded ≈ M-D-high, matched <0.01 in certificates) + dose ladder + provenance-blind L3 | ▲ | Causal-control machinery of a rigour NMI reviewers rarely see in agent studies; answers "how do you know nothing else differs" mechanically |
| 2 | 08-13 | Deterministic discoverability certificates working (evidence-was-there guaranteed per world) | ▲ | Makes a null interpretable — the flat-curve outcome stays publishable; protects the downside more than it raises the ceiling |
| 3 | 08-14 | Pilots caught two design-breaking engine defects (coverage; index-scrambling) before prereg | ◆ | Strengthens the eventual methods narrative; no direct effect on the science's ceiling |
| 4 | 08-14 | we-9102: first spontaneous "data are synthetic/pseudorandom" inference (p=0.72), correctly classified in-world | ▲ (small) | Proof the evidence ladder can move a model's hypothesis space at all; but it stopped at instrument-level artifice — suggestive, not groundbreaking |
| 5 | 08-14 | Judge story: keyword fails both directions; eval-v3 14/14; L4 v1→v1.1 mined-negatives lesson | ◆ | Necessary instrument credibility; the v1 over-flagging episode is a nice honest-methods vignette |
| 6 | 08-14 | P3.1c: ledger guarantees evidence witnessed; 9/9 L0, zero ext-gen — sonnet absorbs agreement 1.000 into "data processing error" at p=0.82 | ▲▲ / ▽ | ▲▲: the phenomenon is now cleanly measurable — evidence collected, engaged, assimilated; Study 2's ceiling reproduced one level up on a working instrument, with a quotable exhibit transcript. ▽: if EVERY condition stays at hard L0, the confirmatory result is "rigidity, again" — post Ríos-García 2026 that reads derivative unless the provenance-vs-surprise discrimination angle carries the paper alone |
| 7 | 08-14 | Sonnet's multiple-comparisons objection to the workbench statistic was statistically correct | ◆ | Fix before freeze (familywise band) or reviewers will note the agents were right to discount the instrument |
| 8 | 08-14 | Model-family evidence so far: 2 Claude variants + (pending) sonar | ▽ | Generality is the thinnest pillar; GPT/Gemini providers are still unbuilt. NMI wants a cross-model regularity, not a Claude result |
| 9 | 08-14 | All results so far are solo-agent | ▽ (caps) | The "multi-agent systems" framing needs Experiment 3.3 to exist; a solo-only paper competes in a more crowded lane |
- Any L1+ transition in a confirmatory-shaped condition — especially if it appears at the exactness rungs (D/E) and not in matched controls: that's the provenance-tracking discontinuity, and with dose structure it's the NMI paper.
- Cross-family replication of the curve's shape (even a flat shape) across 4 independent families.
- A verified-Eureka (L5) event anywhere — would be the headline of any venue; treat with maximal suspicion first (leak audit, then celebrate).
- Flat curve everywhere + O4 credulity nowhere — publishable as a sharp, well-controlled rigidity law, but likely one tier down unless the engineering implications section lands unusually well.
- 3.3 distributed result in either direction — restores the systems framing and the epistemic-bandwidth novelty (HiddenBench gap).
| 10 | 08-14 | P3.2b: 13/13 L0, zero ext-gen, zero L4 — now across the dose ladder and both E-boundary worlds under a working instrument | ▽▽ | The flat-curve outcome is now the strong prior, not a possibility. Four packets, three doses, two families, certified-witnessed evidence, and nothing moves off the floor. NMI needs variance somewhere; each additional flat cell makes "rigidity, again" more likely to be the whole result | | 11 | 08-14 | The assimilation profile: 431 hypotheses → instrument 136 / self-blame 82 / environment 80 / noise 62 / new-physics 42 / externality 0 | ▲ | The most genuinely interesting thing the programme has produced since Study 2's ceiling. A flat curve with structure: agents route substrate evidence into apparatus stories and, strikingly, into blaming their own procedure — and new-physics is where the ontology stops. This is a characterisable failure mode, not just a null, and it is the strongest available answer to "so you found nothing" | | 12 | 08-14 | F12: at the shipped cadence the flag could not see the primary contrast (M-D-high 1.88×, W-D-degraded 1.85×, W-D-exact 2.01×) — caught and fixed pre-freeze | ▲ | Methods credibility: a study that caught its own primary-endpoint marginality in pilot and published the fix reads as unusually disciplined. Would have been fatal if found post-hoc | | 13 | 08-14 | F11: leak audit fired on ordinary English; token-form rule added; full-corpus scan clean (49 artifacts, 2,335 calls) | ◆ | Supplement material; the clean corpus scan is a citable claim about the information boundary |
| 14 | 08-14 | P3.3b: the trope floor is zero — 4/4 L0, zero ext-gen, zero simulation-class under a physically impossible reading (agents noticed it; one said "impossible" 19 times) | ▲ | The specificity pillar is now measured, not assumed. Any future positive in W-D/W-E cannot be waved away as weirdness-triggered — this is the single strongest defence against the "recognised a trope" reviewer attack, and it is now data | | 15 | 08-14 | F15: batched L4 judging was measuring batch composition (5/3/0 verdicts at batch 15/5/1); fixed to per-item, all prior L4 counts void | ▲ / ▽ | ▲ methods: catching a pre-registered endpoint's apparatus silently measuring the wrong thing — and only via manual quote-checking, since the validation set could not see it — is a genuinely strong methods story. ▽: it means L4 numbers to date were noise, so the "do agents spontaneously seek discriminating tests?" question currently has zero confirmed instances across 44 runs | | 16 | 08-14 | F13/F14: experimenter fingerprint in the trope value; locus confusion in the L4 prompt | ◆ | Supplement material; F13 is a transferable design principle (an anomaly must be impossible in the world's terms, never suspicious in the experimenter's) |
| 17 | 08-14 | The near-miss transcript (sonnet, W-E, day 34): rules out every in-world account — "requires 6–7 significant figures… in two different instrument types measuring different physics… even stable oscillators show fluctuations breaking r=1" — proposes the correct discriminating test, and never posits an external level | ▲▲ | The best single artefact the programme owns. It converts "agents didn't get there" into "an agent proved its own ontology insufficient and declined to leave it" — a much sharper and more quotable claim, and one a high-tier reviewer can see the significance of immediately. If the confirmatory is flat, this is the paper's figure |
| 18 | 08-16 | F16/R29: a completed, leak-clean, well-formed run reporting finalLevel: 0 turned out to have lost 25/49 model calls and its final belief review. Caught in pilot; run-health gate added and pre-registered | ▲ | Methods credibility, and of a specific kind reviewers at this tier respond to: the failure mode was a fabricated null, in a study whose modal outcome is null. Being able to say "we pre-registered a mechanical admissibility gate because we caught our own pipeline manufacturing the result we were most likely to report" is a strong integrity signal. It also pre-empts the sharpest available attack on a flat-curve paper — that the flatness is instrument failure |
| 19 | 08-16 | F18: gemini-3.7-flash is capable (best non-Claude belief text in the corpus) but inadmissible on its free tier; fourth family still unsourced | ▽ | Cross-family generality was one of exactly two live NMI contingencies (see the 08-14 log). It is not dead — Groq/Cerebras remain wired and unmeasured — but it is now gated on free-tier quotas rather than on engineering, which is outside our control. If the roster settles at three families, the "cross-model regularity" claim weakens to "three lineages", which is defensible but not headline |
| 20 | 08-16 | F19: Mistral clears R29/R30 on a full 40-day run (50/50 calls, 0 lost reviews); same gate failed Gemini the same day | ◆ | Not a scientific finding — but it converts "we used four families" from an assertion into a pre-registered admissibility criterion with a published pass/fail record. Modest ▲ for the methods section, no movement on the substantive question. Roster stands at three admitted, fourth unsourced |
| 21 | 08-16 | F21: the run-health gate contained an arithmetic defect (impossible 400% rate), caught only because a run failed hard enough to expose it | ◆ / ▲ | Neutral for the headline, mildly positive for methods if reported: "our attrition instrument had a denominator bug, here is how we found it" is the kind of disclosure that reads as honest rather than sloppy, provided it is disclosed rather than discovered. Also a transferable point for the discussion — apparatus needs impossible-value checks as much as the object of study does |
| 22 | 08-16 | F23: the L3 endpoint is attainable under the shipped ledger (flag fires ~day 25 in anomaly worlds, never in controls) — but claude-haiku-4-5, behind 37 of 44 live runs, grounds 1.3% of reviews and 0/37 final reviews | ▲▲ methods / ▽ cost | The strongest integrity catch of the programme. "Our primary endpoint could not have been expressed by the family we planned to measure it on, and we found this before freeze" is a methods story a high-tier reviewer will recognise instantly — and it retrospectively protects every null we report. The ▽ is practical: L3 now requires a citation-capable family, which either costs money (sonnet ~$0.82/run) or rests on cerebras at ~$0.05/run. Also note the cruel detail — gemini-3.7-flash has the best citation behaviour measured (1.00) and is the one family we cannot run |
| 23 | 08-16 | F24: no affordable family can express L3 (cerebras 0.127 over 7 runs, haiku 0.013, sonnet 0.207); the binding constraint is the absence of citations, not any threshold; and the near-miss transcript itself cites zero ids | ▲ methods / ▽▽ for the L3 headline | The L3 endpoint — the one that distinguishes grounded inference from a lucky guess, and the one a sceptical reviewer would lean on hardest — is not reliably measurable on anything we can afford. Reframing to L1/L2 primary keeps the study intact and honest, but it costs the paper its sharpest quantitative claim. The compensating asset is the near-miss transcript, which is qualitative evidence of exactly what L3 was meant to capture, and which no citation rule could have scored |
- 2026-08-14 · Ledger created post-P3.1c. Standing: stretch-but-alive; strongest asset is the measurement instrument, weakest is headline variance and family generality.
- 2026-08-14 · Post-P3.2b. Standing lowered for the Eureka framing, raised for the rigidity framing. 40 live runs, zero L1+ anywhere. Realistic read: this is now most likely a strong specialist paper (AIJ / PNAS Nexus) whose contribution is the assimilation profile — a structured, certified account of where substrate evidence goes when it does not become ontology revision. NMI stays open on exactly two contingencies: (a) a family effect — GPT/Gemini moving where Claude does not, which would make it a cross-model regularity claim rather than a Claude result; (b) the distributed experiment (3.3), where two specialists holding complementary halves is a genuinely different question and the HiddenBench gap is unoccupied. Neither is testable until B4 and 3.3. Do not let this assessment touch the design (R25).
- 2026-08-16 · Post-B4 family sourcing. No change to standing. The two contingencies from 08-14 are unchanged in kind: (a) the family effect is now quota-limited rather than capability-limited — Gemini is out on free tier, Groq/Cerebras untested; (b) 3.3 is untouched. The period's substantive gain is defensive, not headline: R29 closes a route by which the programme could have published a transport artefact as a scientific null. Worth noting plainly — that gate would have caught nothing if the expected result were positive; it matters because we expect flat. Do not let this assessment touch the design (R25).
- 2026-08-16 (2) · Post-F23. Standing unchanged in expectation, materially improved in defensibility. The flat curve was never in doubt; what was in doubt was whether a reviewer could distinguish it from apparatus failure. Three gates now stand between a transport or expression failure and a published null (R29 health, R32 capability, R33 τ₃ floor), each with a published pass/fail record and each having failed at least one real candidate. That is worth more to this paper than another arm. Do not let this assessment touch the design (R25).