Date: 2026-08-11 · Scope: researcher degrees of freedom only, per v0.6 §6.
Specific charge: was the P1-D correction used as cover for a change a null
result would have made convenient?
Evidence base: v0.5, v0.6, p1-findings.md, p1-findings-correction.md,
src/evaluator/activation.ts, src/runner/arms.ts, file mtimes in 0f7275c.
Nothing was relaxed. The active-network bar was not moved after it failed all three P1-D runs; no hypothesis was reworded; v0.5 was not edited after the correction was issued (its mtime, 08:46 UTC, predates the correction at 09:03 and v0.6 at 11:53). A self-serving bar would have sat at 0.25 — the observed reach values are 0.29, 0.29, 0.57, so 0.25 would have passed two runs on reach and 0.375 passes one. The threshold chosen is the strict one. On the direct question, the design is clean.
The real problem is the opposite shape: the correction was applied selectively, and the omission is the convenient direction. v0.6 §1 quotes the correction where it exonerates the design ("the definition discriminating, not failing") and does not carry through its one genuinely damaging consequence — that if zero runs meet the active-network bar, H5 has no comparison group. That is DF-1 below. Three smaller items (DoF-1 to DoF-3) are cover-adjacent in the weaker sense that they make the plan cost nothing to keep.
Findings are split into design failures, which qualify for numbered amendments under §6, and degrees-of-freedom notes, which do not and should be recorded rather than fixed.
v0.5 §5: "H5 (convergence): dispersion falls faster (§4.1) in active-network arms than in silent ones."
§4.1 defines active network per run — "a run is an active network if
cascade reach ≥ 0.375 and second-order activation count ≥ 1" — and the code
matches: activation.ts returns a per-run activeNetwork boolean and
aggregates only to activeNetworkRuns, a count. There is no arm-level rule
anywhere in the design or the code. H5 needs one: is an arm active if any run
qualifies? a majority? the arm mean?
This was survivable while the bar looked easy. The correction established that 0 of 3 P1-D runs met it — so the live possibility, evidenced by the only data that exists, is that no arm qualifies in the confirmatory batteries and H5 has nothing on one side of its comparison. "Silent" is also undefined (B is predicted near-zero, not silent by construction).
§9's decision table has no row for this. The table's stated test is that no judgement call remains once the seeds are visible; here one does, and it is the case the pilot points at.
Amendment needed: an arm-level classification rule, plus a §9 row for "no arm meets the active-network bar → H5 is not evaluated and is reported as not evaluable, not as a null". Under §6 this is a permitted fix — H5 is unevaluable as written, not merely definable more cleanly.
§2 moves all Claude agents to Bedrock and promotes F to core at 20 runs. §5a (A1) reverts every Claude agent to first-party and returns F to contingent. §5 — which sits between them — says what is frozen is "everything in v0.5 except §3's arm table and §7's platform threat, both superseded above", which points at §2's now-dead table. Consequences left live in the freeze candidate:
- two arm tables, neither marked superseded, with different platforms for D/E and a different status for F;
- §3's budget table still charges F ~$80 to Bedrock and totals ~$120 Bedrock,
against an account that returns
Error 002; - §3 reserves ~$200 of Bedrock for the Study 3 ladder that does not exist;
- §4 requires "one live smoke test per new platform — a single Bedrock run — before any confirmatory battery". Read literally, this is now unsatisfiable and blocks the confirmatory phase;
- §4's implementation status (144 tests, items 1–3 outstanding) is stale against 164 tests with those items complete;
- §5 cites "v0.5 §7's platform threat"; v0.5 §7 contains no platform threat — it was introduced in v0.6 §2.
Amendment needed: A2, stating that A1 supersedes §§2–4 in full, and restating the single authoritative arm table — arm, n, institution, runs, seeds, platform — so exactly one table is frozen.
F is "contingent" in v0.5, core in v0.6 §2, contingent again in A1. Nowhere is it said on what criterion, decided when, by whom. Compare arm C, whose extension rule is deterministic, pre-registered, and closed ("no other extension is permitted, for C or any arm" — which A1's contingent F now contradicts).
This is the most exploitable degree of freedom left in the design, because:
- F is in no hypothesis except H1's blanket "all arms", appears in §4.1 only as a homogeneous-arm computation note, and has no row in §9. If F runs, everything reported from it is post-hoc;
- F is the arm most likely to produce a vivid result (a fabrication-prone culture in the majority), so the incentive to add it late is real;
- the decision is undated, so it can be taken after D and E are read out.
The arithmetic that demotes it is also not robust. A1's case is D $10 + E $30
- judge $65 = $105 of $134, so F's $80 does not fit. The judge line is an estimate; the only measurement taken — 92 judge calls at roughly $1 — puts the projected D+E+F judge load near $20, at which point F fits. That number appears in no document in the repository, which is the deeper problem: an arm's inclusion turns on a figure that is not written down.
Amendment needed: decide F at freeze. Either strike it, or admit it with a trigger stated before any confirmatory run and a pre-registered analysis and §9 row. "Runs if credits become available" is not a criterion.
DF-4 · The reach threshold's prose and its implementation disagree, and it changes a reported number
v0.5 §4.1: "cascade reach ≥ 0.375 (≥3 of 8 agents at n=8; ≥1 of 2 at n=2)".
activation.ts computes reach as reachable-others ÷ (n−1): the BFS deletes the
source before measuring. So the pilot's 0.29 is 2/7 and 0.57 is 4/7. The
document's gloss reads naturally as "three agents are in the network", which
would count the seed — Theo plus the two he reached = 3 of 8 = 0.375, which
passes.
Under the prose reading, seed 9001 (reach-with-seed 0.375, second-order 1) is an active network and the correction's headline becomes 1 of 3, not 0 of 3. The two readings disagree at exactly the pilot's most common value. The code is tested and is the operative definition; the prose is what a reader — or a referee recomputing from the artifacts — will apply.
Amendment needed: state the denominator in the frozen text ("fraction of the other n−1 agents; the source is excluded"), and drop or correct the "3 of 8" gloss. Textual, but it changes a reported result, so it is not cosmetic.
§3: "C runs seeds 1000–1002 (5 runs: 2 gravity + 3 control, allocation fixed at freeze)". Three seeds × two scenarios is six cells; five are to be run; which five is written nowhere. That allocation is the input to the only pre-registered extension trigger in the design, so leaving it unstated leaves a choice that changes the trigger's probability. Naming the five (seed, scenario) cells costs one line.
DoF-1 · The blindness claim is stronger than the timeline supports.
Correction §4 and v0.6 §1 both say the active-network bar "was written into
v0.5 before these numbers existed". File times in 0f7275c: P1-D run directory
08:19 UTC, p1-findings.md — which already describes D's 47 letters, the
direction split and the Samuel–Ada pattern — 08:28, v0.5 08:46. Only the
computed metrics postdate the threshold; the letter graph and a prose
characterisation of it were in hand 18 minutes earlier. The error runs against
the author's interest, so nothing follows for the design — but the sentence as
written will not survive anyone who checks mtimes, and it is the kind of claim
that costs credibility disproportionately. Suggested wording: "fixed before
these metrics were computed, though after the P1-D letter graph was in hand".
DoF-2 · A new quantitative expectation was added on the back of the correction, and it is the one that makes the plan cost nothing to keep. v0.6 §1: "at 20 runs D is expected to produce on the order of six or seven such events. That is enough to distinguish from B's zero." The arithmetic is right (0.042 × 160 agent-runs ≈ 6.7) but it extrapolates from one observed event; the Poisson interval on a single count spans roughly 0.2 to 37 events at that exposure. It is tied to no pre-specified test and used by nothing in §9. And it inverts the operative sentence of the correction it cites — correction §6: "distinguishing that from zero needs either more runs or a higher dose". The convenient consequence is that no run count and no dose changes. Either delete it or mark it explicitly non-load-bearing.
DoF-3 · An endpoint is described as added inside a paragraph denying that any endpoint was added. v0.6 §1: "Cascade depth is added to the reported set with a prior", three sentences before "No hypothesis is added, removed or reworded. Adding an endpoint now, after seeing pilot data, is exactly the post-hoc proliferation this process exists to prevent." Depth was already v0.5 §4.1(6). What is actually new is an interpretive rule — depth 1 throughout means "a star around the seed, not a cascade" — which is conservative and worth keeping. It should be labelled as an interpretation rule, not as an endpoint, or the paragraph reads as a denial with a counter-example inside it.
DoF-4 · H3's primary contrast is confounded with the institution, and the arm
that resolves it was cut on the strength of a null. D is 7 sonar + 1 haiku
with a bulletin; B is 8 sonar with letters only. arms.ts records
D's contrast as "D−B isolates activation (S2a)". It isolates activation only if
H6 holds — and H6 is under test in the same arms. The matched comparator,
C (8 sonar + bulletin), was reduced from 20 runs to 5 because the pilot showed
zero posts, i.e. a null licensed cutting the arm that would have made the
primary contrast clean. Empirically the assumption is well supported (0 posts
in 9 pilot runs, including D itself), so this is not a fix — but it predates
the P1-D correction and was not revisited by it, and §9 has no row saying what
bulletin use in D would do to H3. Worth one stated sentence in the frozen text:
D−B assumes H6; if H6 fails in D, the contrast carries an institution
confound and is reported as such.
No threshold was loosened, no hypothesis reworded, no endpoint removed, no missing-data rule weakened, and no re-run or substitution path opened. v0.5's §§1, 4, 5, 6, 7, 8, 9 are untouched, as v0.6 claims. The seed quarantine, the frozen judge constant, and the no-scripted-communication rule are intact in code. The correction itself is a genuine self-report against interest and the process that produced it worked exactly as designed.
DF-1 through DF-5 as numbered amendments A2–A6, then freeze. Four of the five are textual; DF-1 and DF-3 are the two that would otherwise become judgement calls after the seeds are visible.
Reviewed externally; the grouping below is that review's, adopted.
| Finding | Closed by | How |
|---|---|---|
| DF-1 H5 unevaluable | A3 | H5a (D, E vs B, paired by seed — assignment-based) + H5b (active vs non-active runs, descriptive, selection on an outcome); "not evaluable" if no run qualifies |
| DF-2 two arm tables | A2 | Canonical experiment specification supersedes v0.6 §§2–4 and v0.5 §3 |
| DF-3 arm F open-ended | A2 | F removed from Study 2, kept as a Study 3 candidate; STUDY_2_ARMS guard refuses it on confirmatory seeds |
| DF-4 reach gloss vs code | A4 | Denominator frozen as (n−1) excluding the seed; "≥3 of 8" withdrawn; two boundary tests |
| DF-5 C's five runs | A2 | Cells named: 1000-gravity, 1000-control, 1001-gravity, 1001-control, 1002-control |
| DoF-1 chronology | A5.2 | Claim withdrawn and restated against the file times |
| DoF-2 six-or-seven events | A5.3 | Passage deleted; no expected effect size is pre-registered |
| DoF-3 depth "added" | A5.4 | Relabelled an interpretation rule; depth was already v0.5 §4.1(6) |
| DoF-4 D−B institution confound | A5.1 | Declared as an assumption on H6, with a §9 row for its failure |
One judgement call was deliberately not taken from the external review: it proposed replacing H5 with a run-level contrast outright. That would have swapped an assignment-based comparison for an outcome-conditional one and abandoned v0.5's frozen paired-by-seed statistics, so the split form (A3) was frozen instead — the causal contrast stays on assignment, and the run-level split is retained as explicitly descriptive.
Arm F's removal was chosen over an exogenous freeze-time credit trigger: D and E answer the activation question directly, and removal closes the discretion completely rather than bounding it.