Companion code for the essay series On Measurement and Honest Growth in AI Systems by Matthew Childs (July 2026).
The series argues that AI capability measurement needs two-sided error discipline — policing not just inflated claims, but the silent false negatives produced by broken graders — and that the only headline worth trusting for a "compounding" system is a slope of verified reuse, never a count of stored skills.
| # | Essay | Code |
|---|---|---|
| 1 | The Coin Economy: Seeding Towards Generality | planned — transfer instrument + control battery on public atmospheric data |
| 2 | Knowing What You're Good At Beats Being Curious | planned — allocation policies + synthetic portfolio demo |
| 3 | Measurable Compounding Invention Towards Generality | article3-depth-dial/ — the controlled depth dose-response experiment (Figure 5 + budget sweep) |
| 4 | Mutual Training: What Parametric and Non-Parametric Systems Owe Each Other | planned — F6/F7 pre-registered experiment scaffolds |
Essay links will be added here when the series is published.
Everything in this repository is self-contained, runs on one CPU in minutes, uses only the Python standard library unless a directory's README says otherwise, and reproduces the stated figures exactly (seeds are fixed).
The natural-domain systems described in the essays are deliberately reported as sanitized ratios and are not released; the essays explain why. What is released is every controlled experiment the essays claim is reproducible end to end.
MIT — see LICENSE.
Corrections, failed replications, and counter-examples are actively welcome — the series is about exactly that. Open an issue, or email matthew.childs@myyahoo.com.
ORCID: 0009-0007-9258-2579