Model-0
26,653,440 parameters. Independent model birth #1. Official frozen result: STOP, 0 / 8 measurable axes.
This page publishes the official verdict first, then the post-hoc forensic work that examined how the measuring instrument behaved. The forensic work does not change the verdict.
Official result
- Model
- Model-0
- Scientific role
- Independent model birth #1
- Parameters
- 26,653,440
- Architecture
- 8 layers, d=384, 6 heads, FF=1536
- Training steps
- 32,470
- Tokens
- 531,988,480
- Final validation CE
- 2.7905120849609375
- Decision
- STOP
- Measurable axes
- 0 / 8
- Population
- not started
- Thresholds
- unchanged
- Calibration results were used for d* selection only; they did not enter the inferential gate family.
- Confirmation multiplicity family: A∩B replication across the 8 locked axes.
- Within-axis rule: max(p_A, p_B).
- Across-axis correction: Holm-Bonferroni.
This result was not relaxed in order to accommodate interesting signals seen afterwards.
Calibration — selection, not verdict
Calibration selected d* candidates. It was not itself the final inferential gate.
Calibration is part of the official pipeline, but it is the stage that chooses which difficulty level is structurally usable. It does not decide the outcome, and its results did not enter the inferential gate family. It should not be read as epistemically equivalent to the frozen STOP-v1 result.
| Axis | Level | hits/n | Accuracy | Chance | Normalized | Wilson 95% |
|---|---|---|---|---|---|---|
| M | 4 | 293/800 | 0.36625 | 0.25000 | +0.155 | [0.33358, 0.40020] |
| M | 8 | 286/800 | 0.35750 | 0.25000 | +0.143 | [0.32504, 0.39132] |
| M | 16 | 225/800 | 0.28125 | 0.25000 | +0.042 | [0.25120, 0.31339] |
| M | 32 | 228/800 | 0.28500 | 0.25000 | +0.047 | [0.25480, 0.31725] |
| M | 64 | 230/800 | 0.28750 | 0.25000 | +0.050 | [0.25721, 0.31982] |
| R | 1 | 412/800 | 0.51500 | 0.50000 | +0.030 | [0.48038, 0.54948] |
| R | 3 | 221/800 | 0.27625 | 0.25000 | +0.035 | [0.24639, 0.30825] |
| R | 7 | 109/800 | 0.13625 | 0.12500 | +0.013 | [0.11421, 0.16177] |
| R | 15 | 57/800 | 0.07125 | 0.06250 | +0.009 | [0.05540, 0.09120] |
| I | 8 | 158/800 | 0.19750 | 0.03125 | +0.172 | [0.17139, 0.22651] |
| I | 4 | 153/800 | 0.19125 | 0.03125 | +0.165 | [0.16550, 0.21995] |
| I | 2 | 70/800 | 0.08750 | 0.03125 | +0.058 | [0.06984, 0.10910] |
| I | 1 | 33/800 | 0.04125 | 0.03125 | +0.010 | [0.02952, 0.05736] |
| C | True | 33/800 | 0.04125 | 0.03125 | +0.010 | [0.02952, 0.05736] |
| C | False | 20/800 | 0.02500 | 0.03125 | −0.006 | [0.01624, 0.03830] |
| Q | 1 | 16/800 | 0.02000 | 0.06201 | −0.045 | [0.01235, 0.03224] |
| Q | 2 | 14/800 | 0.01750 | 0.05813 | −0.043 | [0.01045, 0.02916] |
| Q | 3 | 13/800 | 0.01625 | 0.05501 | −0.041 | [0.00952, 0.02760] |
| Q | 4 | 11/800 | 0.01375 | 0.05218 | −0.041 | [0.00769, 0.02445] |
| V | 3 | 282/800 | 0.35250 | 0.33333 | +0.029 | [0.32017, 0.38624] |
| V | 5 | 167/800 | 0.20875 | 0.20000 | +0.011 | [0.18201, 0.23827] |
| V | 8 | 101/800 | 0.12625 | 0.12500 | +0.001 | [0.10501, 0.15107] |
| A | 0.0 | 127/800 | 0.15875 | 0.06250 | +0.103 | [0.13506, 0.18570] |
| A | 0.25 | 121/800 | 0.15125 | 0.06250 | +0.095 | [0.12809, 0.17774] |
| A | 0.5 | 85/800 | 0.10625 | 0.06250 | +0.047 | [0.08675, 0.12952] |
| A | 0.75 | 53/800 | 0.06625 | 0.06250 | +0.004 | [0.05100, 0.08564] |
| T | length | 153/800 | 0.19125 | 0.16667 | +0.029 | [0.16550, 0.21995] |
| T | alphabet | 235/800 | 0.29375 | 0.25000 | +0.058 | [0.26323, 0.32624] |
| T | modulus | 72/800 | 0.09000 | 0.01562 | +0.076 | [0.07208, 0.11184] |
| T | composition | 18/800 | 0.02250 | 0.03125 | −0.009 | [0.01428, 0.03529] |
The normalized score is (accuracy − chance) / (1 − chance). It is printed as a number and deliberately not plotted: the source supplies the Wilson interval on the accuracy scale, and silently converting that interval to the normalized scale would be a derivation the source does not provide.
Confirmation — replicated signal, and what it did not mean
In plain language
The model showed a repeatable signal on three of the eight axes. The same effect appeared on two completely separate sets of questions, so it was not a fluke of one question set.
It still did not pass. Passing required the model to also get measurably worse as the questions got harder, by a margin fixed in advance. That did not happen — and on one axis the model did better on the harder level, which is the opposite of what a real difficulty response looks like.
Technical evidence
M, I and A replicated above-chance signal across the disjoint forms A and B, with Holm-adjusted p values down to 1.60×10⁻¹³.
The frozen condition MIN_DIFFICULTY_DROP = 0.10 was applied to the locked d* neighbours inside confirmation_A. M reached +0.01500 and A reached +0.07333, both below the bar. I reached −0.02065, reversing direction.
On the M, I and A axes there is above-chance signal that repeated across two disjoint item forms. None of the three axes met every condition of the frozen gate.
| Axis | Form A | Form B | Holm-adjusted p | Frozen difficulty drop | Threshold 0.10 | Outcome |
|---|---|---|---|---|---|---|
| M | 288/800 | 290/800 | 9.26×10⁻⁴ | +0.015 | not met | Signal replicated; the frozen 0.10 drop was not met |
| I | 131/800 | 156/800 | 1.60×10⁻¹³ | −0.021 | not met | Signal replicated; the difficulty direction reversed |
| A | 161/800 | 165/800 | 1.83×10⁻¹² | +0.073 | not met | Signal replicated; the frozen 0.10 drop was not met |
| Official result | STOP — 0 / 8 measurable axes. Source §5.1. | |||||
Replicated signal does not mean the gate was passed
Correct: On the M, I and A axes there is above-chance signal that repeated across two disjoint item forms. None of the three axes met every condition of the frozen gate.
Incorrect: “The M, I and A axes passed the gate.”
Post-hoc forensic: difficulty-scale audit
After the decision was final, the measuring instrument itself was audited. This work describes how difficulty behaved; it does not reopen the verdict.
| Axis | Slope β | Raw p | Exploratory Holm p | Global normalized drop | Reading |
|---|---|---|---|---|---|
| M | -0.10687 | 4.96×10⁻⁶ | 2.48×10⁻⁵ | +0.10500 | Global ordinal signal present; the locked neighbouring drop is insufficient |
| R | +0.02405 | 0.706 | 0.852 | approx. +0.021 | No ordinal evidence |
| I | -0.54215 | 8.03×10⁻²⁶ | 5.62×10⁻²⁵ | +0.16129 | Strong global trend; reversed direction at the locked neighbour |
| C | -0.51758 | 0.0361 | 0.144 | — | Inconclusive after correction |
| Q | -0.05998 | 0.313 | 0.852 | -0.00424 | No ordinal evidence; floor |
| V | -0.03616 | 0.284 | 0.852 | small | No ordinal evidence |
| A | -0.31090 | 2.96×10⁻¹⁰ | 1.77×10⁻⁹ | +0.09867 | Strong trend; just below the frozen 0.10 boundary |
| T | not applicable | not applicable | not applicable | not applicable | Four categories are not ordinal levels |
A global difficulty trend is not the locked local confirmation drop
MIN_DIFFICULTY_DROP = 0.10 was applied to the locked d* neighbours inside confirmation_A, not to the global easiest-to-hardest calibration difference. The global trend is forensic description only and does not change the official decision.
This distinction is why I can show a very strong global trend (β = −0.54215, Holm p = 5.62×10⁻²⁵) and still fail the frozen condition: the contract asked a narrower question, about specific locked neighbours, and the answer there was negative.
Post-hoc axis classifications
All eight axes received a forensic class. These describe why each axis did not become measurable — a failure of the difficulty scale and a failure of model capacity are different diagnoses with different remedies.
| Axis | Forensic class | Reason |
|---|---|---|
| M | DIFFICULTY_SCALE_FAIL | Real signal and a global decline are present; the locked local drop is not suitable |
| R | MODEL_CAPACITY_FAIL | Despite M working in the same answer format, R sits around chance/floor |
| I | DIFFICULTY_SCALE_FAIL | Strong signal present; the level ordering does not behave locally in the expected direction |
| C | MODEL_CAPACITY_FAIL | At floor despite being in the same numeric scorer family as I |
| Q | SCORING_OR_FORMAT_SUSPECT | Below chance in the official form; a format change moves it near chance but shows no positive capacity |
| V | MODEL_CAPACITY_FAIL | In the format family related to I/C, but the signal is very low |
| A | MEASURABLE_SIGNAL | Clear ordinal signal; slightly below the frozen 0.10 neighbour threshold |
| T | DIFFICULTY_SCALE_FAIL | length/alphabet/modulus/composition are not the same ordinal ladder |
These are post-hoc forensic classes. They do not replace the STOP-v1 result.
Format ablation
- Cells
- 110
- Items per cell
- 800
- Evaluation events
- 88,000
- Purpose
- newline boundary · answer hint · indexed choices · candidate order · scorer / tie behaviour
NOT YET PUBLISHED FROM SOURCE ARTIFACT
Cell-level results for this family are not present in the evidence inventory. They are therefore NOT YET PUBLISHED FROM SOURCE ARTIFACT. A higher event count does not mean more independent scientific replication.
Only the event count, the purpose, the authority and this limitation are published. No cell-level chart or table is shown, because inventing values to fill one would be exactly the failure this site exists to avoid.
Source: Scientific Evidence Inventory v1.0 (evidence cutoff 30 August 2026), §4, §5.1, §5.2, §5.3, §5.4, §5.5, §8.3.
Checkpoint availability
The Model-0 final checkpoint is not available
File not found in current local roots; hash and expected size on record.
- Expected size
- 319,957,754 bytes
- SHA-256
f3ae67d55c1e8f24ae5069e77315aa3dfab3174c68fe378b65ae69b8ca99407e- Status
- NOT FOUND
New ablation or re-evaluation of Model-0 is not possible until the checkpoint is recovered with a matching hash. Nothing on this page is offered as a download. See Evidence.