Model-0

26,653,440 parameters. Independent model birth #1. Official frozen result: STOP, 0 / 8 measurable axes.

This page publishes the official verdict first, then the post-hoc forensic work that examined how the measuring instrument behaved. The forensic work does not change the verdict.

Official result

Evidence authority: Official / frozen result

STOP

0 / 8 measurable axes

The official, frozen STOP-v1 result for Model-0. Population not started. Thresholds unchanged. This verdict was not relaxed to accommodate the replicated signals described further down this page.

Model
Model-0
Scientific role
Independent model birth #1
Parameters
26,653,440
Architecture
8 layers, d=384, 6 heads, FF=1536
Training steps
32,470
Tokens
531,988,480
Final validation CE
2.7905120849609375
Decision
STOP
Measurable axes
0 / 8
Population
not started
Thresholds
unchanged

This result was not relaxed in order to accommodate interesting signals seen afterwards.

Calibration — selection, not verdict

Evidence authority: Official chain — calibration / selection stage

Calibration selected d* candidates. It was not itself the final inferential gate.

Calibration is part of the official pipeline, but it is the stage that chooses which difficulty level is structurally usable. It does not decide the outcome, and its results did not enter the inferential gate family. It should not be read as epistemically equivalent to the frozen STOP-v1 result.

Model-0 calibration: accuracy with Wilson 95% intervals against per-cell chance Thirty calibration cells. Each row shows the measured accuracy as a dot, the source Wilson ninety-five percent interval as a horizontal bar on the accuracy scale, and that cell’s own chance level as a short vertical tick. Where the chance tick falls inside the interval, the cell is not separated from chance. The normalized score is printed as a number and is not plotted, because converting the interval to the normalized scale would be a derivation the source does not supply.accuracy0.00.10.20.30.40.5 M @ 4 M @ 4 — 293/800, accuracy 0.36625, chance 0.25000, Wilson 95% [0.33358, 0.40020], normalized +0.155 293/800 · norm +0.155 M @ 8 M @ 8 — 286/800, accuracy 0.35750, chance 0.25000, Wilson 95% [0.32504, 0.39132], normalized +0.143 286/800 · norm +0.143 M @ 16 M @ 16 — 225/800, accuracy 0.28125, chance 0.25000, Wilson 95% [0.25120, 0.31339], normalized +0.042 225/800 · norm +0.042 M @ 32 M @ 32 — 228/800, accuracy 0.28500, chance 0.25000, Wilson 95% [0.25480, 0.31725], normalized +0.047 228/800 · norm +0.047 M @ 64 M @ 64 — 230/800, accuracy 0.28750, chance 0.25000, Wilson 95% [0.25721, 0.31982], normalized +0.050 230/800 · norm +0.050 R @ 1 R @ 1 — 412/800, accuracy 0.51500, chance 0.50000, Wilson 95% [0.48038, 0.54948], normalized +0.030 412/800 · norm +0.030 R @ 3 R @ 3 — 221/800, accuracy 0.27625, chance 0.25000, Wilson 95% [0.24639, 0.30825], normalized +0.035 221/800 · norm +0.035 R @ 7 R @ 7 — 109/800, accuracy 0.13625, chance 0.12500, Wilson 95% [0.11421, 0.16177], normalized +0.013 109/800 · norm +0.013 R @ 15 R @ 15 — 57/800, accuracy 0.07125, chance 0.06250, Wilson 95% [0.05540, 0.09120], normalized +0.009 57/800 · norm +0.009 I @ 8 I @ 8 — 158/800, accuracy 0.19750, chance 0.03125, Wilson 95% [0.17139, 0.22651], normalized +0.172 158/800 · norm +0.172 I @ 4 I @ 4 — 153/800, accuracy 0.19125, chance 0.03125, Wilson 95% [0.16550, 0.21995], normalized +0.165 153/800 · norm +0.165 I @ 2 I @ 2 — 70/800, accuracy 0.08750, chance 0.03125, Wilson 95% [0.06984, 0.10910], normalized +0.058 70/800 · norm +0.058 I @ 1 I @ 1 — 33/800, accuracy 0.04125, chance 0.03125, Wilson 95% [0.02952, 0.05736], normalized +0.010 33/800 · norm +0.010 C @ True C @ True — 33/800, accuracy 0.04125, chance 0.03125, Wilson 95% [0.02952, 0.05736], normalized +0.010 33/800 · norm +0.010 C @ False C @ False — 20/800, accuracy 0.02500, chance 0.03125, Wilson 95% [0.01624, 0.03830], normalized −0.006 20/800 · norm −0.006 Q @ 1 Q @ 1 — 16/800, accuracy 0.02000, chance 0.06201, Wilson 95% [0.01235, 0.03224], normalized −0.045 16/800 · norm −0.045 Q @ 2 Q @ 2 — 14/800, accuracy 0.01750, chance 0.05813, Wilson 95% [0.01045, 0.02916], normalized −0.043 14/800 · norm −0.043 Q @ 3 Q @ 3 — 13/800, accuracy 0.01625, chance 0.05501, Wilson 95% [0.00952, 0.02760], normalized −0.041 13/800 · norm −0.041 Q @ 4 Q @ 4 — 11/800, accuracy 0.01375, chance 0.05218, Wilson 95% [0.00769, 0.02445], normalized −0.041 11/800 · norm −0.041 V @ 3 V @ 3 — 282/800, accuracy 0.35250, chance 0.33333, Wilson 95% [0.32017, 0.38624], normalized +0.029 282/800 · norm +0.029 V @ 5 V @ 5 — 167/800, accuracy 0.20875, chance 0.20000, Wilson 95% [0.18201, 0.23827], normalized +0.011 167/800 · norm +0.011 V @ 8 V @ 8 — 101/800, accuracy 0.12625, chance 0.12500, Wilson 95% [0.10501, 0.15107], normalized +0.001 101/800 · norm +0.001 A @ 0.0 A @ 0.0 — 127/800, accuracy 0.15875, chance 0.06250, Wilson 95% [0.13506, 0.18570], normalized +0.103 127/800 · norm +0.103 A @ 0.25 A @ 0.25 — 121/800, accuracy 0.15125, chance 0.06250, Wilson 95% [0.12809, 0.17774], normalized +0.095 121/800 · norm +0.095 A @ 0.5 A @ 0.5 — 85/800, accuracy 0.10625, chance 0.06250, Wilson 95% [0.08675, 0.12952], normalized +0.047 85/800 · norm +0.047 A @ 0.75 A @ 0.75 — 53/800, accuracy 0.06625, chance 0.06250, Wilson 95% [0.05100, 0.08564], normalized +0.004 53/800 · norm +0.004 T @ length T @ length — 153/800, accuracy 0.19125, chance 0.16667, Wilson 95% [0.16550, 0.21995], normalized +0.029 153/800 · norm +0.029 T @ alphabet T @ alphabet — 235/800, accuracy 0.29375, chance 0.25000, Wilson 95% [0.26323, 0.32624], normalized +0.058 235/800 · norm +0.058 T @ modulus T @ modulus — 72/800, accuracy 0.09000, chance 0.01562, Wilson 95% [0.07208, 0.11184], normalized +0.076 72/800 · norm +0.076 T @ composition T @ composition — 18/800, accuracy 0.02250, chance 0.03125, Wilson 95% [0.01428, 0.03529], normalized −0.009 18/800 · norm −0.009
Each row is one calibration cell. The dot is accuracy, the bar is the Wilson 95% interval from the source, and the short vertical tick is that cell’s own chance level — chance is not shared across axes, it ranges from 0.01562 to 0.50000. The normalized score is shown as text only. Source §5.2.
Model-0 calibration, all 30 cells — source §5.2, official chain, calibration / selection stage, n=800 per cell. Wilson intervals are on the accuracy scale exactly as the source gives them.
AxisLevelhits/nAccuracyChanceNormalizedWilson 95%
M4293/8000.366250.25000+0.155[0.33358, 0.40020]
M8286/8000.357500.25000+0.143[0.32504, 0.39132]
M16225/8000.281250.25000+0.042[0.25120, 0.31339]
M32228/8000.285000.25000+0.047[0.25480, 0.31725]
M64230/8000.287500.25000+0.050[0.25721, 0.31982]
R1412/8000.515000.50000+0.030[0.48038, 0.54948]
R3221/8000.276250.25000+0.035[0.24639, 0.30825]
R7109/8000.136250.12500+0.013[0.11421, 0.16177]
R1557/8000.071250.06250+0.009[0.05540, 0.09120]
I8158/8000.197500.03125+0.172[0.17139, 0.22651]
I4153/8000.191250.03125+0.165[0.16550, 0.21995]
I270/8000.087500.03125+0.058[0.06984, 0.10910]
I133/8000.041250.03125+0.010[0.02952, 0.05736]
CTrue33/8000.041250.03125+0.010[0.02952, 0.05736]
CFalse20/8000.025000.03125−0.006[0.01624, 0.03830]
Q116/8000.020000.06201−0.045[0.01235, 0.03224]
Q214/8000.017500.05813−0.043[0.01045, 0.02916]
Q313/8000.016250.05501−0.041[0.00952, 0.02760]
Q411/8000.013750.05218−0.041[0.00769, 0.02445]
V3282/8000.352500.33333+0.029[0.32017, 0.38624]
V5167/8000.208750.20000+0.011[0.18201, 0.23827]
V8101/8000.126250.12500+0.001[0.10501, 0.15107]
A0.0127/8000.158750.06250+0.103[0.13506, 0.18570]
A0.25121/8000.151250.06250+0.095[0.12809, 0.17774]
A0.585/8000.106250.06250+0.047[0.08675, 0.12952]
A0.7553/8000.066250.06250+0.004[0.05100, 0.08564]
Tlength153/8000.191250.16667+0.029[0.16550, 0.21995]
Talphabet235/8000.293750.25000+0.058[0.26323, 0.32624]
Tmodulus72/8000.090000.01562+0.076[0.07208, 0.11184]
Tcomposition18/8000.022500.03125−0.009[0.01428, 0.03529]

The normalized score is (accuracy − chance) / (1 − chance). It is printed as a number and deliberately not plotted: the source supplies the Wilson interval on the accuracy scale, and silently converting that interval to the normalized scale would be a derivation the source does not provide.

Confirmation — replicated signal, and what it did not mean

Evidence authority: Official / frozen result

In plain language

The model showed a repeatable signal on three of the eight axes. The same effect appeared on two completely separate sets of questions, so it was not a fluke of one question set.

It still did not pass. Passing required the model to also get measurably worse as the questions got harder, by a margin fixed in advance. That did not happen — and on one axis the model did better on the harder level, which is the opposite of what a real difficulty response looks like.

Technical evidence

M, I and A replicated above-chance signal across the disjoint forms A and B, with Holm-adjusted p values down to 1.60×10⁻¹³.

The frozen condition MIN_DIFFICULTY_DROP = 0.10 was applied to the locked d* neighbours inside confirmation_A. M reached +0.01500 and A reached +0.07333, both below the bar. I reached −0.02065, reversing direction.

On the M, I and A axes there is above-chance signal that repeated across two disjoint item forms. None of the three axes met every condition of the frozen gate.

Model-0 confirmation: replicated signal did not pass the frozen gate Three axes reached confirmation. Each shows its Form A and Form B hit counts, its Holm-adjusted p value, and its frozen difficulty drop measured against the locked threshold of zero point one zero. All three replicated above-chance signal across two disjoint item forms. None met the frozen drop condition, and on axis I the difficulty direction reversed. The official decision remained STOP with zero of eight measurable axes.frozen difficulty drop-0.050.000.050.10frozen threshold 0.10 M A 288/800 B 290/800 M — frozen difficulty drop +0.015; threshold 0.10 not met drop +0.015 · Holm p 9.26×10⁻⁴ Signal replicated; the frozen 0.10 drop was not met I A 131/800 B 156/800 I — frozen difficulty drop −0.021; threshold 0.10 not met drop −0.021 · Holm p 1.60×10⁻¹³ Signal replicated; the difficulty direction reversed A A 161/800 B 165/800 A — frozen difficulty drop +0.073; threshold 0.10 not met drop +0.073 · Holm p 1.83×10⁻¹² Signal replicated; the frozen 0.10 drop was not met Official result: STOP — 0 / 8 measurable axes Replicated signal on M, I and A did not become a gate pass.
Each row is one axis that reached confirmation. The dot is the frozen difficulty drop; the vertical rule is the locked 0.10 threshold. A filled dot would mean the threshold was met — none are filled. Replicated signal and a gate pass are different things. Source §5.3, official result §5.1.
Model-0 confirmation — source §5.3, official / frozen result, n=800 per form. Form A and Form B are disjoint item forms tested against the same checkpoint; they are not two independent models.
AxisForm AForm BHolm-adjusted pFrozen difficulty dropThreshold 0.10Outcome
M288/800290/8009.26×10⁻⁴+0.015not metSignal replicated; the frozen 0.10 drop was not met
I131/800156/8001.60×10⁻¹³−0.021not metSignal replicated; the difficulty direction reversed
A161/800165/8001.83×10⁻¹²+0.073not metSignal replicated; the frozen 0.10 drop was not met
Official resultSTOP — 0 / 8 measurable axes. Source §5.1.

Replicated signal does not mean the gate was passed

Correct: On the M, I and A axes there is above-chance signal that repeated across two disjoint item forms. None of the three axes met every condition of the frozen gate.

Incorrect: “The M, I and A axes passed the gate.”

Post-hoc forensic: difficulty-scale audit

Evidence authority: Post-hoc forensic

After the decision was final, the measuring instrument itself was audited. This work describes how difficulty behaved; it does not reopen the verdict.

Item-level logistic / binomial slope results — source §5.4, post-hoc forensic.
AxisSlope βRaw pExploratory Holm pGlobal normalized dropReading
M-0.106874.96×10⁻⁶2.48×10⁻⁵+0.10500Global ordinal signal present; the locked neighbouring drop is insufficient
R+0.024050.7060.852approx. +0.021No ordinal evidence
I-0.542158.03×10⁻²⁶5.62×10⁻²⁵+0.16129Strong global trend; reversed direction at the locked neighbour
C-0.517580.03610.144Inconclusive after correction
Q-0.059980.3130.852-0.00424No ordinal evidence; floor
V-0.036160.2840.852smallNo ordinal evidence
A-0.310902.96×10⁻¹⁰1.77×10⁻⁹+0.09867Strong trend; just below the frozen 0.10 boundary
Tnot applicablenot applicablenot applicablenot applicableFour categories are not ordinal levels

A global difficulty trend is not the locked local confirmation drop

MIN_DIFFICULTY_DROP = 0.10 was applied to the locked d* neighbours inside confirmation_A, not to the global easiest-to-hardest calibration difference. The global trend is forensic description only and does not change the official decision.

This distinction is why I can show a very strong global trend (β = −0.54215, Holm p = 5.62×10⁻²⁵) and still fail the frozen condition: the contract asked a narrower question, about specific locked neighbours, and the answer there was negative.

Post-hoc axis classifications

Evidence authority: Post-hoc forensic

All eight axes received a forensic class. These describe why each axis did not become measurable — a failure of the difficulty scale and a failure of model capacity are different diagnoses with different remedies.

Forensic axis classifications — source §5.5. These do not overwrite STOP-v1.
AxisForensic classReason
MDIFFICULTY_SCALE_FAILReal signal and a global decline are present; the locked local drop is not suitable
RMODEL_CAPACITY_FAILDespite M working in the same answer format, R sits around chance/floor
IDIFFICULTY_SCALE_FAILStrong signal present; the level ordering does not behave locally in the expected direction
CMODEL_CAPACITY_FAILAt floor despite being in the same numeric scorer family as I
QSCORING_OR_FORMAT_SUSPECTBelow chance in the official form; a format change moves it near chance but shows no positive capacity
VMODEL_CAPACITY_FAILIn the format family related to I/C, but the signal is very low
AMEASURABLE_SIGNALClear ordinal signal; slightly below the frozen 0.10 neighbour threshold
TDIFFICULTY_SCALE_FAILlength/alphabet/modulus/composition are not the same ordinal ladder

These are post-hoc forensic classes. They do not replace the STOP-v1 result.

Format ablation

Evidence authority: Post-hoc forensic

Cells
110
Items per cell
800
Evaluation events
88,000
Purpose
newline boundary · answer hint · indexed choices · candidate order · scorer / tie behaviour

NOT YET PUBLISHED FROM SOURCE ARTIFACT

Cell-level results for this family are not present in the evidence inventory. They are therefore NOT YET PUBLISHED FROM SOURCE ARTIFACT. A higher event count does not mean more independent scientific replication.

Only the event count, the purpose, the authority and this limitation are published. No cell-level chart or table is shown, because inventing values to fill one would be exactly the failure this site exists to avoid.

Source: Scientific Evidence Inventory v1.0 (evidence cutoff 30 August 2026), §4, §5.1, §5.2, §5.3, §5.4, §5.5, §8.3.

Checkpoint availability

The Model-0 final checkpoint is not available

File not found in current local roots; hash and expected size on record.

Expected size
319,957,754 bytes
SHA-256
f3ae67d55c1e8f24ae5069e77315aa3dfab3174c68fe378b65ae69b8ca99407e
Status
NOT FOUND

New ablation or re-evaluation of Model-0 is not possible until the checkpoint is recovered with a matching hash. Nothing on this page is offered as a download. See Evidence.