Methods
The measurement contract that every result on this site is produced under.
Why this exists
These pages define the measurement vocabulary once. Without them, each run page would have to restate the same scientific terms, and small restatements are where meaning quietly drifts.
Nothing here is a result. Everything here is the contract under which the results were produced.
The contract
- The eight axes
What M, R, I, C, Q, V, A and T denote, the answer-space family each belongs to, why their chance levels differ, and why T is not an ordinal difficulty ladder.
- The frozen decision contract
STOP-v1: calibration selects d*, confirmation replicates across two disjoint forms, thresholds are locked in advance, and the protocol fails closed.
- Evidence authority
Six epistemic roles a measurement can hold. These are roles, not quality grades.
- Counting and independence
What an item-cell evaluation event is, and why a large event count is not a large independent sample.
What exactly is being measured
In plain language
We give a model a set of fixed questions and score whether each answer is right. We do this at several points during training, using the same questions each time, so that a change in the score reflects a change in the model rather than a change in the test.
A score on its own means very little. A model can get a question right by guessing, and different question types are guessable to very different degrees. So each result is always compared against the chance level of that specific question type.
Technical evidence
An evaluation cell is one axis at one difficulty level. Scoring one checkpoint on one item within one cell is a single item-cell evaluation event.
Matched items are held fixed across checkpoints so that longitudinal comparison is not confounded by item resampling. The consequence is that repeated measurements are not independent replications.
The normalized score is (accuracy − chance) / (1 − chance), which places zero at chance for every axis regardless of its answer space. The normalized score is a comparison aid, not an inferential quantity: intervals and tests are computed on the accuracy scale over hits / n_items.
Diagnostic and inferential roles are not interchangeable
Two measurement layers exist and they answer different questions.
n=128 is the diagnostic layer that scans many checkpoints and 30 cells at reasonable cost. n=800 is the layer that strengthens confidence intervals and exact binomial inference on a narrower set of final/confirmation tests.
A diagnostic scan can describe what happened in a run. It cannot deliver a verdict, because it never entered a decision gate. Conflating the two is the single most common way a research surface overstates itself, so the two layers are labelled separately everywhere on this site.
Source: Scientific Evidence Inventory v1.0 (evidence cutoff 30 August 2026), §2, §5, §8.