Methods

The measurement contract that every result on this site is produced under.

Why this exists

These pages define the measurement vocabulary once. Without them, each run page would have to restate the same scientific terms, and small restatements are where meaning quietly drifts.

Nothing here is a result. Everything here is the contract under which the results were produced.

The contract

What exactly is being measured

In plain language

We give a model a set of fixed questions and score whether each answer is right. We do this at several points during training, using the same questions each time, so that a change in the score reflects a change in the model rather than a change in the test.

A score on its own means very little. A model can get a question right by guessing, and different question types are guessable to very different degrees. So each result is always compared against the chance level of that specific question type.

Technical evidence

An evaluation cell is one axis at one difficulty level. Scoring one checkpoint on one item within one cell is a single item-cell evaluation event.

Matched items are held fixed across checkpoints so that longitudinal comparison is not confounded by item resampling. The consequence is that repeated measurements are not independent replications.

The normalized score is (accuracy − chance) / (1 − chance), which places zero at chance for every axis regardless of its answer space. The normalized score is a comparison aid, not an inferential quantity: intervals and tests are computed on the accuracy scale over hits / n_items.

Diagnostic and inferential roles are not interchangeable

Two measurement layers exist and they answer different questions.

n=128 is the diagnostic layer that scans many checkpoints and 30 cells at reasonable cost. n=800 is the layer that strengthens confidence intervals and exact binomial inference on a narrower set of final/confirmation tests.

A diagnostic scan can describe what happened in a run. It cannot deliver a verdict, because it never entered a decision gate. Conflating the two is the single most common way a research surface overstates itself, so the two layers are labelled separately everywhere on this site.

Source: Scientific Evidence Inventory v1.0 (evidence cutoff 30 August 2026), §2, §5, §8.