The eight axes
Eight measurement axes, their answer spaces, and the reason they cannot be compared on a single scale.
Notation, not invented names
The evidence inventory does not formally define an expanded capability name for each letter. This page therefore publishes the notation and the answer-space family that the source does define, rather than inventing names.
Naming an axis is a scientific claim about what it measures. The evidence inventory does not make that claim, so this site does not make it either. The letters are used as the source uses them.
Answer-space families
- M, R
- Two-token "v" answer family.
- I, C, V
- Single-token numeric answer family.
- Q
- Variable candidate count, with boundary and format sensitivity.
- A
- Item endpoint is primary; positions are dependent sub-observations.
- T
- Not an ordinal level set; a different out-of-distribution category family.
Forcing an equal number of confirmations across axes does not create scientific symmetry. It manufactures false independence and a false difficulty assumption.
Chance is not shared across axes
Chance is not shared across axes. In the Model-0 calibration it ranges from 0.01562 to 0.50000, so each cell is compared against its own chance level and never against a single global baseline.
| Axis | Levels measured | Chance level(s) |
|---|---|---|
| M | 4, 8, 16, 32, 64 | 0.25000 |
| R | 1, 3, 7, 15 | 0.50000, 0.25000, 0.12500, 0.06250 |
| I | 8, 4, 2, 1 | 0.03125 |
| C | True, False | 0.03125 |
| Q | 1, 2, 3, 4 | 0.06201, 0.05813, 0.05501, 0.05218 |
| V | 3, 5, 8 | 0.33333, 0.20000, 0.12500 |
| A | 0.0, 0.25, 0.5, 0.75 | 0.06250 |
| T | length, alphabet, modulus, composition | 0.16667, 0.25000, 0.01562, 0.03125 |
This is why a raw accuracy comparison across axes is meaningless here, and why no chart on this site plots two axes against a single shared baseline.
T is not an ordinal difficulty ladder
T is not a valid ordinal difficulty ladder in its current four-category construction. length, alphabet, modulus and composition are not steps on one scale.
length, alphabet, modulus and composition are different out-of-distribution category families. Treating them as ordered steps would manufacture a difficulty gradient that does not exist, and every conclusion drawn from that gradient would be an artefact of the ordering.
This has a direct consequence in the results: no d* could be established for T, so T never entered confirmation. That is the protocol behaving correctly, not missing data.
Source: Scientific Evidence Inventory v1.0 (evidence cutoff 30 August 2026), §8.9, §5.4, §5.5.