Counting and independence
There are at least 193,440 item-cell evaluation events behind this research. This page explains exactly what that number is, and what it is not.
What an evaluation event is
Scoring one model checkpoint on one item, at one difficulty level of one axis, is a single evaluation event.
- 30 cells × 800 items
- 24,000 evaluation events source §2.1
- 8 checkpoints × 30 cells × 128 items
- 30,720 evaluation events source §8.4
Neither figure means 24,000 or 30,720 independent models, and neither means that many independent experiments.
The full evaluation inventory
At least 193,440 item-cell evaluation events
This is a count of item-cell evaluation events. It is not an independent sample size, not 193,440 independent experiments, and not 193,440 independent models. It includes the same frozen items being re-measured at different checkpoints, in different forms, and under different format variants.
| Package | Calculation | Event count | Evidence authority | Sample semantics |
|---|---|---|---|---|
| Model-0 calibration | 30 cells × 800 items | 24,000 | OFFICIAL CHAIN — CALIBRATION / SELECTION | Selection stage of the official chain |
| Model-0 confirmation | 3 axes × (2 A levels + 1 B retest) × 800 | 7,200 | OFFICIAL / FROZEN | Official inferential confirmation |
| Model-0 format ablation | 110 format cells × 800 items | 88,000 | POST-HOC FORENSIC | Post-hoc, outside the gate |
| Pilot-0.2 longitudinal Birth Map | 8 checkpoints × 30 cells × 128 items | 30,720 | DIAGNOSTIC-ONLY | Matched items re-measured across checkpoints |
| Pilot-0.2 final decision chain | (30 calibration + 10 confirmation) × 800 | 32,000 | COMPATIBILITY-ADAPTER | canonical_stop_v1_artifact: false |
| Pilot-0.3 continuation scan | 3 checkpoints × 30 cells × 128 items | 11,520 | DIAGNOSTIC-ONLY | Matched items |
| Minimum total | — | at least 193,440 | Mixed authority | Not a single n; not an independent sample size |
The total does not separately add candidate-order and batch-perturbation records, so nested derivatives of the same computation are not counted twice into a headline number.
Why the total is not an independent n
- The inferential unit of the A axis is the item. Multiple positions inside A are not independent binomial observations.
- Wilson intervals and binomial tests are computed over hits / n_items. Position-level accuracy is descriptive only.
- Measuring the same matched item set at different checkpoints is a matched longitudinal measurement across checkpoints, not independent replication.
One checkpoint contains 30 cells × 128 = 3,840 item-cell observations. The 30,720 events across eight checkpoints cannot be treated as eight separate independent samples of 3,840.
confirmation_A and confirmation_B are disjoint item forms with zero surface and true-latent hash intersection. They test the same model checkpoint. They are not two independent models and not two independent initializations.
Running the full n=800 decision chain at all eight checkpoints would have required 8 × 30 × 800 = 192,000 evaluations for calibration alone, before confirmation and format controls. The design deliberately separates the developmental scan from the official inferential test.
How to read a large event count
A larger event count is not more independent replication
The single largest package in the inventory is the 88,000-event format ablation family. It is post-hoc forensic work that sits outside the decision gate, and it exists to probe the measuring instrument rather than to test the model.
That is why the Model-0 raw data file is far larger than the others, and why file size is not a proxy for scientific weight anywhere on this site.
Source: Scientific Evidence Inventory v1.0 (evidence cutoff 30 August 2026), §2.1, §2.2, §2.3, §4, §8.2, §8.3, §8.4.