Counting and independence

There are at least 193,440 item-cell evaluation events behind this research. This page explains exactly what that number is, and what it is not.

What an evaluation event is

Scoring one model checkpoint on one item, at one difficulty level of one axis, is a single evaluation event.

30 cells × 800 items
24,000 evaluation events source §2.1
8 checkpoints × 30 cells × 128 items
30,720 evaluation events source §8.4

Neither figure means 24,000 or 30,720 independent models, and neither means that many independent experiments.

The full evaluation inventory

At least 193,440 item-cell evaluation events

This is a count of item-cell evaluation events. It is not an independent sample size, not 193,440 independent experiments, and not 193,440 independent models. It includes the same frozen items being re-measured at different checkpoints, in different forms, and under different format variants.

Evaluation inventory by data package, in absolute event counts Six data packages sized by absolute item-cell evaluation event count. The bars are absolute counts, not shares of a whole, because the packages carry different evidence authorities and are not parts of one comparable total. The largest bar is the eighty-eight thousand event format ablation family, which is post-hoc forensic work outside the decision gate.item-cell evaluation events022,00044,00066,00088,000 Model-0 calibration □ OFFICIAL CHAIN — CALIBRATION / SELECTION Model-0 calibration — 30 cells × 800 items = 24,000 item-cell evaluation events; Official chain — calibration / selection stage 24,000 Model-0 confirmation ■ OFFICIAL / FROZEN Model-0 confirmation — 3 axes × (2 A levels + 1 B retest) × 800 = 7,200 item-cell evaluation events; Official / frozen result 7,200 Model-0 format ablation ▲ POST-HOC FORENSIC Model-0 format ablation — 110 format cells × 800 items = 88,000 item-cell evaluation events; Post-hoc forensic 88,000 Pilot-0.2 longitudinal Birth Map ● DIAGNOSTIC-ONLY Pilot-0.2 longitudinal Birth Map — 8 checkpoints × 30 cells × 128 items = 30,720 item-cell evaluation events; Diagnostic-only 30,720 Pilot-0.2 final decision chain ◆ COMPATIBILITY-ADAPTER Pilot-0.2 final decision chain — (30 calibration + 10 confirmation) × 800 = 32,000 item-cell evaluation events; Compatibility-adapter 32,000 Pilot-0.3 continuation scan ● DIAGNOSTIC-ONLY Pilot-0.3 continuation scan — 3 checkpoints × 30 cells × 128 items = 11,520 item-cell evaluation events; Diagnostic-only 11,520
Absolute event counts per data package. Deliberately not percentage-normalised and never a pie chart: the packages carry different evidence authorities and are not slices of one comparable whole. Source §4.
Evaluation inventory — source §4. Counts are item-cell evaluation events, not independent samples.
PackageCalculationEvent countEvidence authoritySample semantics
Model-0 calibration30 cells × 800 items24,000 OFFICIAL CHAIN — CALIBRATION / SELECTIONSelection stage of the official chain
Model-0 confirmation3 axes × (2 A levels + 1 B retest) × 8007,200 OFFICIAL / FROZENOfficial inferential confirmation
Model-0 format ablation110 format cells × 800 items88,000 POST-HOC FORENSICPost-hoc, outside the gate
Pilot-0.2 longitudinal Birth Map8 checkpoints × 30 cells × 128 items30,720 DIAGNOSTIC-ONLYMatched items re-measured across checkpoints
Pilot-0.2 final decision chain(30 calibration + 10 confirmation) × 80032,000 COMPATIBILITY-ADAPTERcanonical_stop_v1_artifact: false
Pilot-0.3 continuation scan3 checkpoints × 30 cells × 128 items11,520 DIAGNOSTIC-ONLYMatched items
Minimum totalat least 193,440Mixed authorityNot a single n; not an independent sample size

The total does not separately add candidate-order and batch-perturbation records, so nested derivatives of the same computation are not counted twice into a headline number.

Why the total is not an independent n

One checkpoint contains 30 cells × 128 = 3,840 item-cell observations. The 30,720 events across eight checkpoints cannot be treated as eight separate independent samples of 3,840.

confirmation_A and confirmation_B are disjoint item forms with zero surface and true-latent hash intersection. They test the same model checkpoint. They are not two independent models and not two independent initializations.

Running the full n=800 decision chain at all eight checkpoints would have required 8 × 30 × 800 = 192,000 evaluations for calibration alone, before confirmation and format controls. The design deliberately separates the developmental scan from the official inferential test.

How to read a large event count

A larger event count is not more independent replication

The single largest package in the inventory is the 88,000-event format ablation family. It is post-hoc forensic work that sits outside the decision gate, and it exists to probe the measuring instrument rather than to test the model.

That is why the Model-0 raw data file is far larger than the others, and why file size is not a proxy for scientific weight anywhere on this site.

Source: Scientific Evidence Inventory v1.0 (evidence cutoff 30 August 2026), §2.1, §2.2, §2.3, §4, §8.2, §8.3, §8.4.