Limits

What the current evidence supports, what it only hints at, and what it does not support at all.

This page is scientific integrity content, not a disclaimer. The thirteen unsupported claims are listed because they are the misreadings the results most naturally invite.

What we can claim

Nine findings that are strongly defensible on the current evidence.

  1. Model-0 did not pass the official STOP-v1 gate.
  2. In Model-0, M, I and A produced above-chance signal in two disjoint confirmation forms.
  3. The frozen difficulty-drop condition prevented those signals from being declared measurable axes.
  4. The current four categories of T are not an ordinal difficulty ladder.
  5. The Q measurement carries format/scoring suspicion, but the available ablation showed no positive Q capacity.
  6. In Pilot-0.2, the A signal became distinct before the I signal.
  7. In Pilot-0.2, A and I peaked before the final checkpoint.
  8. From Pilot-0.2 to Pilot-0.3, validation CE improved while I and A declined and easy R rose.
  9. The general training metric and the eight-axis profile did not follow the same developmental curve.

What is exploratory

Five observations that generate hypotheses and are not results.

  1. There may be a developmental order in which A forms early and I forms later.
  2. T-modulus behaviour may be a separate, late-developing transfer subtype.
  3. Continued training may strengthen some axes while weakening others.
  4. The fixed V scores may point to scorer or task insensitivity.
  5. In small models, the assumption that lower validation loss means higher scores on every capability may not hold.

These items generate hypotheses because they rest on a single run and repeated matched items. They are not population results.

What we do not claim

Thirteen statements that this research does not support. They are listed explicitly because several of them are the natural misreadings of the findings above.

  1. The initialization seed determined the cognitive phenotype.
  2. A particular seed produces an M/I/A or T/Q/V model.
  3. The theory has been proven.
  4. Pilot-0.3 is an independent third model.
  5. Two A/B forms are two independent model replications.
  6. M, I and A passed the gate.
  7. Q reasoning capacity was found.
  8. T failed because of model capacity.
  9. The inferential n of the A axis is 6,400 positions.
  10. The 69.97M model has 88.9M parameters.
  11. Model-1..5 were trained, or the population was completed.
  12. Retrospective Observatory records are a canonical birth ledger.
  13. LAB_READY means a scientific result or a confirmed theory.

These are not disclaimers

This list is not legal boilerplate and it is not modesty. Each line is a specific claim that someone could reasonably infer from the published results, and each one is unsupported by the evidence behind them.

Unmeasured and missing evidence

Next evidence needed

Ten requirements, in priority order, for moving these findings toward population and causal claims.

  1. Hash-matched recovery of the Model-0 checkpoint.
  2. Running the same frozen n=128 diagnostic scan on the Pilot-0.3 100% final checkpoint.
  3. A separately pre-written n=800 confirmation contract for Pilot-0.3 if one is needed, without retrofitting the existing STOP-v1 result.
  4. Actually training Model-1..5 or a new canonical cohort.
  5. A population in which only init_seed varies under the same architecture, data, data order and training contract.
  6. A matched-item mixed-effects / binomial model for longitudinal change, with a pre-defined multiplicity family.
  7. Splitting the T axis into category-level subtests instead of using a false ordinal order.
  8. An independent, format-invariant measurement contract for Q.
  9. Reference and random baselines, and where possible external model controls on the same battery.
  10. New Model Birth Observatory runs that begin with a canonical θ₀ record.

Closing scientific verdict

The current data does not prove the Cognitive Birth approach as a whole, nor initialization-to-phenotype causality. The strongest available inference is not "a given seed produces a given capability".

Three concrete outcomes were produced:

The narrower claim that is currently defensible

Model development may not be adequately represented by final loss or a single benchmark score alone. Within the same model, different measurement axes can rise and fall at different times, and the difficulty and format structure of the measuring instrument can materially affect the observed phenotype.

Reaching the population and causality level requires an independent initialization cohort, canonical Observatory records, and pre-defined longitudinal inference.

Source: Scientific Evidence Inventory v1.0 (evidence cutoff 30 August 2026), §10.1, §10.2, §10.3, §11, §15.