Pilot-0.2 / Local-70M-0
69,970,848 parameters. Independent model birth #2. 531,988,480 tokens over 32,470 steps, final validation CE 2.6221997812390327.
Eight checkpoints were measured on the same fixed cells, from the untrained starting state to the end of training.
Training result
- Scientific role
- Independent model birth #2
- Parameters
- 69,970,848 unique trainable parameters — not 88.9M, see below
- Tokens
- 531,988,480
- Training steps
- 32,470
- Precision
- fp16
- Final validation CE
- 2.6221997812390327
- Final checkpoint size
- 915,598,984 bytes
- Final checkpoint SHA-256
cd53d21bf67f1ab23951152542bc13119cee8532bb5fe4fb9ce93178bdb9f309
Completed through verified checkpoint recovery
Training completed through a deterministic recovery branch that resumed from verified checkpoints after interruptions.
This run is not presented as a single uninterrupted process.
This is a 69,970,848-parameter model, not an 88.9M model
Some Observatory outputs report roughly 88.9 million state elements. That figure counts the tied embedding weight twice, because it is materialised in the state dictionary under both tok_emb.weight and lm_head.weight. The unique trainable parameter count is 69,970,848. The model must not be presented as an 88.9M model.
The Birth Map
In plain language
We measured the same eight fixed cells at eight points during training, from the untrained starting state to the end.
The measured signals did not rise together. One measure (A) appeared early in this run and then fell back before the end. Another (I) stayed flat for a long time and rose late. Some measures peaked before the final checkpoint. Meanwhile the model’s overall language-modelling loss fell steadily the whole time.
So in this run, the general training metric and the measured profile were not the same thing.
Technical evidence
Matched items, n=128 per cell, held fixed across all eight checkpoints so that a change in score reflects a change in the model rather than a change in the test. Diagnostic-only: this scan never entered a decision gate.
Each panel below is on its own scale with zero at that cell’s own chance level. The eight axes are never averaged into a combined score, because they have different answer spaces and different chance levels.
No confidence intervals are drawn: the source does not provide them for these n=128 cells, and computing them here would introduce a statistic the evidence does not contain.
| Checkpoint | Step | Tokens | Validation CE | M@4 | R@1 | I@8 | C@True | Q@1 | V@8 | A@0 | T@modulus |
|---|---|---|---|---|---|---|---|---|---|---|---|
| θ₀ | 0 | 0 | 10.503032 | 35/128 +0.031 | 66/128 +0.031 | 4/128 +0.000 | 6/128 +0.016 | 2/128 −0.049 | 17/128 +0.009 | 4/128 −0.033 | 4/128 +0.016 |
| 1% | 325 | 5,324,800 | 5.210447 | 26/128 −0.062 | 65/128 +0.016 | 8/128 +0.032 | 0/128 −0.032 | 8/128 +0.001 | 13/128 −0.027 | 9/128 +0.008 | 1/128 −0.008 |
| 5% | 1,624 | 26,607,616 | 3.860014 | 34/128 +0.021 | 66/128 +0.031 | 5/128 +0.008 | 5/128 +0.008 | 14/128 +0.051 | 20/128 +0.036 | 11/128 +0.025 | 3/128 +0.008 |
| 10% | 3,248 | 53,215,232 | 3.311519 | 27/128 −0.052 | 65/128 +0.016 | 3/128 −0.008 | 4/128 +0.000 | 5/128 −0.024 | 21/128 +0.045 | 45/128 +0.308 | 1/128 −0.008 |
| 25% | 8,118 | 133,005,312 | 2.950929 | 36/128 +0.042 | 78/128 +0.219 | 22/128 +0.145 | 2/128 −0.016 | 1/128 −0.057 | 20/128 +0.036 | 36/128 +0.233 | 6/128 +0.032 |
| 50% | 16,236 | 266,010,624 | 2.760658 | 35/128 +0.031 | 67/128 +0.047 | 47/128 +0.347 | 7/128 +0.024 | 3/128 −0.041 | 20/128 +0.036 | 59/128 +0.425 | 11/128 +0.071 |
| 75% | 24,353 | 398,999,552 | 2.655592 | 32/128 +0.000 | 76/128 +0.188 | 57/128 +0.427 | 3/128 −0.008 | 2/128 −0.049 | 20/128 +0.036 | 48/128 +0.333 | 14/128 +0.095 |
| 100% | 32,470 | 531,988,480 | 2.622200 | 32/128 +0.000 | 74/128 +0.156 | 53/128 +0.395 | 7/128 +0.024 | 3/128 −0.041 | 20/128 +0.036 | 42/128 +0.283 | 19/128 +0.135 |
Development observations
- A became distinct early. The A@0 normalized score moved from −0.033 at θ₀ to +0.308 at 10%, peaked at +0.425 at 50%, and fell back to +0.283 at the final checkpoint.
- I rose later. I@8 stayed around chance up to 10%, then reached +0.145 at 25%, +0.347 at 50% and +0.427 at 75%.
- In this model the measured A signal appeared before the I signal. This single-run pattern does not support the hypothesis that I is a required precursor of the other axes, and it does not refute it either.
- T-modulus rose late, from around 0 to +0.135 at the final checkpoint. This does not mean the whole T axis is ordinal or measurable.
- Q showed no positive capacity. The selected Q cell stayed below chance at most checkpoints.
- V was largely fixed. V@8 repeating 20/128 across later checkpoints keeps measurement insensitivity on the table alongside model invariance.
- Validation CE and the axes did not move together. CE fell steadily while A peaked before the end, M stayed volatile, and I opened late.
Descriptive, single-run, and not a general law
These are exploratory longitudinal measurements over a matched item set. No formal change-point, mixed-effects or population model was applied.
In particular, “A appeared before I” is a statement about this run. It does not establish a developmental order that other models must follow, and it does not by itself refute one either.
Final n=800 decision chain
This chain is not a canonical STOP-v1 artifact
It ran through an architectural compatibility adapter in order to load the 70M architecture, and is recorded as canonical_stop_v1_artifact: false. It must not be visually or epistemically merged with the diagnostic Birth Map above, and it is not equivalent to the Model-0 official result.
| Axis | Main result | Forms A / B | Multiplicity | Gate status |
|---|---|---|---|---|
| M | No usable calibration level | No confirmation | Not tested | Not measurable |
| R | d*=3; A=252/800, B=248/800; chance=0.25 | Forms consistent | Not significant after Holm | Not measurable |
| I | d*=8; A=347/800, B=332/800; chance=0.03125 | Very strong replicated signal | adjusted p = 1.0817×10⁻¹⁴⁴ | Difficulty sensitivity not met |
| C | No usable calibration level | No confirmation | Not tested | Not measurable |
| Q | No usable calibration level | No confirmation | Not tested | Not measurable |
| V | No usable calibration level | No confirmation | Not tested | Not measurable |
| A | d*=0.0; A=229/800, B=253/800; chance=0.0625 | Strong replicated signal | adjusted p = 4.4794×10⁻⁴⁰ | Measurable |
| T | Non-ordinal category set | No confirmation | Not tested | Not measurable under the current contract |
- Verdict
- RESCOPE
- Measurable axes
- A only
- Count
- 1
- Requirement
- At least 3 measurable axes and at least 2 higher-axis
- Canonical artifact
- No
The general birth claim was not supported; the run was held at the boundary of a low-level learning phenotype.
Source: Scientific Evidence Inventory v1.0 (evidence cutoff 30 August 2026), §6.1, §6.2, §6.3, §6.4, §3.1.
What happened next
This model was later trained further, to 1.00 billion cumulative tokens. That continuation is a separate page — and it is the same model, not a new one.
- Pilot-0.3 — continued training to 1B tokens
Continued training of this model. Not a third independent birth.