Pilot-0.2 / Local-70M-0

69,970,848 parameters. Independent model birth #2. 531,988,480 tokens over 32,470 steps, final validation CE 2.6221997812390327.

Eight checkpoints were measured on the same fixed cells, from the untrained starting state to the end of training.

Training result

Scientific role
Independent model birth #2
Parameters
69,970,848 unique trainable parameters — not 88.9M, see below
Tokens
531,988,480
Training steps
32,470
Precision
fp16
Final validation CE
2.6221997812390327
Final checkpoint size
915,598,984 bytes
Final checkpoint SHA-256
cd53d21bf67f1ab23951152542bc13119cee8532bb5fe4fb9ce93178bdb9f309

Completed through verified checkpoint recovery

Training completed through a deterministic recovery branch that resumed from verified checkpoints after interruptions.

This run is not presented as a single uninterrupted process.

This is a 69,970,848-parameter model, not an 88.9M model

Some Observatory outputs report roughly 88.9 million state elements. That figure counts the tied embedding weight twice, because it is materialised in the state dictionary under both tok_emb.weight and lm_head.weight. The unique trainable parameter count is 69,970,848. The model must not be presented as an 88.9M model.

The Birth Map

Evidence authority: Diagnostic-only

In plain language

We measured the same eight fixed cells at eight points during training, from the untrained starting state to the end.

The measured signals did not rise together. One measure (A) appeared early in this run and then fell back before the end. Another (I) stayed flat for a long time and rose late. Some measures peaked before the final checkpoint. Meanwhile the model’s overall language-modelling loss fell steadily the whole time.

So in this run, the general training metric and the measured profile were not the same thing.

Technical evidence

Matched items, n=128 per cell, held fixed across all eight checkpoints so that a change in score reflects a change in the model rather than a change in the test. Diagnostic-only: this scan never entered a decision gate.

Each panel below is on its own scale with zero at that cell’s own chance level. The eight axes are never averaged into a combined score, because they have different answer spaces and different chance levels.

No confidence intervals are drawn: the source does not provide them for these n=128 cells, and computing them here would introduce a statistic the evidence does not contain.

M@4 across eight Pilot-0.2 checkpoints Normalized score for cell M@4 at eight checkpoints, from theta-zero to one hundred percent, over n=128 matched items. The zero line is chance. Filled dots are at or above chance; hollow dots are below chance. No confidence interval is drawn because the source does not provide one for these cells.0 = chance-0.05M@4 at θ₀ — 35/128, normalized +0.031M@4 at 1% — 26/128, normalized −0.062M@4 at 5% — 34/128, normalized +0.021M@4 at 10% — 27/128, normalized −0.052M@4 at 25% — 36/128, normalized +0.042M@4 at 50% — 35/128, normalized +0.031M@4 at 75% — 32/128, normalized +0.000M@4 at 100% — 32/128, normalized +0.000θ₀1%5%10%25%50%75%100%M@4
R@1 across eight Pilot-0.2 checkpoints Normalized score for cell R@1 at eight checkpoints, from theta-zero to one hundred percent, over n=128 matched items. The zero line is chance. Filled dots are at or above chance; hollow dots are below chance. No confidence interval is drawn because the source does not provide one for these cells.0 = chance0.20-0.05R@1 at θ₀ — 66/128, normalized +0.031R@1 at 1% — 65/128, normalized +0.016R@1 at 5% — 66/128, normalized +0.031R@1 at 10% — 65/128, normalized +0.016R@1 at 25% — 78/128, normalized +0.219R@1 at 50% — 67/128, normalized +0.047R@1 at 75% — 76/128, normalized +0.188R@1 at 100% — 74/128, normalized +0.156θ₀1%5%10%25%50%75%100%R@1
I@8 across eight Pilot-0.2 checkpoints Normalized score for cell I@8 at eight checkpoints, from theta-zero to one hundred percent, over n=128 matched items. The zero line is chance. Filled dots are at or above chance; hollow dots are below chance. No confidence interval is drawn because the source does not provide one for these cells.0 = chance0.200.40-0.05I@8 at θ₀ — 4/128, normalized +0.000I@8 at 1% — 8/128, normalized +0.032I@8 at 5% — 5/128, normalized +0.008I@8 at 10% — 3/128, normalized −0.008I@8 at 25% — 22/128, normalized +0.145I@8 at 50% — 47/128, normalized +0.347I@8 at 75% — 57/128, normalized +0.427I@8 at 100% — 53/128, normalized +0.395θ₀1%5%10%25%50%75%100%I@8
C@True across eight Pilot-0.2 checkpoints Normalized score for cell C@True at eight checkpoints, from theta-zero to one hundred percent, over n=128 matched items. The zero line is chance. Filled dots are at or above chance; hollow dots are below chance. No confidence interval is drawn because the source does not provide one for these cells.0 = chance-0.05C@True at θ₀ — 6/128, normalized +0.016C@True at 1% — 0/128, normalized −0.032C@True at 5% — 5/128, normalized +0.008C@True at 10% — 4/128, normalized +0.000C@True at 25% — 2/128, normalized −0.016C@True at 50% — 7/128, normalized +0.024C@True at 75% — 3/128, normalized −0.008C@True at 100% — 7/128, normalized +0.024θ₀1%5%10%25%50%75%100%C@True
Q@1 across eight Pilot-0.2 checkpoints Normalized score for cell Q@1 at eight checkpoints, from theta-zero to one hundred percent, over n=128 matched items. The zero line is chance. Filled dots are at or above chance; hollow dots are below chance. No confidence interval is drawn because the source does not provide one for these cells.0 = chance-0.05Q@1 at θ₀ — 2/128, normalized −0.049Q@1 at 1% — 8/128, normalized +0.001Q@1 at 5% — 14/128, normalized +0.051Q@1 at 10% — 5/128, normalized −0.024Q@1 at 25% — 1/128, normalized −0.057Q@1 at 50% — 3/128, normalized −0.041Q@1 at 75% — 2/128, normalized −0.049Q@1 at 100% — 3/128, normalized −0.041θ₀1%5%10%25%50%75%100%Q@1
V@8 across eight Pilot-0.2 checkpoints Normalized score for cell V@8 at eight checkpoints, from theta-zero to one hundred percent, over n=128 matched items. The zero line is chance. Filled dots are at or above chance; hollow dots are below chance. No confidence interval is drawn because the source does not provide one for these cells.0 = chance-0.05V@8 at θ₀ — 17/128, normalized +0.009V@8 at 1% — 13/128, normalized −0.027V@8 at 5% — 20/128, normalized +0.036V@8 at 10% — 21/128, normalized +0.045V@8 at 25% — 20/128, normalized +0.036V@8 at 50% — 20/128, normalized +0.036V@8 at 75% — 20/128, normalized +0.036V@8 at 100% — 20/128, normalized +0.036θ₀1%5%10%25%50%75%100%V@8
A@0 across eight Pilot-0.2 checkpoints Normalized score for cell A@0 at eight checkpoints, from theta-zero to one hundred percent, over n=128 matched items. The zero line is chance. Filled dots are at or above chance; hollow dots are below chance. No confidence interval is drawn because the source does not provide one for these cells.0 = chance0.200.40-0.05A@0 at θ₀ — 4/128, normalized −0.033A@0 at 1% — 9/128, normalized +0.008A@0 at 5% — 11/128, normalized +0.025A@0 at 10% — 45/128, normalized +0.308A@0 at 25% — 36/128, normalized +0.233A@0 at 50% — 59/128, normalized +0.425A@0 at 75% — 48/128, normalized +0.333A@0 at 100% — 42/128, normalized +0.283θ₀1%5%10%25%50%75%100%A@0
T@modulus across eight Pilot-0.2 checkpoints Normalized score for cell T@modulus at eight checkpoints, from theta-zero to one hundred percent, over n=128 matched items. The zero line is chance. Filled dots are at or above chance; hollow dots are below chance. No confidence interval is drawn because the source does not provide one for these cells.0 = chance-0.05T@modulus at θ₀ — 4/128, normalized +0.016T@modulus at 1% — 1/128, normalized −0.008T@modulus at 5% — 3/128, normalized +0.008T@modulus at 10% — 1/128, normalized −0.008T@modulus at 25% — 6/128, normalized +0.032T@modulus at 50% — 11/128, normalized +0.071T@modulus at 75% — 14/128, normalized +0.095T@modulus at 100% — 19/128, normalized +0.135θ₀1%5%10%25%50%75%100%T@modulus
Eight measured cells across eight Pilot-0.2 checkpoints, one panel per cell. Each panel is on its own scale; the axes are never averaged together. Zero is that cell’s own chance level. Filled dots sit at or above chance, hollow dots below it. Source §6.2, n=128 matched items per cell.
Pilot-0.2 Birth Map — source §6.2, diagnostic-only, n=128 matched items per cell. Values are hits over n with the normalized score below.
CheckpointStepTokensValidation CEM@4R@1I@8C@TrueQ@1V@8A@0T@modulus
θ₀0010.50303235/128
+0.031
66/128
+0.031
4/128
+0.000
6/128
+0.016
2/128
−0.049
17/128
+0.009
4/128
−0.033
4/128
+0.016
1%3255,324,8005.21044726/128
−0.062
65/128
+0.016
8/128
+0.032
0/128
−0.032
8/128
+0.001
13/128
−0.027
9/128
+0.008
1/128
−0.008
5%1,62426,607,6163.86001434/128
+0.021
66/128
+0.031
5/128
+0.008
5/128
+0.008
14/128
+0.051
20/128
+0.036
11/128
+0.025
3/128
+0.008
10%3,24853,215,2323.31151927/128
−0.052
65/128
+0.016
3/128
−0.008
4/128
+0.000
5/128
−0.024
21/128
+0.045
45/128
+0.308
1/128
−0.008
25%8,118133,005,3122.95092936/128
+0.042
78/128
+0.219
22/128
+0.145
2/128
−0.016
1/128
−0.057
20/128
+0.036
36/128
+0.233
6/128
+0.032
50%16,236266,010,6242.76065835/128
+0.031
67/128
+0.047
47/128
+0.347
7/128
+0.024
3/128
−0.041
20/128
+0.036
59/128
+0.425
11/128
+0.071
75%24,353398,999,5522.65559232/128
+0.000
76/128
+0.188
57/128
+0.427
3/128
−0.008
2/128
−0.049
20/128
+0.036
48/128
+0.333
14/128
+0.095
100%32,470531,988,4802.62220032/128
+0.000
74/128
+0.156
53/128
+0.395
7/128
+0.024
3/128
−0.041
20/128
+0.036
42/128
+0.283
19/128
+0.135

Development observations

Evidence authority: Diagnostic-only

  1. A became distinct early. The A@0 normalized score moved from −0.033 at θ₀ to +0.308 at 10%, peaked at +0.425 at 50%, and fell back to +0.283 at the final checkpoint.
  2. I rose later. I@8 stayed around chance up to 10%, then reached +0.145 at 25%, +0.347 at 50% and +0.427 at 75%.
  3. In this model the measured A signal appeared before the I signal. This single-run pattern does not support the hypothesis that I is a required precursor of the other axes, and it does not refute it either.
  4. T-modulus rose late, from around 0 to +0.135 at the final checkpoint. This does not mean the whole T axis is ordinal or measurable.
  5. Q showed no positive capacity. The selected Q cell stayed below chance at most checkpoints.
  6. V was largely fixed. V@8 repeating 20/128 across later checkpoints keeps measurement insensitivity on the table alongside model invariance.
  7. Validation CE and the axes did not move together. CE fell steadily while A peaked before the end, M stayed volatile, and I opened late.

Descriptive, single-run, and not a general law

These are exploratory longitudinal measurements over a matched item set. No formal change-point, mixed-effects or population model was applied.

In particular, “A appeared before I” is a statement about this run. It does not establish a developmental order that other models must follow, and it does not by itself refute one either.

Final n=800 decision chain

Evidence authority: Compatibility-adapter

This chain is not a canonical STOP-v1 artifact

It ran through an architectural compatibility adapter in order to load the 70M architecture, and is recorded as canonical_stop_v1_artifact: false. It must not be visually or epistemically merged with the diagnostic Birth Map above, and it is not equivalent to the Model-0 official result.

Pilot-0.2 final n=800 decision chain — source §6.4, compatibility-adapter authority.
AxisMain resultForms A / BMultiplicityGate status
MNo usable calibration levelNo confirmationNot testedNot measurable
Rd*=3; A=252/800, B=248/800; chance=0.25Forms consistentNot significant after HolmNot measurable
Id*=8; A=347/800, B=332/800; chance=0.03125Very strong replicated signaladjusted p = 1.0817×10⁻¹⁴⁴Difficulty sensitivity not met
CNo usable calibration levelNo confirmationNot testedNot measurable
QNo usable calibration levelNo confirmationNot testedNot measurable
VNo usable calibration levelNo confirmationNot testedNot measurable
Ad*=0.0; A=229/800, B=253/800; chance=0.0625Strong replicated signaladjusted p = 4.4794×10⁻⁴⁰Measurable
TNon-ordinal category setNo confirmationNot testedNot measurable under the current contract
Verdict
RESCOPE
Measurable axes
A only
Count
1
Requirement
At least 3 measurable axes and at least 2 higher-axis
Canonical artifact
No

The general birth claim was not supported; the run was held at the boundary of a low-level learning phenotype.

Source: Scientific Evidence Inventory v1.0 (evidence cutoff 30 August 2026), §6.1, §6.2, §6.3, §6.4, §3.1.