Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 4 additions & 4 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -90,10 +90,10 @@ If a relevant check cannot be run, state why and what remains unverified.

## Current Technical Status

Milestone 022 is the latest modeling evidence. Patience-3 validation early stopping reproduces
ByteBPE512's best full-validation BPC `2.0083`, stops at step 2,500, and roughly halves runtime.
Weight decay `0.01` reaches `2.0080`, too small a difference to claim improvement. Future work should
test the tokenizer result across seeds or corpus splits rather than add another narrow hyperparameter.
Milestone 023 is the latest modeling evidence. Across preregistered seeds 1337, 2027, and 4242,
ByteBPE512 early stopping reaches best full-validation BPC `2.0225 ± 0.0124` (population SD), range
`2.0083–2.0384`; all seeds beat the character control. Future work should test split or corpus
robustness rather than select a favorable seed or add another narrow hyperparameter.

## Safety And Security

Expand Down
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,7 @@ chronological 90/10 split unless noted otherwise.
| [020](experiments/020-bpe-context-and-learning-rate.md) | BPE context/LR controls | Best BPE128 bits/character | `2.0976` remains best | Matched character context and `5e-4` LR changed timing/diversity but did not improve held-out BPC. |
| [021](experiments/021-boundary-aware-byte-bpe.md) | Boundary-aware ByteBPE320/512 | Full-validation best bits/character | `2.0286` / `2.0083` | Both beat the corrected character and BPE128 controls; ByteBPE512 overfit after step 1750. |
| [022](experiments/022-early-stopping-and-regularization.md) | ByteBPE512 early stopping / weight decay | Actual steps and best BPC | `2500`, `2.0083` / `2.0080` | Early stopping halves runtime and limits final degradation; weight decay `0.01` is effectively neutral. |
| [023](experiments/023-multi-seed-robustness.md) | ByteBPE512 across three seeds | Best BPC mean ± population SD | `2.0225 ± 0.0124` | All tested seeds beat the character control; stopping varies from 2500–3000 steps. |

Token-level loss and perplexity are not directly comparable between character
and BPE tokenizers because they predict different units. The tokenizer
Expand Down
38 changes: 38 additions & 0 deletions configs/gptiny_bytebpe512_5k_lr1e-3_ctx37_earlystop_seed2027.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
data:
input_path: data/raw/input.txt
prepared_path: data/processed/corpus.txt
manifest_path: data/processed/corpus_manifest.json
tokenizer_path: data/processed/tokenizer_bytebpe512.json
tokenizer_type: byte_bpe
bpe_vocab_size: 512
bpe_min_frequency: 2
block_size: 37
train_split: 0.9

model:
vocab_size: 512
block_size: 37
n_layer: 4
n_head: 4
n_embd: 128
dropout: 0.1

train:
run_name: gptiny_bytebpe512_5k_lr1e-3_ctx37_earlystop_seed2027
runs_dir: runs
batch_size: 27
max_steps: 5000
learning_rate: 0.001
weight_decay: 0.0
log_interval: 100
eval_interval: 250
eval_batches: null
early_stopping_patience: 3
early_stopping_min_delta: 0.0
sample_prompt: Once
sample_max_new_tokens: 100
sample_temperature: 1.0
sample_top_k: null
sample_seed: 1337
sample_greedy: false
seed: 2027
38 changes: 38 additions & 0 deletions configs/gptiny_bytebpe512_5k_lr1e-3_ctx37_earlystop_seed4242.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
data:
input_path: data/raw/input.txt
prepared_path: data/processed/corpus.txt
manifest_path: data/processed/corpus_manifest.json
tokenizer_path: data/processed/tokenizer_bytebpe512.json
tokenizer_type: byte_bpe
bpe_vocab_size: 512
bpe_min_frequency: 2
block_size: 37
train_split: 0.9

model:
vocab_size: 512
block_size: 37
n_layer: 4
n_head: 4
n_embd: 128
dropout: 0.1

train:
run_name: gptiny_bytebpe512_5k_lr1e-3_ctx37_earlystop_seed4242
runs_dir: runs
batch_size: 27
max_steps: 5000
learning_rate: 0.001
weight_decay: 0.0
log_interval: 100
eval_interval: 250
eval_batches: null
early_stopping_patience: 3
early_stopping_min_delta: 0.0
sample_prompt: Once
sample_max_new_tokens: 100
sample_temperature: 1.0
sample_top_k: null
sample_seed: 1337
sample_greedy: false
seed: 4242
2 changes: 1 addition & 1 deletion docs/codex/build-and-test.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,7 +50,7 @@ The individual commands remain canonical and are listed below.
python -m pytest
```

Current expected result after milestone 022+: at least 121 tests passing with at least 90% coverage.
Current expected result after milestone 023+: at least 141 tests passing with at least 90% coverage.

### Compile Check

Expand Down
4 changes: 4 additions & 0 deletions docs/codex/experiments.md
Original file line number Diff line number Diff line change
Expand Up @@ -157,3 +157,7 @@ controlled question is early stopping or modest regularization rather than more
Milestone 022 adds patience-3 validation early stopping. It reproduces ByteBPE512's step-1,750 best
checkpoint and stops at step 2,500, roughly halving runtime. Weight decay `0.01` is effectively
neutral at best BPC `2.0080`; future work should test robustness across seeds or corpus splits.

Milestone 023 runs the unregularized ByteBPE512 early-stopping setup at seeds 1337, 2027, and 4242.
Best BPC is `2.0225 ± 0.0124` (population SD), range `2.0083–2.0384`; every seed beats the
character control. The next modeling question should test split/corpus robustness, not select a seed.
6 changes: 6 additions & 0 deletions docs/experiments.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,7 @@ not a replacement for the original reports.
| [020 BPE Context and Learning Rate](../experiments/020-bpe-context-and-learning-rate.md) | Tokenizer diagnostics | Matching character context and lowering BPE learning rate did not beat the corrected BPE128 control or character model. |
| [021 Boundary-Aware Byte BPE](../experiments/021-boundary-aware-byte-bpe.md) | Tokenizer design | Lossless boundary-aware ByteBPE320/512 beat both corrected controls on best BPC; ByteBPE512 reached `2.0083` but overfit early. |
| [022 Early Stopping and Regularization](../experiments/022-early-stopping-and-regularization.md) | Training control | Patience-3 stopping reproduced the step-1750 optimum and halved runtime; weight decay `0.01` was effectively neutral. |
| [023 Multi-Seed Robustness](../experiments/023-multi-seed-robustness.md) | Robustness | Three preregistered seeds average best BPC `2.0225 ± 0.0124`; every seed beats the corrected character control. |

## Topic Shortcuts

Expand Down Expand Up @@ -79,3 +80,8 @@ The 512-token run's final BPC rose to `2.2450`, making early best-checkpoint sel
Milestone 022 operationalizes that result. Patience-3 early stopping terminates at step 2,500 and
retains best BPC `2.0083`, while stopped-final BPC improves from `2.2450` to `2.0554`. Weight decay
`0.01` reaches best BPC `2.0080`; the `0.00025` difference is too small to interpret as a real gain.

Milestone 023 measures seed sensitivity directly. Across seeds 1337, 2027, and 4242, best BPC is
`2.0225 ± 0.0124` with range `2.0083–2.0384`; best step ranges 1,750–2,250 and stop step
2,500–3,000. The tokenizer result survives all tested seeds, while the observed seed spread confirms
that milestone 022's tiny weight-decay delta was not decision-grade evidence.
12 changes: 12 additions & 0 deletions docs/training.md
Original file line number Diff line number Diff line change
Expand Up @@ -117,6 +117,15 @@ the saved sample.
inspectable. The final checkpoint is the model at the stop step, while `best_checkpoint.pt` remains
the lowest observed validation-loss model.

Aggregate completed runs without selecting a winner:

```bash
python scripts/summarize_runs.py runs/<name-a>/<id> runs/<name-b>/<id> runs/<name-c>/<id>
```

The command reads each run's config and summary, requires distinct training seeds, and prints every
observation plus mean, population standard deviation, minimum, and maximum.

## Generation Controls

```bash
Expand Down Expand Up @@ -167,6 +176,9 @@ control, though ByteBPE512 overfit sharply after step 1,750.
Experiment 022 adds patience-3 early stopping: it ends at step 2,500, halves runtime, and avoids most
final-checkpoint degradation. Weight decay `0.01` changes best BPC by only `0.00025`, which is not
meaningful evidence of improvement.
Experiment 023 repeats the unregularized early-stopping run across seeds 1337, 2027, and 4242.
Best BPC is `2.0225 ± 0.0124` (population SD), and all three runs remain better than the corrected
character control. Stop steps vary from 2,500 to 3,000.

## Artifact Policy

Expand Down
91 changes: 91 additions & 0 deletions experiments/023-multi-seed-robustness.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,91 @@
# 023 — Multi-Seed Robustness

## Goal

Measure whether ByteBPE512's advantage and early-stopping behavior survive training randomness.
Seeds 1337, 2027, and 4242 were fixed before the additional runs; all completed seeds are reported,
and no result is selected or discarded based on quality.

## Setup

Every run uses the milestone-022 unregularized configuration: the same training-only ByteBPE512
tokenizer, corpus checksum `a4c81ef23eb9…`, chronological 90/10 split, 37-token context, batch 27,
4-layer/4-head/128-wide GPTiny, dropout 0.1, AdamW `lr=1e-3`, zero weight decay, full validation
every 250 steps, and patience-3 early stopping under a 5,000-step ceiling. Training seed is the only
changed field. Generation sampling holds seed 1337 fixed to isolate model-training variation.

Population standard deviation describes this complete preregistered seed set; with only three seeds,
it is descriptive uncertainty, not a confidence interval or population estimate.
The aggregation CLI verified canonical experiment fingerprint
`ed9650e463830d0d85dbbc32a40164b6699d7033017386a2cb22f2ab035431b0` across all runs.

## Validation Results

| seed | actual steps | best step | best loss | best BPC | stopped-final BPC | duration |
| ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| 1337 | 2,500 | 1,750 | 2.375155 | 2.008269 | 2.055447 | 209.3s |
| 2027 | 3,000 | 2,250 | 2.410806 | 2.038414 | 2.099879 | 308.3s |
| 4242 | 2,500 | 1,750 | 2.390150 | 2.020948 | 2.052917 | 244.1s |

| metric | mean | population SD | minimum | maximum |
| --- | ---: | ---: | ---: | ---: |
| actual steps | 2,666.67 | 235.70 | 2,500 | 3,000 |
| best step | 1,916.67 | 235.70 | 1,750 | 2,250 |
| best BPC | **2.022544** | **0.012358** | 2.008269 | 2.038414 |
| stopped-final BPC | 2.069414 | 0.021567 | 2.052917 | 2.099879 |
| duration seconds | 253.93 | 41.01 | 209.32 | 308.33 |

All three seeds beat the corrected character control (`2.075981`) and BPE128 control (`2.097552`)
on best BPC. The worst tested ByteBPE512 seed retains a 0.03757 BPC advantage over character. The
tokenizer conclusion therefore survives this seed set, although its effect is smaller than the
single best seed suggested.

Seed-to-seed best-BPC SD (`0.012358`) is about 49 times milestone 022's apparent weight-decay gain
(`0.000251`). That comparison strengthens the earlier conclusion that weight decay `0.01` was
effectively neutral. Early stopping is stable in direction but not timing: seed 2027 improves later
and needs 500 additional steps.

## Controlled Generation

Prompt `Once`, 100 new tokens. Seeded decoding uses temperature 0.8, top-k 10, seed 1337.

| training seed | best greedy d-2 | best seeded d-2 | final greedy d-2 | final seeded d-2 |
| ---: | ---: | ---: | ---: | ---: |
| 1337 | 0.5217 | 0.5417 | 0.5723 | 0.5980 |
| 2027 | 0.5562 | 0.6836 | 0.5595 | 0.6591 |
| 4242 | 0.3657 | 0.6548 | 0.5056 | 0.5670 |
| mean | 0.4812 | 0.6267 | 0.5458 | 0.6080 |
| population SD | 0.0829 | 0.0612 | 0.0289 | 0.0383 |

Generation varies materially across training seeds. Seeded distinct-2 is usually higher than greedy,
but neither it nor final-checkpoint diversity tracks best validation BPC. Samples remain locally
plausible and globally incoherent, so the robustness claim is limited to held-out likelihood.

## Exact Commands

```bash
uv run --frozen --extra dev python scripts/prepare_data.py --config configs/gptiny_bytebpe512_5k_lr1e-3_ctx37_earlystop_seed2027.yaml
uv run --frozen --extra dev python scripts/train.py --config configs/gptiny_bytebpe512_5k_lr1e-3_ctx37_earlystop_seed2027.yaml
uv run --frozen --extra dev python scripts/train.py --config configs/gptiny_bytebpe512_5k_lr1e-3_ctx37_earlystop_seed4242.yaml
uv run --frozen --extra dev python scripts/summarize_runs.py \
runs/gptiny_bytebpe512_5k_lr1e-3_ctx37_earlystop/2026-07-12_22-12-08 \
runs/gptiny_bytebpe512_5k_lr1e-3_ctx37_earlystop_seed2027/2026-07-12_22-37-30 \
runs/gptiny_bytebpe512_5k_lr1e-3_ctx37_earlystop_seed4242/2026-07-12_22-43-16
uv run --frozen --extra dev make check
```

Generation used `scripts/generate.py` for seeds 2027 and 4242 at `best` and `final`, once with
`--greedy --diagnostics` and once with
`--temperature 0.8 --top-k 10 --seed 1337 --diagnostics`. Seed-1337 diagnostics are the controlled
milestone-022 values.

## Limitations And Next Step

- Three seeds reveal variation but do not estimate a broad seed distribution precisely.
- Every seed shares one chronological validation split and one small English corpus.
- Repeated validation drives early stopping and model selection.
- Distinct-n is not semantic or human evaluation.

The next strong test changes the data axis: evaluate the fixed ByteBPE512 early-stopping protocol on
multiple deterministic corpus splits or an additional public-domain corpus. That would test whether
the tokenizer advantage is distribution-robust rather than merely seed-robust.
21 changes: 21 additions & 0 deletions notes/06-reproducibility.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,3 +60,24 @@ Implementation: [`artifacts.py`](../src/smallm/training/artifacts.py),

Checks: reconstruct a result from its summary; identify which failures occur before any run
directory exists; explain why old and corrected BPC values cannot be compared as one series.
### Seed ensembles and descriptive uncertainty

A random seed fixes initialization, minibatch order, dropout masks, and sampling streams; it is an
experimental condition, not a hyperparameter to optimize. For preregistered seeds
(s_1,\ldots,s_n) and metric (x_i), report every observation plus

\[
\bar{x}=\frac1n\sum_{i=1}^n x_i,
\qquad
\sigma_{\mathrm{pop}}=\sqrt{\frac1n\sum_{i=1}^n(x_i-\bar{x})^2}.
\]

smaLLM uses population standard deviation because the report describes the complete, explicitly
chosen seed set; it does not pretend three seeds estimate a universal sampling distribution. The
minimum and maximum expose asymmetry that a mean and deviation can hide. Stop step is itself a
random outcome under early stopping and must be summarized alongside model quality.

Never discard a completed seed because it weakens the conclusion, and never report the best seed as
the expected result. A hyperparameter difference much smaller than seed-to-seed spread is not robust
evidence. Decoding randomness is held fixed when comparing training seeds so observed generation
variation comes from model training rather than a second uncontrolled random stream.
Loading
Loading