Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 5 additions & 5 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -90,11 +90,11 @@ If a relevant check cannot be run, state why and what remains unverified.

## Current Technical Status

Milestone 024 is the latest modeling evidence. On a deterministically extracted, near-size-matched
Peter Pan corpus, ByteBPE512 reaches best BPC `2.1539` versus `2.1721` for character. This
replicates the direction found on Alice, but the 0.83% margin from one seed is not a stable
effect-size estimate. Milestone 023 remains the seed-robustness reference (`2.0225 ± 0.0124` on
Alice); the next strong test is a corpus-by-seed matrix.
Milestone 025 is the latest modeling evidence. In a balanced 2-tokenizer × 2-corpus × 3-seed
matrix, ByteBPE512 beats character in all six paired comparisons. Its mean advantage is `0.0619`
BPC on Alice and `0.0252` on Peter Pan; the `+0.0367` corpus interaction means effect magnitude is
not distribution-invariant. The next contract change should add a sealed chronological test segment
before further model selection.

## Safety And Security

Expand Down
5 changes: 3 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,8 +17,8 @@ For a fast technical review, inspect:
1. [`docs/architecture.md`](docs/architecture.md) for boundaries and the
end-to-end pipeline.
2. [`docs/experiments.md`](docs/experiments.md) for the milestone index.
3. [`experiments/024-cross-corpus-robustness.md`](experiments/024-cross-corpus-robustness.md)
for the latest cross-corpus modeling evidence and its limitations.
3. [`experiments/025-corpus-by-seed-matrix.md`](experiments/025-corpus-by-seed-matrix.md)
for the latest balanced robustness evidence and its limitations.
4. [`src/smallm/model/`](src/smallm/model/) for the from-scratch GPTiny model.
5. [`src/smallm/training/`](src/smallm/training/) for training and run artifacts.
6. [`src/smallm/data/`](src/smallm/data/) for corpus and tokenizer contracts.
Expand Down Expand Up @@ -49,6 +49,7 @@ chronological 90/10 split unless noted otherwise.
| [022](experiments/022-early-stopping-and-regularization.md) | ByteBPE512 early stopping / weight decay | Actual steps and best BPC | `2500`, `2.0083` / `2.0080` | Early stopping halves runtime and limits final degradation; weight decay `0.01` is effectively neutral. |
| [023](experiments/023-multi-seed-robustness.md) | ByteBPE512 across three seeds | Best BPC mean ± population SD | `2.0225 ± 0.0124` | All tested seeds beat the character control; stopping varies from 2500–3000 steps. |
| [024](experiments/024-cross-corpus-robustness.md) | Peter Pan character vs ByteBPE512 | Full-validation best bits/character | `2.1721` vs `2.1539` | ByteBPE512 replicates the direction on a second book, but only by 0.83%. |
| [025](experiments/025-corpus-by-seed-matrix.md) | 2 tokenizers × 2 corpora × 3 seeds | Paired ByteBPE512 minus character BPC | `-0.0619` Alice; `-0.0252` Peter Pan | ByteBPE512 wins all six pairs; effect magnitude is corpus-dependent. |

Token-level loss and perplexity are not directly comparable between character
and BPE tokenizers because they predict different units. The tokenizer
Expand Down
36 changes: 36 additions & 0 deletions configs/gptiny_char_5k_lr1e-3_earlystop.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
data:
input_path: data/raw/input.txt
prepared_path: data/processed/corpus.txt
manifest_path: data/processed/corpus_manifest.json
tokenizer_path: data/processed/tokenizer.json
tokenizer_type: char
block_size: 64
train_split: 0.9

model:
vocab_size: 256
block_size: 64
n_layer: 4
n_head: 4
n_embd: 128
dropout: 0.1

train:
run_name: gptiny_char_5k_lr1e-3_earlystop
runs_dir: runs
batch_size: 16
max_steps: 5000
learning_rate: 0.001
weight_decay: 0.0
log_interval: 100
eval_interval: 250
eval_batches: null
early_stopping_patience: 3
early_stopping_min_delta: 0.0
sample_prompt: Once
sample_max_new_tokens: 100
sample_temperature: 1.0
sample_top_k: null
sample_seed: 1337
sample_greedy: false
seed: 1337
36 changes: 36 additions & 0 deletions configs/gptiny_char_5k_lr1e-3_earlystop_seed2027.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
data:
input_path: data/raw/input.txt
prepared_path: data/processed/corpus.txt
manifest_path: data/processed/corpus_manifest.json
tokenizer_path: data/processed/tokenizer.json
tokenizer_type: char
block_size: 64
train_split: 0.9

model:
vocab_size: 256
block_size: 64
n_layer: 4
n_head: 4
n_embd: 128
dropout: 0.1

train:
run_name: gptiny_char_5k_lr1e-3_earlystop_seed2027
runs_dir: runs
batch_size: 16
max_steps: 5000
learning_rate: 0.001
weight_decay: 0.0
log_interval: 100
eval_interval: 250
eval_batches: null
early_stopping_patience: 3
early_stopping_min_delta: 0.0
sample_prompt: Once
sample_max_new_tokens: 100
sample_temperature: 1.0
sample_top_k: null
sample_seed: 1337
sample_greedy: false
seed: 2027
36 changes: 36 additions & 0 deletions configs/gptiny_char_5k_lr1e-3_earlystop_seed4242.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
data:
input_path: data/raw/input.txt
prepared_path: data/processed/corpus.txt
manifest_path: data/processed/corpus_manifest.json
tokenizer_path: data/processed/tokenizer.json
tokenizer_type: char
block_size: 64
train_split: 0.9

model:
vocab_size: 256
block_size: 64
n_layer: 4
n_head: 4
n_embd: 128
dropout: 0.1

train:
run_name: gptiny_char_5k_lr1e-3_earlystop_seed4242
runs_dir: runs
batch_size: 16
max_steps: 5000
learning_rate: 0.001
weight_decay: 0.0
log_interval: 100
eval_interval: 250
eval_batches: null
early_stopping_patience: 3
early_stopping_min_delta: 0.0
sample_prompt: Once
sample_max_new_tokens: 100
sample_temperature: 1.0
sample_top_k: null
sample_seed: 1337
sample_greedy: false
seed: 4242
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
data:
input_path: data/raw/peter_pan_body.txt
prepared_path: data/processed/peter_pan_corpus.txt
manifest_path: data/processed/peter_pan_corpus_manifest.json
tokenizer_path: data/processed/peter_pan_tokenizer_bytebpe512.json
tokenizer_type: byte_bpe
bpe_vocab_size: 512
bpe_min_frequency: 2
block_size: 37
train_split: 0.9

model:
vocab_size: 512
block_size: 37
n_layer: 4
n_head: 4
n_embd: 128
dropout: 0.1

train:
run_name: gptiny_peterpan_bytebpe512_5k_lr1e-3_ctx37_earlystop_seed2027
runs_dir: runs
batch_size: 27
max_steps: 5000
learning_rate: 0.001
weight_decay: 0.0
log_interval: 100
eval_interval: 250
eval_batches: null
early_stopping_patience: 3
early_stopping_min_delta: 0.0
sample_prompt: Once
sample_max_new_tokens: 100
sample_temperature: 1.0
sample_top_k: null
sample_seed: 1337
sample_greedy: false
seed: 2027
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
data:
input_path: data/raw/peter_pan_body.txt
prepared_path: data/processed/peter_pan_corpus.txt
manifest_path: data/processed/peter_pan_corpus_manifest.json
tokenizer_path: data/processed/peter_pan_tokenizer_bytebpe512.json
tokenizer_type: byte_bpe
bpe_vocab_size: 512
bpe_min_frequency: 2
block_size: 37
train_split: 0.9

model:
vocab_size: 512
block_size: 37
n_layer: 4
n_head: 4
n_embd: 128
dropout: 0.1

train:
run_name: gptiny_peterpan_bytebpe512_5k_lr1e-3_ctx37_earlystop_seed4242
runs_dir: runs
batch_size: 27
max_steps: 5000
learning_rate: 0.001
weight_decay: 0.0
log_interval: 100
eval_interval: 250
eval_batches: null
early_stopping_patience: 3
early_stopping_min_delta: 0.0
sample_prompt: Once
sample_max_new_tokens: 100
sample_temperature: 1.0
sample_top_k: null
sample_seed: 1337
sample_greedy: false
seed: 4242
36 changes: 36 additions & 0 deletions configs/gptiny_peterpan_char_5k_lr1e-3_earlystop_seed2027.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
data:
input_path: data/raw/peter_pan_body.txt
prepared_path: data/processed/peter_pan_corpus.txt
manifest_path: data/processed/peter_pan_corpus_manifest.json
tokenizer_path: data/processed/peter_pan_tokenizer_char.json
tokenizer_type: char
block_size: 64
train_split: 0.9

model:
vocab_size: 256
block_size: 64
n_layer: 4
n_head: 4
n_embd: 128
dropout: 0.1

train:
run_name: gptiny_peterpan_char_5k_lr1e-3_earlystop_seed2027
runs_dir: runs
batch_size: 16
max_steps: 5000
learning_rate: 0.001
weight_decay: 0.0
log_interval: 100
eval_interval: 250
eval_batches: null
early_stopping_patience: 3
early_stopping_min_delta: 0.0
sample_prompt: Once
sample_max_new_tokens: 100
sample_temperature: 1.0
sample_top_k: null
sample_seed: 1337
sample_greedy: false
seed: 2027
36 changes: 36 additions & 0 deletions configs/gptiny_peterpan_char_5k_lr1e-3_earlystop_seed4242.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
data:
input_path: data/raw/peter_pan_body.txt
prepared_path: data/processed/peter_pan_corpus.txt
manifest_path: data/processed/peter_pan_corpus_manifest.json
tokenizer_path: data/processed/peter_pan_tokenizer_char.json
tokenizer_type: char
block_size: 64
train_split: 0.9

model:
vocab_size: 256
block_size: 64
n_layer: 4
n_head: 4
n_embd: 128
dropout: 0.1

train:
run_name: gptiny_peterpan_char_5k_lr1e-3_earlystop_seed4242
runs_dir: runs
batch_size: 16
max_steps: 5000
learning_rate: 0.001
weight_decay: 0.0
log_interval: 100
eval_interval: 250
eval_batches: null
early_stopping_patience: 3
early_stopping_min_delta: 0.0
sample_prompt: Once
sample_max_new_tokens: 100
sample_temperature: 1.0
sample_top_k: null
sample_seed: 1337
sample_greedy: false
seed: 4242
2 changes: 1 addition & 1 deletion docs/codex/build-and-test.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,7 +50,7 @@ The individual commands remain canonical and are listed below.
python -m pytest
```

Current expected result after milestone 024+: at least 146 tests passing with at least 90% coverage.
Current expected result after milestone 025+: at least 155 tests passing with at least 90% coverage.

### Compile Check

Expand Down
5 changes: 5 additions & 0 deletions docs/codex/experiments.md
Original file line number Diff line number Diff line change
Expand Up @@ -165,3 +165,8 @@ character control. The next modeling question should test split/corpus robustnes
Milestone 024 uses a deterministically extracted, near-size-matched Peter Pan corpus. ByteBPE512
reaches best BPC `2.1539` versus `2.1721` for character. The direction replicates, but the 0.0181
BPC margin from one shared seed is weaker evidence than a corpus-by-seed matrix.

Milestone 025 completes the balanced corpus-by-seed matrix. ByteBPE512 beats character in all six
same-seed pairs. Mean candidate-minus-reference BPC is `-0.0619` on Alice and `-0.0252` on Peter
Pan; the corpus interaction is `+0.0367`, so report a robust direction but corpus-dependent size.
The next contract change should introduce a sealed test segment before further model selection.
8 changes: 7 additions & 1 deletion docs/experiments.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,7 @@ not a replacement for the original reports.
| [022 Early Stopping and Regularization](../experiments/022-early-stopping-and-regularization.md) | Training control | Patience-3 stopping reproduced the step-1750 optimum and halved runtime; weight decay `0.01` was effectively neutral. |
| [023 Multi-Seed Robustness](../experiments/023-multi-seed-robustness.md) | Robustness | Three preregistered seeds average best BPC `2.0225 ± 0.0124`; every seed beats the corrected character control. |
| [024 Cross-Corpus Robustness](../experiments/024-cross-corpus-robustness.md) | External validity | On near-size-matched Peter Pan, ByteBPE512 narrowly beats character at `2.1539` versus `2.1721` BPC. |
| [025 Corpus-by-Seed Matrix](../experiments/025-corpus-by-seed-matrix.md) | Factorial robustness | ByteBPE512 wins all six paired comparisons; mean advantage is `0.0619` BPC on Alice and `0.0252` on Peter Pan. |

## Topic Shortcuts

Expand All @@ -56,7 +57,8 @@ not a replacement for the original reports.
- Tokenization: [016](../experiments/016-tokenization-study.md),
[020](../experiments/020-bpe-context-and-learning-rate.md),
[021](../experiments/021-boundary-aware-byte-bpe.md),
[024](../experiments/024-cross-corpus-robustness.md).
[024](../experiments/024-cross-corpus-robustness.md),
[025](../experiments/025-corpus-by-seed-matrix.md).

## Current Status

Expand Down Expand Up @@ -92,3 +94,7 @@ Milestone 024 changes the data distribution to a near-size-matched Peter Pan cor
reaches best BPC `2.1539` versus `2.1721` for character, reproducing the direction with a much
smaller 0.83% margin. This supports limited cross-corpus robustness, not a universal tokenizer
advantage; a corpus-by-seed matrix is the next stronger test.

Milestone 025 completes that balanced matrix. ByteBPE512 beats character for seeds 1337, 2027, and
4242 on both Alice and Peter Pan. Its paired mean advantage is `0.0619` BPC on Alice and `0.0252`
on Peter Pan; the `+0.0367` BPC interaction shows that effect magnitude remains corpus-dependent.
20 changes: 20 additions & 0 deletions docs/training.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,8 @@ compileall, and Markdown links. `make smoke` writes an ignored smoke run, and
| `configs/gptiny_bpe128_5k_lr5e-4_ctx64.yaml` | Longer-context lower-learning-rate BPE128 run. |
| `configs/gptiny_peterpan_char_5k_lr1e-3_earlystop.yaml` | Peter Pan character control. |
| `configs/gptiny_peterpan_bytebpe512_5k_lr1e-3_ctx37_earlystop.yaml` | Context-matched Peter Pan ByteBPE512 run. |
| `configs/gptiny_char_5k_lr1e-3_earlystop*.yaml` | Three-seed Alice character early-stopping controls. |
| `configs/gptiny_peterpan_*_earlystop_seed*.yaml` | Additional Peter Pan matrix seeds. |

## Corpus Preparation

Expand Down Expand Up @@ -128,6 +130,21 @@ python scripts/summarize_runs.py runs/<name-a>/<id> runs/<name-b>/<id> runs/<nam
The command reads each run's config and summary, requires distinct training seeds, and prints every
observation plus mean, population standard deviation, minimum, and maximum.

For a balanced two-corpus, two-tokenizer design, label every run explicitly:

```bash
python scripts/summarize_matrix.py \
--reference-tokenizer char --candidate-tokenizer byte_bpe \
--cell alice char runs/<name>/<id> \
--cell alice byte_bpe runs/<name>/<id> \
--cell peter_pan char runs/<name>/<id> \
--cell peter_pan byte_bpe runs/<name>/<id>
```

Repeat `--cell` for every seed. The analyzer requires exactly two distinct corpus checksums, both
tokenizers on both corpora, a shared seed set, within-cell experiment fingerprints, and complete
schema-v2 summaries. Reported contrasts are candidate minus reference BPC and are paired by seed.

## Generation Controls

```bash
Expand Down Expand Up @@ -185,6 +202,9 @@ Experiment 024 tests a second book. On the near-size-matched Peter Pan corpus, B
best BPC `2.1539` versus `2.1721` for character and stops at step 2,750. The 0.83% advantage is a
cross-corpus replication in direction, but too small and sparsely sampled to establish a stable
effect size.
Experiment 025 completes the 2-tokenizer × 2-corpus × 3-seed matrix. ByteBPE512 wins every paired
comparison. Its mean advantage is `0.0619` BPC on Alice and `0.0252` on Peter Pan, so the direction
is robust within the matrix while the effect magnitude remains corpus-dependent.

## Artifact Policy

Expand Down
Loading
Loading