Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 4 additions & 4 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -90,10 +90,10 @@ If a relevant check cannot be run, state why and what remains unverified.

## Current Technical Status

Milestone 021 is the latest modeling evidence. Boundary-aware UTF-8 ByteBPE320 and ByteBPE512 reach
best full-validation BPC `2.0286` and `2.0083`, beating the corrected character (`2.0760`) and
BPE128 (`2.0976`) controls. ByteBPE512 overfits after step 1,750 and ends at BPC `2.2450`; future
work should test early stopping or regularization before increasing model scale.
Milestone 022 is the latest modeling evidence. Patience-3 validation early stopping reproduces
ByteBPE512's best full-validation BPC `2.0083`, stops at step 2,500, and roughly halves runtime.
Weight decay `0.01` reaches `2.0080`, too small a difference to claim improvement. Future work should
test the tokenizer result across seeds or corpus splits rather than add another narrow hyperparameter.

## Safety And Security

Expand Down
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,7 @@ chronological 90/10 split unless noted otherwise.
| [019](experiments/019-professionalization-and-corrected-evaluation.md) | Corrected character vs BPE128 | Full-validation best bits/character | `2.0760` char vs `2.0976` BPE128 | BPE128 shortened sequences but remained narrowly worse after removing leakage and coverage bias. |
| [020](experiments/020-bpe-context-and-learning-rate.md) | BPE context/LR controls | Best BPE128 bits/character | `2.0976` remains best | Matched character context and `5e-4` LR changed timing/diversity but did not improve held-out BPC. |
| [021](experiments/021-boundary-aware-byte-bpe.md) | Boundary-aware ByteBPE320/512 | Full-validation best bits/character | `2.0286` / `2.0083` | Both beat the corrected character and BPE128 controls; ByteBPE512 overfit after step 1750. |
| [022](experiments/022-early-stopping-and-regularization.md) | ByteBPE512 early stopping / weight decay | Actual steps and best BPC | `2500`, `2.0083` / `2.0080` | Early stopping halves runtime and limits final degradation; weight decay `0.01` is effectively neutral. |

Token-level loss and perplexity are not directly comparable between character
and BPE tokenizers because they predict different units. The tokenizer
Expand Down Expand Up @@ -81,7 +82,7 @@ shapes link directly to the implementation and its tests.
| Tokenization | Character, educational character-BPE, and lossless boundary-aware UTF-8 byte-BPE tokenizers. |
| Model | Decoder-only GPT-style Transformer with causal self-attention. |
| Evaluation | Uniform, unigram, and add-one smoothed bigram baselines. |
| Training | Config-driven training with validation loss, progress logging, checkpoints, metrics, summaries, and samples. |
| Training | Config-driven training with validation loss, optional early stopping, progress logging, checkpoints, metrics, summaries, and samples. |
| Run records | Preserved run directories with copied dataset manifests and selected provenance fields in `summary.json`. |
| Generation | `max_new_tokens`, `temperature`, `top_k`, `seed`, and greedy decoding. |
| Tests | Focused tests for data, baselines, model shape, training artifacts, run utilities, and generation behavior. |
Expand Down
38 changes: 38 additions & 0 deletions configs/gptiny_bytebpe512_5k_lr1e-3_ctx37_earlystop.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
data:
input_path: data/raw/input.txt
prepared_path: data/processed/corpus.txt
manifest_path: data/processed/corpus_manifest.json
tokenizer_path: data/processed/tokenizer_bytebpe512.json
tokenizer_type: byte_bpe
bpe_vocab_size: 512
bpe_min_frequency: 2
block_size: 37
train_split: 0.9

model:
vocab_size: 512
block_size: 37
n_layer: 4
n_head: 4
n_embd: 128
dropout: 0.1

train:
run_name: gptiny_bytebpe512_5k_lr1e-3_ctx37_earlystop
runs_dir: runs
batch_size: 27
max_steps: 5000
learning_rate: 0.001
weight_decay: 0.0
log_interval: 100
eval_interval: 250
eval_batches: null
early_stopping_patience: 3
early_stopping_min_delta: 0.0
sample_prompt: Once
sample_max_new_tokens: 100
sample_temperature: 1.0
sample_top_k: null
sample_seed: 1337
sample_greedy: false
seed: 1337
38 changes: 38 additions & 0 deletions configs/gptiny_bytebpe512_5k_lr1e-3_ctx37_wd0.01_earlystop.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
data:
input_path: data/raw/input.txt
prepared_path: data/processed/corpus.txt
manifest_path: data/processed/corpus_manifest.json
tokenizer_path: data/processed/tokenizer_bytebpe512.json
tokenizer_type: byte_bpe
bpe_vocab_size: 512
bpe_min_frequency: 2
block_size: 37
train_split: 0.9

model:
vocab_size: 512
block_size: 37
n_layer: 4
n_head: 4
n_embd: 128
dropout: 0.1

train:
run_name: gptiny_bytebpe512_5k_lr1e-3_ctx37_wd0.01_earlystop
runs_dir: runs
batch_size: 27
max_steps: 5000
learning_rate: 0.001
weight_decay: 0.01
log_interval: 100
eval_interval: 250
eval_batches: null
early_stopping_patience: 3
early_stopping_min_delta: 0.0
sample_prompt: Once
sample_max_new_tokens: 100
sample_temperature: 1.0
sample_top_k: null
sample_seed: 1337
sample_greedy: false
seed: 1337
2 changes: 1 addition & 1 deletion docs/codex/build-and-test.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,7 +50,7 @@ The individual commands remain canonical and are listed below.
python -m pytest
```

Current expected result after milestone 019+: 90 tests passing with at least 90% coverage.
Current expected result after milestone 022+: at least 121 tests passing with at least 90% coverage.

### Compile Check

Expand Down
4 changes: 4 additions & 0 deletions docs/codex/experiments.md
Original file line number Diff line number Diff line change
Expand Up @@ -153,3 +153,7 @@ tokenizer design rather than continue fine tuning around BPE128.
Milestone 021 altered tokenizer design with boundary-aware ByteBPE320/512. Best full-validation BPC
improved to `2.0286` and `2.0083`; the 512-token model overfit sharply after step 1,750, so the next
controlled question is early stopping or modest regularization rather than more scale.

Milestone 022 adds patience-3 validation early stopping. It reproduces ByteBPE512's step-1,750 best
checkpoint and stops at step 2,500, roughly halving runtime. Weight decay `0.01` is effectively
neutral at best BPC `2.0080`; future work should test robustness across seeds or corpus splits.
5 changes: 5 additions & 0 deletions docs/experiments.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,7 @@ not a replacement for the original reports.
| [019 Professionalization and Corrected Evaluation](../experiments/019-professionalization-and-corrected-evaluation.md) | Scientific hardening | Corrects tokenizer leakage and validation coverage, hardens artifacts and quality gates, adds the theory handbook, and reruns the controlled comparison. |
| [020 BPE Context and Learning Rate](../experiments/020-bpe-context-and-learning-rate.md) | Tokenizer diagnostics | Matching character context and lowering BPE learning rate did not beat the corrected BPE128 control or character model. |
| [021 Boundary-Aware Byte BPE](../experiments/021-boundary-aware-byte-bpe.md) | Tokenizer design | Lossless boundary-aware ByteBPE320/512 beat both corrected controls on best BPC; ByteBPE512 reached `2.0083` but overfit early. |
| [022 Early Stopping and Regularization](../experiments/022-early-stopping-and-regularization.md) | Training control | Patience-3 stopping reproduced the step-1750 optimum and halved runtime; weight decay `0.01` was effectively neutral. |

## Topic Shortcuts

Expand Down Expand Up @@ -74,3 +75,7 @@ reported, but it did not beat the character control.
Milestone 021 adds a lossless UTF-8 byte fallback and whitespace-boundary-aware merges. ByteBPE320
reached best BPC `2.0286`; ByteBPE512 reached `2.0083`, the strongest corrected result so far.
The 512-token run's final BPC rose to `2.2450`, making early best-checkpoint selection essential.

Milestone 022 operationalizes that result. Patience-3 early stopping terminates at step 2,500 and
retains best BPC `2.0083`, while stopped-final BPC improves from `2.2450` to `2.0554`. Weight decay
`0.01` reaches best BPC `2.0080`; the `0.00025` difference is too small to interpret as a real gain.
9 changes: 9 additions & 0 deletions docs/training.md
Original file line number Diff line number Diff line change
Expand Up @@ -111,6 +111,12 @@ python scripts/show_run.py --run latest --run-name gptiny
`show_run.py` prints run paths, dataset summary fields, the latest metric, and
the saved sample.

`summary.json` records the configured `max_steps` ceiling and `actual_steps`. Optional
`early_stopping_patience` counts consecutive validation events without an improvement larger than
`early_stopping_min_delta`; `stopped_early`, `stop_reason`, and the terminal state make the decision
inspectable. The final checkpoint is the model at the stop step, while `best_checkpoint.pt` remains
the lowest observed validation-loss model.

## Generation Controls

```bash
Expand Down Expand Up @@ -158,6 +164,9 @@ Experiment 020 found that matching BPE character context and halving its learnin
the experiment-019 BPE control. Experiment 021 changed the tokenizer itself: boundary-aware
ByteBPE320 and ByteBPE512 reached best BPC `2.0286` and `2.0083`, both beating the character
control, though ByteBPE512 overfit sharply after step 1,750.
Experiment 022 adds patience-3 early stopping: it ends at step 2,500, halves runtime, and avoids most
final-checkpoint degradation. Weight decay `0.01` changes best BPC by only `0.00025`, which is not
meaningful evidence of improvement.

## Artifact Policy

Expand Down
81 changes: 81 additions & 0 deletions experiments/022-early-stopping-and-regularization.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,81 @@
# 022 — Early Stopping and Regularization

## Goal

Turn milestone 021's early ByteBPE512 optimum into explicit, inspectable training behavior, then test
whether modest AdamW weight decay delays overfitting. The intervention stays narrow: the same seed,
tokenizer, model, data order, validation contract, and 5,000-step ceiling.

## Implementation

`TrainConfig` now accepts `early_stopping_patience` and `early_stopping_min_delta`. Patience counts
full validation events without a meaningful improvement. The lowest numerical validation checkpoint
is still saved independently. Run summaries add `actual_steps`, `stopped_early`, `stop_reason`, and
the terminal early-stopping state; the final checkpoint represents the actual stop step.

Both experiment configs use patience 3, minimum delta 0, full validation every 250 steps, and the
milestone-021 ByteBPE512 setup. One keeps weight decay 0; the other changes only weight decay to
`0.01`.

## Results

| setting | actual / ceiling steps | best step | best loss | best BPC | stopped-final BPC | duration |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| 021 control, no stopping | 5,000 / 5,000 | 1,750 | 2.375155 | 2.008269 | 2.244955 | 422.7s |
| early stopping | 2,500 / 5,000 | 1,750 | 2.375155 | 2.008269 | 2.055447 | 209.3s |
| early stopping + WD 0.01 | 2,500 / 5,000 | 1,750 | 2.374858 | 2.008018 | 2.053365 | 204.6s |

The unregularized run reproduces every observed milestone-021 loss through step 2,500, showing that
the stopping feature does not perturb optimization. It cuts runtime by 50.5% and avoids most of the
late final-checkpoint degradation. This is an orchestration win, not a new best model.

Weight decay improves best BPC by only `0.000251` (0.0125%) and stopped-final BPC by `0.002082`.
With one deterministic seed, those differences are practically neutral and do not justify changing
the default research conclusion. Both runs stop after the three non-improving evaluations at steps
2,000, 2,250, and 2,500.

## Controlled Generation

Prompt `Once`, 100 new tokens. Seeded decoding uses temperature 0.8, top-k 10, seed 1337.

| setting/checkpoint | greedy distinct-2 | seeded distinct-2 |
| --- | ---: | ---: |
| early-stop best | 0.5217 | 0.5417 |
| early-stop final | 0.5723 | 0.5980 |
| WD 0.01 best | 0.4326 | 0.5439 |
| WD 0.01 final | 0.4253 | 0.6380 |

Best-checkpoint likelihood and distinct-2 again disagree. Weight decay does not visibly solve
coherence or repetition; the samples remain locally plausible but semantically unstable.

## Exact Commands

```bash
uv run --frozen --extra dev python scripts/prepare_data.py --config configs/gptiny_bytebpe512_5k_lr1e-3_ctx37_earlystop.yaml
uv run --frozen --extra dev python scripts/prepare_data.py --config configs/gptiny_bytebpe512_5k_lr1e-3_ctx37_wd0.01_earlystop.yaml
uv run --frozen --extra dev python scripts/train.py --config configs/gptiny_bytebpe512_5k_lr1e-3_ctx37_earlystop.yaml
uv run --frozen --extra dev python scripts/train.py --config configs/gptiny_bytebpe512_5k_lr1e-3_ctx37_wd0.01_earlystop.yaml
uv run --frozen --extra dev python scripts/show_run.py --run latest --run-name gptiny_bytebpe512_5k_lr1e-3_ctx37_earlystop
uv run --frozen --extra dev python scripts/show_run.py --run latest --run-name gptiny_bytebpe512_5k_lr1e-3_ctx37_wd0.01_earlystop
uv run --frozen --extra dev make check
```

Generation used `scripts/generate.py` for both runs and `best`/`final` checkpoint kinds, once with
`--greedy --diagnostics` and once with
`--temperature 0.8 --top-k 10 --seed 1337 --diagnostics`.

Run paths:

- `runs/gptiny_bytebpe512_5k_lr1e-3_ctx37_earlystop/2026-07-12_22-12-08`
- `runs/gptiny_bytebpe512_5k_lr1e-3_ctx37_wd0.01_earlystop/2026-07-12_22-16-28`

## Limitations And Next Step

- Patience is evaluated only every 250 steps, so stopping reacts with bounded delay.
- One seed and one chronological split cannot establish whether the tiny WD difference is stable.
- Early stopping observes the validation set repeatedly and is part of model selection.
- Generation diagnostics are surface statistics, not human evaluation.

The next useful milestone is robustness rather than another single-run hyperparameter: repeat the
ByteBPE512 early-stopping control across multiple seeds and report mean, spread, stop-step variation,
and generation diagnostics without selecting the best seed.
24 changes: 24 additions & 0 deletions notes/04-training.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,30 @@ Using sampled validation adds estimator variance. Official configs therefore use
deterministic validation; `eval_batches` remains available for quick experiments and records its
coverage explicitly.

### Validation-based early stopping

Let validation be observed at evaluation index (j), with loss (L_j). Given minimum meaningful
improvement (delta\geq0), maintain a reference (R) and stale count (q):

\[
(R,q)\leftarrow
\begin{cases}
(L_j,0), & L_j < R-\delta,\\
(R,q+1), & \text{otherwise}.
\end{cases}
\]

Training stops when (q\geq P), where (P) is patience measured in validation events—not gradient
steps or epochs. smaLLM still stores the numerically lowest validation checkpoint independently of
the early-stopping reference. This matters when improvements smaller than (delta) are real enough
to preserve but intentionally too small to reset patience.

Early stopping is a sequential model-selection rule, not regularization: it limits exposure to
overfitting but does not change the objective or gradients before the stop. Its latency is bounded by
(P\times\texttt{eval_interval}), and sampled validation can make the stopping time noisy. Official
modeling configs therefore use full deterministic validation. Summaries record the step ceiling,
actual steps, stop reason, patience, minimum delta, and terminal stale count.

## Determinism limits

Setting Python and PyTorch seeds controls many random choices, but bitwise identity can still fail
Expand Down
43 changes: 37 additions & 6 deletions src/smallm/config.py
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
from __future__ import annotations

from dataclasses import dataclass
from math import isfinite
from pathlib import Path
from typing import Any

Expand All @@ -10,6 +11,15 @@
yaml = None


def _is_bounded_finite_number(value: object) -> bool:
return (
isinstance(value, int | float)
and not isinstance(value, bool)
and (not isinstance(value, float) or isfinite(value))
and abs(value) <= 1_000_000_000
)


@dataclass(frozen=True)
class DataConfig:
input_path: str = "data/raw/input.txt"
Expand Down Expand Up @@ -77,6 +87,8 @@ class TrainConfig:
log_interval: int = 10
eval_interval: int = 100
eval_batches: int | None = 5
early_stopping_patience: int | None = None
early_stopping_min_delta: float = 0.0
sample_prompt: str = "Once"
sample_max_new_tokens: int = 100
sample_temperature: float = 1.0
Expand All @@ -93,16 +105,35 @@ def __post_init__(self) -> None:
raise ValueError(f"train.{name} must be positive")
if self.eval_batches is not None and self.eval_batches <= 0:
raise ValueError("train.eval_batches must be positive or null")
if self.learning_rate <= 0:
raise ValueError("train.learning_rate must be positive")
if self.weight_decay < 0:
raise ValueError("train.weight_decay must be non-negative")
if self.early_stopping_patience is not None and (
not isinstance(self.early_stopping_patience, int)
or isinstance(self.early_stopping_patience, bool)
or self.early_stopping_patience <= 0
):
raise ValueError("train.early_stopping_patience must be a positive integer or null")
if (
not isinstance(self.early_stopping_min_delta, int | float)
or isinstance(self.early_stopping_min_delta, bool)
or (
isinstance(self.early_stopping_min_delta, float)
and not isfinite(self.early_stopping_min_delta)
)
or self.early_stopping_min_delta < 0
or self.early_stopping_min_delta > 1_000_000_000
):
raise ValueError(
"train.early_stopping_min_delta must be finite and between 0 and 1000000000"
)
if not _is_bounded_finite_number(self.learning_rate) or self.learning_rate <= 0:
raise ValueError("train.learning_rate must be finite and positive")
if not _is_bounded_finite_number(self.weight_decay) or self.weight_decay < 0:
raise ValueError("train.weight_decay must be finite and non-negative")
if not self.sample_prompt:
raise ValueError("train.sample_prompt must not be empty")
if self.sample_max_new_tokens < 0:
raise ValueError("train.sample_max_new_tokens must be non-negative")
if self.sample_temperature <= 0:
raise ValueError("train.sample_temperature must be positive")
if not _is_bounded_finite_number(self.sample_temperature) or self.sample_temperature <= 0:
raise ValueError("train.sample_temperature must be finite and positive")
if self.sample_top_k is not None and self.sample_top_k <= 0:
raise ValueError("train.sample_top_k must be positive or null")

Expand Down
6 changes: 3 additions & 3 deletions src/smallm/training/artifacts.py
Original file line number Diff line number Diff line change
Expand Up @@ -48,7 +48,7 @@ def write_config_snapshot(path: str | Path, config: ExperimentConfig) -> None:
for section, values in data.items():
lines.append(f"{section}:")
for key, value in values.items():
lines.append(f" {key}: {json.dumps(value, ensure_ascii=False)}")
lines.append(f" {key}: {json.dumps(value, ensure_ascii=False, allow_nan=False)}")
lines.append("")
atomic_write_text(output, "\n".join(lines).rstrip() + "\n")

Expand All @@ -60,7 +60,7 @@ def __init__(self, path: str | Path) -> None:
self._handle = self.path.open("w", encoding="utf-8")

def write(self, record: dict[str, Any]) -> None:
self._handle.write(json.dumps(record, sort_keys=True) + "\n")
self._handle.write(json.dumps(record, sort_keys=True, allow_nan=False) + "\n")
self._handle.flush()
os.fsync(self._handle.fileno())

Expand All @@ -82,7 +82,7 @@ def __exit__(
def write_json(path: str | Path, payload: dict[str, Any]) -> None:
output = Path(path)
output.parent.mkdir(parents=True, exist_ok=True)
atomic_write_text(output, json.dumps(payload, indent=2, sort_keys=True) + "\n")
atomic_write_text(output, json.dumps(payload, indent=2, sort_keys=True, allow_nan=False) + "\n")


def load_dataset_manifest(path: str | Path) -> dict[str, Any]:
Expand Down
Loading
Loading