This guide groups the experiment reports by milestone. It is a navigation aid, not a replacement for the original reports.
| Experiment | Area | Result |
|---|---|---|
| 001 Tiny GPT Smoke Test | First model path | Initial tokenizer, dataset, model, train script, generate script, and tests passed under the original name. |
| 002 Training Logger and Tiny Config | Training visibility | Tiny config and progress logger produced readable loss/throughput output. |
| 003 Run Artifacts and Metrics | Run records | Training began preserving config, metrics, summary, checkpoint, and sample files. |
| 004 Run Discovery | Run navigation | Added run listing, latest-run resolution, and run inspection. |
| 005 Real Corpus Tiny GPT | Real text | GPTiny, then named Tiny GPT, trained on a 4,838-character prose corpus with vocab size 53. |
| 006 Baseline Evaluation | Baselines | GPTiny validation loss 2.4505 beat add-one bigram 2.5562 on the small corpus. |
| 007 Reproducible Corpus Preparation | Corpus prep | Baselines and training moved to a normalized prepared corpus path. |
| 008 Dataset Manifest and Checksums | Dataset identity | Corpus manifests began recording source metadata, checksums, counts, and normalization rules. |
| 009 Run Dataset Provenance | Run provenance | Runs copied dataset_manifest.json and stored dataset fields in summary.json. |
| 010 Sampling Controls | Generation | Added max_new_tokens, temperature, top_k, seed, and greedy decoding. |
| 011 Larger Public-Domain Corpus Experiment | Model quality | Larger corpus reached 144,530 prepared characters, but GPTiny loss 2.5914 trailed bigram 2.4340. |
| 012 Documentation and Portfolio Narrative | Documentation | Reframed the project docs around the current reproducible language-model lab. |
| 013 GPTiny Training Budget and Optimization | Training budget | Renamed the model family to GPTiny and found 2k/5k-step runs beat the larger-corpus bigram baseline. |
| 014 Optimizer and Sampling Diagnostics | Optimizer diagnostics | A 5k lr=0.001 run beat the 5k control, but greedy generation still collapsed. |
| 015 GPTiny Capacity and Generation Diagnostics | Capacity diagnostics | Wider/deeper GPTiny improved validation and generation diversity metrics, but prose remained incoherent. |
| 016 Tokenization Study | Tokenization | Simple BPE128 shortened sequences and made some greedy text more word-like, but underperformed the character control on estimated bits per character. |
| 017 Best-Checkpoint Evaluation | Checkpoint diagnostics | Added best-validation checkpoints; best checkpoints improved validation but not controlled generation for BPE128 or the character control. |
| 018 Launch Polish and Public Evidence | Launch polish | Added reviewer navigation, CI, local checks, and a public evidence summary. |
| 019 Professionalization and Corrected Evaluation | Scientific hardening | Corrects tokenizer leakage and validation coverage, hardens artifacts and quality gates, adds the theory handbook, and reruns the controlled comparison. |
| 020 BPE Context and Learning Rate | Tokenizer diagnostics | Matching character context and lowering BPE learning rate did not beat the corrected BPE128 control or character model. |
| 021 Boundary-Aware Byte BPE | Tokenizer design | Lossless boundary-aware ByteBPE320/512 beat both corrected controls on best BPC; ByteBPE512 reached 2.0083 but overfit early. |
| 022 Early Stopping and Regularization | Training control | Patience-3 stopping reproduced the step-1750 optimum and halved runtime; weight decay 0.01 was effectively neutral. |
| 023 Multi-Seed Robustness | Robustness | Three preregistered seeds average best BPC 2.0225 ± 0.0124; every seed beats the corrected character control. |
| 024 Cross-Corpus Robustness | External validity | On near-size-matched Peter Pan, ByteBPE512 narrowly beats character at 2.1539 versus 2.1721 BPC. |
| 025 Corpus-by-Seed Matrix | Factorial robustness | ByteBPE512 wins all six paired comparisons; mean advantage is 0.0619 BPC on Alice and 0.0252 on Peter Pan. |
| 026 Sealed Test Evaluation | Confirmatory evaluation | On untouched terminal segments, ByteBPE512 beats character by 0.0614 BPC on Alice and 0.0258 on Peter Pan. |
| 027 Hamlet External-Distribution Replication | Preregistered external validity | ByteBPE512 beats character by 0.0673 sealed-test BPC on a dramatic play; the terminal region is easier than validation for both models. |
| 028 Preregistered External Corpus Panel | Multi-corpus confirmatory panel | ByteBPE512 wins all six same-seed sealed-test pairs, averaging -0.1134 BPC on Art of War and -0.1543 on Lincoln. |
| 029 Final Capacity Panel Preregistration | Frozen final protocol | Commits the three-corpus × three-arm × three-seed capacity-control design before source access. |
| 030 Final Capacity Panel and Completion | Final evidence and closure | Reports the final near-parameter-matched sealed-test panel and permanently completes the project scope. |
- Model path: 001, 002, 005.
- Run artifacts: 003, 004, 009.
- Data provenance: 007, 008, 009.
- Evaluation: 006, 011, 016.
- Generation: 010, 011, 013, 014, 015, 016, 017.
- Tokenization: 016, 020, 021, 024, 025, 026, 027, 028, 029, and 030.
Milestone 030 records the complete preregistered capacity-controlled panel and closes the modeling roadmap. ByteBPE512 beats near-parameter-matched char136 in eight of nine sealed pairs; the narrow Douglass reversal and corpus-dependent effects bound the claim. The table and topic shortcuts above lead to the immutable reports; the durable lifecycle statement is Project Completion.