Just me going through some AI papers w/ code implementation to understand math behind them.
Model I use GLM 5.2 w/ PI Harness
-
- XGBoost: A Scalable Tree Boosting System — Boosted decision trees with regularized split scoring.
-
- Learning Representations by Back-Propagating Errors — Manual backpropagation for a small neural network.
-
- Deep Residual Learning for Image Recognition — Residual blocks and the optimization benefit of skip connections.
-
- Batch Normalization + Adam — Two training tools: normalized activations and adaptive updates.
-
- Attention Is All You Need — Self-attention and the core Transformer block.
-
- An Image is Worth 16x16 Words — Image patches treated as tokens for classification.
-
- LLaMA: Open and Efficient Foundation Language Models — RMSNorm, SwiGLU, RoPE, and a modern decoder block.
-
- FlashAttention — Exact attention with less memory movement.
-
- LoRA: Low-Rank Adaptation of Large Language Models — Low-rank adapters for efficient fine-tuning.
-
- Direct Preference Optimization — Preference learning without PPO or a separate reward model.
-
- Dense Passage Retrieval — Dual-encoder retrieval with contrastive training.
-
- Switch Transformers — Sparse experts with simple top-1 routing.
Each paper folder stays flat: README.md, a few Python files, one optional notebook, and result images when they exist.
Core code belongs in Python files. The notebook is only for quick checks, plots, or visual inspection.