# PROGRESS — BLT-LLM Last updated: 2026-06-08 ## Current phase **Phase 3 — End-to-end training: DONE & verified.** Full 20k-step run on the 2.2GB TinyStories TRAIN split finished: held-out **BPB 8.0 → 0.71** (best 0.7077 @ step 19500), plateaued; causality gate passes; samples are coherent TinyStories prose. All three Phase 3 done-criteria met. Next up: Phase 4 (scale config) — **rented GPU only, not local**. ## Done & verified - Repo under git, pushed to private GitHub. `.gitignore` excludes `reference/`, `data/`, `logs/`, caches. - `blt/` package scaffolded: `config.py`, `data.py`, `entropy_model.py`, `train_entropy.py`, `patcher.py`, `tests/`. - **Entropy byte-LM** (`blt/entropy_model.py`): **1,705,216 params**. 2 layers × dim 256, 4 heads, sliding-window (local block-causal) attention window 128, RoPE θ=500000, RMSNorm, SwiGLU, tied embeddings. - **Data** (`blt/data.py`): `TinyStoriesV2-GPT4-valid.txt` via `huggingface_hub` → ~22.3M train bytes / 225K val bytes. - **Training** (`blt/train_entropy.py`): 8000 steps; AdamW β(0.9,0.95) wd 0.1 clip 1.0; cosine LR 4e-4, 300 warmup. **Final val BPB = 0.9278** (step-0 baseline 8.06 = log₂256). ~30 min on RTX 3050 6GB. → `checkpoints/entropy_model.pt`. - **Patcher** (`blt/patcher.py`): entropy in nats (`-Σ p·log p`); causal global-threshold rule (byte *i* starts a patch iff `H(x_i | x_ θ`); threshold tuning by bisection. Tuned **θ = 1.0917** (official repo default ≈ 1.335). → `checkpoints/patcher_threshold.json`. - **Phase 0 isolation test** (`blt/tests/test_phase0_patcher.py`): **ALL PASS** - achieved avg patch on held-out 100 kB = **4.36** (target 4–6) ✓ - reproducible ✓, causal / no future leakage ✓, patch lengths sum to n ✓, first byte = boundary ✓ - in-dist "Once upon a time…" → big patches (low entropy); OOD GoT sentence over-segments (expected). - `artifacts/phase0_entropy.png` (paper Fig. 4 style). - **BLT model** (`blt/model.py`): **55,392,768 params** (~55M). Assembles encoder → global → decoder. hE=hD=256, hG=768 (k=3), encoder 1×256/4h, global 6×768/12h, decoder 4×256/4h, cross 4h. Tied byte embed/output. - **Shared layers** (`blt/layers.py`): RMSNorm, RoPE θ=500000, SwiGLU, SelfAttention (windowed or full causal), TransformerBlock. - **Hash n-grams** (`blt/ngram_hash.py`): RollPolyHash (reference primes), precomputed IDs, per-n tables (4096×256), e_i normalized by (n_sizes+1). - **Encoder** (`blt/encoder.py`): windowed-causal byte layer; **max-pool** seeds patch queries; encoder cross-attn (patch↔own bytes) AFTER the layer. Returns byte_hidden (carry-over, not detached) + patch_reps. - **Global** (`blt/global_transformer.py`): block-causal over patches + key-padding for padded patches. - **Decoder** (`blt/decoder.py`): cross-attn BEFORE layers; **byte i attends previous patch o_{j-1}** (causal shift — first-patch bytes get no global ctx); windowed-causal byte layers; output → 256. - **Dummy patcher** (`blt/patcher.py:strided_patch_lengths`): fixed stride, but emits the same [B,M] format as the entropy patcher (model never assumes equal size). - **Tests PASS**: `test_shapes.py` (hand-checked 2-patch masks + T→M→T transitions + padded patches all finite); `test_overfit.py` (single batch, loss 5.63→0.047 in 41 steps). ### Phase 2 (real entropy patcher) — DONE & verified - Added batched patching to `blt/patcher.py`: `batch_entropies`, `batch_boundaries`, `boundaries_to_patch_lengths` ([B,Mmax] right-padded), `entropy_patch_lengths` (drop-in for `strided_patch_lengths`). - `test_phase2.py` PASS: real patcher gives variable per-row patches (counts e.g. [5,19,23,23], lengths 1..30), exercises the padded-patch path with real data (finite), and the single-batch overfit still collapses (5.63→0.0496). - **HARD GATE — `test_causality.py` PASS (no future leakage), required before Phase 2 done:** full model + real patcher, 4 batches × 7 t-values × 2 tests: - gradient leakage: ∂logit_t/∂embed_s == 0 (exact) for all s>t and all other batch rows; - perturbation: scramble bytes>t (recompute hashes) → logits≤t bit-identical (diff exactly 0). Added `inputs_embeds` path to `model.py`/`encoder.py` to put grad on input byte embeds. - Phase 1 strided overfit + shapes still PASS (no regression). ### Phase 3 (train end-to-end) — DONE & verified - **Full run result** (`python -m blt.train --data train --steps 20000 --batch-size 8 --seq-len 512 --grad-accum 2`, ~2.4 h on RTX 3050 6GB, ~2.33 it/s): held-out BPB **8.07 → 0.71** (best **0.7077** @ step 19500), curve flat over the last ~4k steps → plateaued. Untrained baseline = 8.0 (log₂256); brief's rough target ≤ 1.5 → beaten. - BPB curve: 1.39(500) 1.16(1k) 1.04(2k) 0.96(3k) 0.87(5k) 0.80(9k) 0.73(13k) 0.71(17k+). - Best weights+optimizer → `checkpoints/blt_model.pt` (635M, gitignored; see below). - **Final sample** (greedy-ish, temp 0.7, from the best checkpoint): *"Once upon a time, in a small house, there was a boy named Tim. One day, Tim went to the store with his mom. They needed to buy a toy… Tim said to his mom, \"Mom, can I give t…"* → coherent, named characters, dialogue, paragraph breaks. Recognizable text ✓. - **Causality leakage gate**: PASS (exact-zero; architecture-level, re-run after the Phase 3 code and unchanged since). Three Phase-3 done-criteria all met. ### Phase 3 (train end-to-end) — code (built earlier in the phase) - **`blt/config.py:BLTTrainConfig`**: AdamW β(0.9,0.95) wd 0.1 clip 1.0; cosine LR 4e-4, 300 warmup; defaults sized for 6GB (batch 8 × grad_accum 2, seq_len 512). Eval = BPB. - **`blt/train.py`**: full loop. Each batch → FROZEN entropy patcher decides per-row patch boundaries (causal, no grad) → precomputed n-gram hash IDs → BLT forward → next-byte CE. Held-out BPB eval, best+last checkpointing, `--resume`, in-loop sampling, `--smoke`, `--data train|valid`. Smoke (40 steps, valid): BPB 8.07 → 4.6; ckpt+resume+sample OK. Target batch (8×2, seq 512) fits 6GB with headroom. - **`blt/generate.py`**: autoregressive byte generation. Recomputes patch boundaries over the running prefix each step (correct-first; causal patcher → no future leak). Temperature / top-k. Loads BLT ckpt + frozen entropy model + θ. Verified end-to-end on the smoke ckpt. - **`test_causality.py` HARD GATE re-run after the Phase 3 code** — still ALL PASS (worst future-grad / cross-batch / perturb |diff| all 0.000e+00). Safe to train. - **Real run launched** (background, detached): `python -m blt.train --data train --steps 20000 --batch-size 8 --seq-len 512 --grad-accum 2` → `logs/train.log`, ckpts in `checkpoints/`. Downloads the 2.2GB TinyStories TRAIN split first (cached), then trains (~2.2B bytes ≫ 55M params, so it won't memorize). Resumable with `--resume`. ## In progress - Nothing. Phases 0–3 complete and **publicly released** on both GitHub (public) and Hugging Face. Best model weights backed up via **git-LFS** (`checkpoints/blt_model_weights.pt`, 212M, model-only, val_bpb 0.7078 @ step 19500). The 635M full checkpoint (with optimizer) stays local + gitignored. git-lfs binary in `~/.local/bin` (no system install). - **Public artifacts** (anyone can clone + `pip install -r requirements.txt` + `python -m blt.generate`): - GitHub (public): https://github.com/shaikh-saud705/blt-llm — code + weights (LFS) + README. - HF model (public): https://huggingface.co/sssssaud/blt-llm-tinystories-55m — code mirror + weights + entropy patcher + θ + model card. - `README.md` (root) + `requirements.txt` added; `blt/generate.py` now defaults `--ckpt` to `checkpoints/blt_model_weights.pt` so it runs straight from a clone. ## Next action - **Phase 4 (scale-up) — rented GPU only, NEVER local** (BLT_LLM.md sec 4 / gotcha #10): swap TinyConfig for the ~1.5B config (paper 1B row widened), add bf16/mixed precision, gradient checkpointing, FlexAttention for the patch-masked cross-attn, robust ckpt save/resume. Done when it trains stably on the rented GPU and checkpoints resume. The architecture itself is now proven end-to-end (this was the whole point of Phases 0–3). - Optional polish before scaling: trim model toward 50M (ngram_table_size or one global layer) if desired. (Tiny checkpoint is already backed up via git-LFS — see In progress.) ## Blockers - None. ## Reproduce ``` python -m blt.train_entropy --steps 8000 # Phase 0: train entropy model -> checkpoints/entropy_model.pt python -m blt.tests.test_phase0_patcher # Phase 0 isolation test -> artifacts/ + threshold json python -m blt.tests.test_shapes # Phase 1 shapes + masks python -m blt.tests.test_overfit # Phase 1 single-batch overfit python -m blt.tests.test_phase2 # Phase 2 real patcher + overfit python -m blt.tests.test_causality # HARD GATE: no future leakage (re-run after any mask change) python -m blt.train --smoke # Phase 3 quick wiring check (valid split, 40 steps) python -m blt.train --data train --steps 20000 --batch-size 8 --seq-len 512 --grad-accum 2 # Phase 3 full run python -m blt.generate --prompt "Once upon a time" # sample from the trained model ```