--- license: apache-2.0 language: - en pipeline_tag: text-generation tags: - byte-level - cascade - reasoning - pretraining library_name: pytorch --- # Cascade-1B-Logic Public experimental byte-level causal language model. Training is in progress and checkpoints are uploaded continuously. ## Architecture - Approximately 1B stored parameters and approximately one third active per token. - UTF-8 bytes, vocabulary 256, no external tokenizer. - Four local causal Transformer layers process every byte. - Entropy boundaries compress the stream to about one patch per eight bytes, with capacity for 192 patches per 1024-byte sequence. Over-cap boundaries are uniformly coalesced across the sequence instead of truncating its tail. - Three independent eight-layer global routes are trained cyclically, one full batch per route. - The previous completed global patch conditions the local prediction, preventing future-token leakage. - Muon optimizes hidden matrices; AdamW optimizes embeddings and normalization parameters. - BF16 training on one NVIDIA RTX PRO 6000 Blackwell, with batch size selected by a live memory/throughput probe. Default context is 1024 bytes. This is a research checkpoint, not an instruction model. ## Data mixture The stream is not redistributed in this repository: - 58% [FineWeb-Edu sample-10BT](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) - 22% [FineMath finemath-4plus](https://huggingface.co/datasets/HuggingFaceTB/finemath) - 10% [Cosmopedia v2](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus) - 7% [OpenR1-Math-220k](https://huggingface.co/datasets/open-r1/OpenR1-Math-220k) - 3% [decontaminated Codeforces reasoning traces](https://huggingface.co/datasets/open-r1/codeforces-cots) Review each upstream dataset card and license before redistribution or commercial use. ## Checkpoints `checkpoints/step-XXXXXXXX/model.safetensors` contains BF16 model weights. `training_state.json` reports live loss, throughput, token count, batch size, routing and patch statistics. ## Evidence and limits The precursor 998.7M-parameter prototype reached validation byte loss 1.235 versus 1.360 for a 988.1M dense baseline on the same TinyStories test. It used about 4.15x fewer estimated model FLOPs and ran 2.06x faster wall-clock. This short controlled result motivated full training; it is not evidence of 100x speedup, broad reasoning ability or production readiness. Training code and architecture are included for auditability. Results may change while training is active.