File size: 10,432 Bytes
2e9d390 fef393c dd25e82 fef393c dd25e82 2e9d390 fef393c dd25e82 fef393c dd25e82 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 | ---
library_name: custom
pipeline_tag: text-generation
license: other
license_name: cc-by-nc-4.0-weights-apache-2.0-code
license_link: LICENSE.md
tags:
- mamba2
- byte-level
- multimodal
- custom-code
- base_model
language:
- en
datasets:
- HuggingFaceM4/FineVision
- HuggingFaceFW/finepdfs
- HuggingFaceFW/fineweb-edu
- webshart/suno-various-94k
---
# Three-Level Nested Byte Mamba-2
This repository contains a research checkpoint for a **2.478B-parameter causal byte model** with three nested Mamba-2 resolutions. It predicts raw bytes rather than tokenizer IDs and was trained on a mixture of web/PDF text, serialized image-text examples, and serialized audio.
It is not a Transformers `AutoModel` checkpoint and is not instruction-formatted as a conventional chat model. Use the included cached inference script.

## Checkpoint contents
The published weights are sharded SafeTensors containing only the 980 model tensors. The original optimizer, scaler, training phase, data cursor, dataset paths, source fingerprints, and other training-only checkpoint objects were removed.
- Parameters: **2,478,820,575**
- Weight precision on disk: **FP32**
- Raw tensor size: **9,915,282,300 bytes**
- Source checkpoint step: **889,000**
- Recommended runtime precision: **BF16**
- Recommended placement: fine/decoder on `cuda:0`, level 2 on `cuda:1`, level 3 on `cuda:2`
- Last training count: 15GB **Absoloutly undertrained**
The source checkpoint step is documentation only; it is not embedded in the SafeTensors weights or inference configuration.
## Latest validation results
The latest recorded validation event is step **890,000**, one scheduled validation event after the packaged `last.pt` weight step.
| Validation stream | Cross entropy (nats/byte) | Bits per byte | Scored bytes |
|---|---:|---:|---:|
| Aggregate mixed validation | **3.828962** | **5.524025** | 14,530,840 |
| JSONL text | 0.937809 | 1.352973 | 1,246,101 |
| Parquet text | 0.820114 | 1.183175 | 1,929,612 |
| Image + text multimodal | 1.397195 | 2.015726 | 1,626,324 |
| Audio objectives | 5.189134 | 7.486338 | 9,824,803 |
The aggregate should not be interpreted as a pure language score: audio accounts for most evaluated bytes and has a substantially different entropy scale. For text use, the JSONL and Parquet rows are the relevant measurements.
## Architecture
### Byte vocabulary
There is no learned tokenizer:
```text
PAD=0, BOS=1, EOS=2, UNK=3
raw byte 0..255 -> ID 4..259
vocabulary size = 260
```
UTF-8 text and serialized binary modalities therefore share one next-byte objective.
### Three causal resolutions
1. **Fine level:** a local causal convolutional encoder and 6 Mamba-2 blocks operate at byte resolution. A learned causal boundary head closes variable pools between 1 and 96 bytes.
2. **Level 2:** 20 Mamba-2 blocks consume completed fine-pool states. A learned boundary head groups 4–16 completed fine pools.
3. **Level 3:** 30 Mamba-2 blocks consume completed level-2 states and group 2–16 level-2 pools.
Every Mamba block uses model width 2,000, Mamba-2 `d_state=64`, and head dimension 100. A pool can use only states already available in its causal prefix. A closure never revises an earlier prediction.
### Fusion decoder
For each byte, the decoder concatenates four 2,000-dimensional signals:
- byte-local contextual state;
- current fine latent;
- latest level-2 latent;
- latest level-3 latent.
The 10,000-dimensional concatenation is normalized, projected through an 8,000-wide GELU fusion layer, reduced to width 2,000, and mapped to 260 next-byte logits. This enlarged decoder was added to avoid choking the information arriving from three recurrent resolutions.
### Pool-density fallback
Fine pooling includes a rolling short-pool quota. Among the most recent 6,000 completed fine pools, at most 3,000 may be shorter than 6 bytes. When that quota fills, the next pool must reach the secondary minimum; short closures become eligible again as older short pools leave the rolling window. The quota counts completed pools, not raw bytes.
### Delayed decoder controller
The checkpoint includes an optional hold/refresh/compress controller. Its output at time `t` can influence closure only at `t+1`:
```text
decode byte t -> controller C[t] -> choose closure at t+1 -> decode byte t+1
```
The included cached inference path applies this without future leakage or a second full-model pass.
### Parameter distribution
| Component | Parameters |
|---|---:|
| Fine level, shared byte modules, decoder, and LM head | 411,954,571 |
| Level 2 pooler and 20 Mamba-2 blocks | 831,559,202 |
| Level 3 pooler and 30 Mamba-2 blocks | 1,235,306,802 |
Pooling reduces sequence activations and recurrent update frequency, not layer-weight storage. This is why the deepest level remains the largest parameter group even though it updates least frequently.
## Inference
### Dependencies
Use Linux, CUDA, and versions of PyTorch, `mamba-ssm`, Triton, and `causal-conv1d` that are mutually compatible:
```bash
pip install -r requirements.txt
```
BF16 is strongly recommended. FP16 cached rollouts can become numerically unstable on some Mamba-2 builds.
### Three-GPU inference
From the downloaded repository:
```bash
python infer_nested_model.py \
--checkpoint . \
--prompt "The history of state space models begins" \
--max-new-bytes 512 \
--precision bf16 \
--fine-device cuda:0 \
--nested-devices cuda:1 \
--tertiary-device cuda:2 \
--temperature 0.8 \
--top-p 0.9
```
The script accepts `--prompt-file` for arbitrary byte prefixes, `--output` for raw generated bytes, `--html-output` for hierarchy-attribution output, and `--image-output-dir` to extract complete generated P6 images.
Single-GPU inference is supported when the GPU can hold the requested precision:
```bash
python infer_nested_model.py --checkpoint . --device cuda:0 \
--prompt "Once upon a time" --max-new-bytes 256 --precision bf16
```
### Stateful generation
Generation prefills the prompt once, then caches the convolution and SSM states for the fine, level-2, and level-3 stacks. New bytes advance those caches token by token; the entire prefix is not reprocessed for every generated byte.
## Training mixture and modality representation
The training run mixed educational web text, PDF-derived text, image/question
and instruction examples, and paired music/cover data from the following
repositories:
| Training source | Use in this model | Upstream licensing and rights notice |
|---|---|---|
| [HuggingFaceM4/FineVision](https://huggingface.co/datasets/HuggingFaceM4/FineVision) | Image, document, question, and instruction examples | FineVision is an aggregation. Each constituent dataset retains its own license; rights in prompts contributed by FineVision are offered under CC BY 4.0. Consult the license metadata for the constituent subsets. |
| [HuggingFaceFW/finepdfs](https://huggingface.co/datasets/HuggingFaceFW/finepdfs) | PDF-derived document text | ODC-By 1.0; use is also subject to applicable Common Crawl terms and upstream-content rights. |
| [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) | Educational web text | ODC-By 1.0; source pages retain their applicable rights. |
| [webshart/suno-various-94k](https://huggingface.co/datasets/webshart/suno-various-94k) | Music, captions, lyrics, and generated cover pairs | Marked `source-rights-retained`. Rights in source audio and lyrics remain with their creators; the dataset does not grant rights over the underlying content. |
These datasets are not redistributed in this repository. Their upstream terms
continue to apply independently and are not replaced by this repository's
license.
- Text and code are UTF-8 bytes.
- Images are complete RGB PPM byte sequences plus associated text.
- Audio uses 24 kHz EnCodec payloads with generation and detection objectives.
- Instruction and dialogue fields present in source records were serialized in full rather than using assistant-response-only loss.
This mixture makes the checkpoint experimental and general-purpose at the byte level; it does not guarantee strong image or audio generation quality.
## Limitations
- This is custom research code, not an official Mamba or Transformers architecture.
- The model is not a safety-aligned chat assistant.
- Raw-byte sampling can produce invalid UTF-8, malformed images, or incomplete audio containers.
- Image training used small PPM rasters, limiting fine visual detail.
- Audio validation remains much weaker than text validation.
- The audio corpus includes third-party creator material whose source rights are retained. The model license does not grant rights to reproduce protected training content, lyrics, compositions, voices, or recordings.
- The current weights are FP32 and large; practical use generally requires BF16 casting.
- The delayed pooling controller is causal but makes exact routing inherently sequential.
- The latest CSV validation event is at step 890,000, while the packaged `last.pt` weights identify step 889,000; the table must therefore be read as the latest run validation, not an evaluation re-run performed directly on this exported artifact.
## Intended use
Intended for research into byte-level modeling, hierarchical state-space models, adaptive causal pooling, long recurrent context, and mixed text/binary generation. Validate outputs independently before using them in downstream systems.
## License
The model weights, model card, and visual assets are available under
[CC BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/). The Python
inference source is available under
[Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0). Training
datasets and third-party content are not covered by either grant. See
[LICENSE.md](LICENSE.md) for the precise repository scope and notices.
## Repository files
- `model-*.safetensors`: inference-only model shards
- `model.safetensors.index.json`: tensor-to-shard map
- `config.json`: architecture-only inference configuration
- `modeling_nested_mamba.py`: custom model implementation
- `nested_inference_tools.py`: SafeTensors loading, state caching, and sampling
- `infer_nested_model.py`: command-line generator
- `architecture.svg`: architecture visualization
- `LICENSE.md`: weight, documentation, and code license scope |