File size: 10,432 Bytes
2e9d390
fef393c
 
 
 
 
 
dd25e82
 
 
 
 
fef393c
dd25e82
 
 
 
 
 
2e9d390
fef393c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
dd25e82
fef393c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
dd25e82
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
---
library_name: custom
pipeline_tag: text-generation
license: other
license_name: cc-by-nc-4.0-weights-apache-2.0-code
license_link: LICENSE.md
tags:
- mamba2
- byte-level
- multimodal
- custom-code
- base_model
language:
- en
datasets:
- HuggingFaceM4/FineVision
- HuggingFaceFW/finepdfs
- HuggingFaceFW/fineweb-edu
- webshart/suno-various-94k
---

# Three-Level Nested Byte Mamba-2

This repository contains a research checkpoint for a **2.478B-parameter causal byte model** with three nested Mamba-2 resolutions. It predicts raw bytes rather than tokenizer IDs and was trained on a mixture of web/PDF text, serialized image-text examples, and serialized audio.

It is not a Transformers `AutoModel` checkpoint and is not instruction-formatted as a conventional chat model. Use the included cached inference script.

![Three-level nested Mamba-2 architecture](architecture.svg)

## Checkpoint contents

The published weights are sharded SafeTensors containing only the 980 model tensors. The original optimizer, scaler, training phase, data cursor, dataset paths, source fingerprints, and other training-only checkpoint objects were removed.

- Parameters: **2,478,820,575**
- Weight precision on disk: **FP32**
- Raw tensor size: **9,915,282,300 bytes**
- Source checkpoint step: **889,000**
- Recommended runtime precision: **BF16**
- Recommended placement: fine/decoder on `cuda:0`, level 2 on `cuda:1`, level 3 on `cuda:2`
- Last training count: 15GB **Absoloutly undertrained**

The source checkpoint step is documentation only; it is not embedded in the SafeTensors weights or inference configuration.

## Latest validation results

The latest recorded validation event is step **890,000**, one scheduled validation event after the packaged `last.pt` weight step.

| Validation stream | Cross entropy (nats/byte) | Bits per byte | Scored bytes |
|---|---:|---:|---:|
| Aggregate mixed validation | **3.828962** | **5.524025** | 14,530,840 |
| JSONL text | 0.937809 | 1.352973 | 1,246,101 |
| Parquet text | 0.820114 | 1.183175 | 1,929,612 |
| Image + text multimodal | 1.397195 | 2.015726 | 1,626,324 |
| Audio objectives | 5.189134 | 7.486338 | 9,824,803 |

The aggregate should not be interpreted as a pure language score: audio accounts for most evaluated bytes and has a substantially different entropy scale. For text use, the JSONL and Parquet rows are the relevant measurements.

## Architecture

### Byte vocabulary

There is no learned tokenizer:

```text
PAD=0, BOS=1, EOS=2, UNK=3
raw byte 0..255 -> ID 4..259
vocabulary size = 260
```

UTF-8 text and serialized binary modalities therefore share one next-byte objective.

### Three causal resolutions

1. **Fine level:** a local causal convolutional encoder and 6 Mamba-2 blocks operate at byte resolution. A learned causal boundary head closes variable pools between 1 and 96 bytes.
2. **Level 2:** 20 Mamba-2 blocks consume completed fine-pool states. A learned boundary head groups 4–16 completed fine pools.
3. **Level 3:** 30 Mamba-2 blocks consume completed level-2 states and group 2–16 level-2 pools.

Every Mamba block uses model width 2,000, Mamba-2 `d_state=64`, and head dimension 100. A pool can use only states already available in its causal prefix. A closure never revises an earlier prediction.

### Fusion decoder

For each byte, the decoder concatenates four 2,000-dimensional signals:

- byte-local contextual state;
- current fine latent;
- latest level-2 latent;
- latest level-3 latent.

The 10,000-dimensional concatenation is normalized, projected through an 8,000-wide GELU fusion layer, reduced to width 2,000, and mapped to 260 next-byte logits. This enlarged decoder was added to avoid choking the information arriving from three recurrent resolutions.

### Pool-density fallback

Fine pooling includes a rolling short-pool quota. Among the most recent 6,000 completed fine pools, at most 3,000 may be shorter than 6 bytes. When that quota fills, the next pool must reach the secondary minimum; short closures become eligible again as older short pools leave the rolling window. The quota counts completed pools, not raw bytes.

### Delayed decoder controller

The checkpoint includes an optional hold/refresh/compress controller. Its output at time `t` can influence closure only at `t+1`:

```text
decode byte t -> controller C[t] -> choose closure at t+1 -> decode byte t+1
```

The included cached inference path applies this without future leakage or a second full-model pass.

### Parameter distribution

| Component | Parameters |
|---|---:|
| Fine level, shared byte modules, decoder, and LM head | 411,954,571 |
| Level 2 pooler and 20 Mamba-2 blocks | 831,559,202 |
| Level 3 pooler and 30 Mamba-2 blocks | 1,235,306,802 |

Pooling reduces sequence activations and recurrent update frequency, not layer-weight storage. This is why the deepest level remains the largest parameter group even though it updates least frequently.

## Inference

### Dependencies

Use Linux, CUDA, and versions of PyTorch, `mamba-ssm`, Triton, and `causal-conv1d` that are mutually compatible:

```bash
pip install -r requirements.txt
```

BF16 is strongly recommended. FP16 cached rollouts can become numerically unstable on some Mamba-2 builds.

### Three-GPU inference

From the downloaded repository:

```bash
python infer_nested_model.py \
  --checkpoint . \
  --prompt "The history of state space models begins" \
  --max-new-bytes 512 \
  --precision bf16 \
  --fine-device cuda:0 \
  --nested-devices cuda:1 \
  --tertiary-device cuda:2 \
  --temperature 0.8 \
  --top-p 0.9
```

The script accepts `--prompt-file` for arbitrary byte prefixes, `--output` for raw generated bytes, `--html-output` for hierarchy-attribution output, and `--image-output-dir` to extract complete generated P6 images.

Single-GPU inference is supported when the GPU can hold the requested precision:

```bash
python infer_nested_model.py --checkpoint . --device cuda:0 \
  --prompt "Once upon a time" --max-new-bytes 256 --precision bf16
```

### Stateful generation

Generation prefills the prompt once, then caches the convolution and SSM states for the fine, level-2, and level-3 stacks. New bytes advance those caches token by token; the entire prefix is not reprocessed for every generated byte.

## Training mixture and modality representation

The training run mixed educational web text, PDF-derived text, image/question
and instruction examples, and paired music/cover data from the following
repositories:

| Training source | Use in this model | Upstream licensing and rights notice |
|---|---|---|
| [HuggingFaceM4/FineVision](https://huggingface.co/datasets/HuggingFaceM4/FineVision) | Image, document, question, and instruction examples | FineVision is an aggregation. Each constituent dataset retains its own license; rights in prompts contributed by FineVision are offered under CC BY 4.0. Consult the license metadata for the constituent subsets. |
| [HuggingFaceFW/finepdfs](https://huggingface.co/datasets/HuggingFaceFW/finepdfs) | PDF-derived document text | ODC-By 1.0; use is also subject to applicable Common Crawl terms and upstream-content rights. |
| [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) | Educational web text | ODC-By 1.0; source pages retain their applicable rights. |
| [webshart/suno-various-94k](https://huggingface.co/datasets/webshart/suno-various-94k) | Music, captions, lyrics, and generated cover pairs | Marked `source-rights-retained`. Rights in source audio and lyrics remain with their creators; the dataset does not grant rights over the underlying content. |

These datasets are not redistributed in this repository. Their upstream terms
continue to apply independently and are not replaced by this repository's
license.

- Text and code are UTF-8 bytes.
- Images are complete RGB PPM byte sequences plus associated text.
- Audio uses 24 kHz EnCodec payloads with generation and detection objectives.
- Instruction and dialogue fields present in source records were serialized in full rather than using assistant-response-only loss.

This mixture makes the checkpoint experimental and general-purpose at the byte level; it does not guarantee strong image or audio generation quality.

## Limitations

- This is custom research code, not an official Mamba or Transformers architecture.
- The model is not a safety-aligned chat assistant.
- Raw-byte sampling can produce invalid UTF-8, malformed images, or incomplete audio containers.
- Image training used small PPM rasters, limiting fine visual detail.
- Audio validation remains much weaker than text validation.
- The audio corpus includes third-party creator material whose source rights are retained. The model license does not grant rights to reproduce protected training content, lyrics, compositions, voices, or recordings.
- The current weights are FP32 and large; practical use generally requires BF16 casting.
- The delayed pooling controller is causal but makes exact routing inherently sequential.
- The latest CSV validation event is at step 890,000, while the packaged `last.pt` weights identify step 889,000; the table must therefore be read as the latest run validation, not an evaluation re-run performed directly on this exported artifact.

## Intended use

Intended for research into byte-level modeling, hierarchical state-space models, adaptive causal pooling, long recurrent context, and mixed text/binary generation. Validate outputs independently before using them in downstream systems.

## License

The model weights, model card, and visual assets are available under
[CC BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/). The Python
inference source is available under
[Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0). Training
datasets and third-party content are not covered by either grant. See
[LICENSE.md](LICENSE.md) for the precise repository scope and notices.

## Repository files

- `model-*.safetensors`: inference-only model shards
- `model.safetensors.index.json`: tensor-to-shard map
- `config.json`: architecture-only inference configuration
- `modeling_nested_mamba.py`: custom model implementation
- `nested_inference_tools.py`: SafeTensors loading, state caching, and sampling
- `infer_nested_model.py`: command-line generator
- `architecture.svg`: architecture visualization
- `LICENSE.md`: weight, documentation, and code license scope