File size: 12,747 Bytes
2847929
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
---
license: other
license_name: bananamind-community-license-1.0
license_link: LICENSE
language:
- en
library_name: transformers
pipeline_tag: text-generation
datasets:
- HuggingFaceFW/fineweb-edu
- mlfoundations/dclm-baseline-1.0
- HuggingFaceTB/smollm-corpus
- HuggingFaceTB/finemath
tags:
- causal-lm
- language-model
- base-model
- small-language-model
- bananamind
- bananamind2
- bananamind2-pro
- preview-checkpoint
- digit-tokenizer
- pytorch
- safetensors
- custom-code
- trust-remote-code
- custom-architecture
---

![BananaMind 2 Pro](banner.png)

# BananaMind-2-Pro-Preview

BananaMind-2-Pro-Preview is the first public checkpoint preview of BananaMind 2 Pro, a decoder-only base causal language model trained from scratch by BananaMind. This checkpoint was captured after **96,000 completed optimizer steps** and **51,904,512,000 training tokens** in an ongoing 100B-token pretraining run.

The model has **138,971,520 parameters**, a **3,072-token context window**, and a custom **32,768-token digit-aware byte-level BPE tokenizer**. It uses grouped-query attention, QK normalization, RoPE, SwiGLU, RMSNorm, tied input/output embeddings, and a KV cache for generation.

This is a base model, not an instruction-tuned or chat model. Use continuation-style prompts and load the repository with `trust_remote_code=True`.

![BananaMind 2 Pro Preview benchmark comparison](benchmarks.png)

## Preview Status

| Field | Value |
|---|---:|
| Release type | First public preview checkpoint |
| Checkpoint step | 95,999 |
| Optimizer steps completed | 96,000 |
| Tokens seen | 51,904,512,000 |
| Full-run target | 100B tokens |
| Training phase | Reasoning core |
| Training status | Ongoing |

Benchmark scores describe this exact 96K preview checkpoint. They should not be treated as final BananaMind 2 Pro results.

## Model Details

| Field | Value |
|---|---:|
| Parameters | 138,971,520 |
| Architecture | BananaMind2Pro decoder-only Transformer |
| Layers | 24 |
| Hidden size | 640 |
| Intermediate size | 1,920 |
| Attention heads | 8 |
| KV heads | 4 |
| Head dimension | 80 |
| Attention style | Grouped-query attention with QK norm |
| MLP | SwiGLU |
| Position embeddings | RoPE |
| RoPE theta | 100,000 |
| Normalization | RMSNorm |
| RMSNorm epsilon | 1e-6 |
| Vocabulary size | 32,768 |
| Context length | 3,072 |
| Embeddings | Tied input/output embeddings |
| Generation cache | KV cache supported |
| Weight format | safetensors |
| HF architecture | `BananaMind2ProForCausalLM` |
| HF model type | `bananamind2_pro` |

## Evaluation

The BananaMind 2 Pro scores below were measured on the exported 96K checkpoint. ARC Easy, ARC Challenge, PIQA, and HellaSwag use `acc_norm,none`; ArithMark 3 uses length-normalized continuation accuracy; ArithMark 2 uses raw continuation accuracy. INT Index uses the Open SLM Leaderboard-style aggregate. Code Only is the Base Bench 1.1 code-completion category Elo, while Base Bench 1.1 reports overall fixed-item Elo.

| Model | Parameters | ARC Easy | ARC Challenge | PIQA | HellaSwag | ArithMark 3 | ArithMark 2 | INT Index | Code Only | Base Bench 1.1 |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| **BananaMind-2-Pro-Preview 96K** | **139M** | **51.01%** | **27.13%** | **66.76%** | **39.83%** | **38.90%** | 28.60% | 23.04 | **1295** | **1106** |
| [GPT-X2-125M](https://huggingface.co/AxiomicLabs/GPT-X2-125M) | 125M | 51.47% | 27.82% | 67.30% | 40.41% | 37.20% | **30.68%** | 23.36 | 1078 | 1062 |
| [GPT-X-125M](https://huggingface.co/AxiomicLabs/GPT-X-125M) | 125M | 50.76% | 26.62% | 64.96% | 36.57% | 35.60% | 30.24% | 19.94 | 916 | 1013 |
| [SmolLM-135M](https://huggingface.co/HuggingFaceTB/SmolLM-135M) | 135M | **56.31%** | **29.01%** | **68.28%** | **42.70%** | 36.80% | 28.84% | **25.74** | **1585** | **1125** |
| [BananaMind-2-Medium](https://huggingface.co/BananaMind/BananaMind-2-Medium) | 49.6M | 43.81% | 25.34% | 61.86% | 32.43% | 36.20% | 28.20% | 15.37 | 1269 | 1034 |
| [GPT-2](https://huggingface.co/openai-community/gpt2) | 124M | 39.35% | 22.35% | 62.08% | 31.26% | 35.70% | 26.48% | N/A | 1052 | 996 |
| [Pythia-160M](https://huggingface.co/EleutherAI/pythia-160m) | 160M | 39.81% | 24.23% | 61.75% | 30.05% | N/A | N/A | N/A | N/A | N/A |

The preview row is bold for emphasis; the strongest reported score in each metric is also bold. Base Bench comparison values for GPT-X2-125M, GPT-X-125M, SmolLM-135M, BananaMind-2-Medium, and GPT-2 are taken from the BananaMind Base Bench leaderboard. Metrics without a supplied or leaderboard result are marked N/A.

**Code Only (Base Bench 1.1): 1295 Elo | 38/50 correct (76.00%) | 75.48% weighted accuracy**

### INT Index vs Training Compute

![INT Index versus estimated training compute](int_index_vs_compute.png)

Comparison-model training compute is estimated as `6 x parameters x training tokens`, matching the referenced GPT-X2 chart methodology. The Pro Preview point uses the supplied run estimate rather than recomputing it with the comparison approximation. GPT-2 and Pythia are excluded.

| Model | Training compute | INT Index |
|---|---:|---:|
| BananaMind-2-Pro-Preview | 72,669.44 PFLOPs | 23.04 |
| GPT-X2-125M | 56,286.75 PFLOPs | 23.36 |
| GPT-X-125M | 11,210.56 PFLOPs | 19.94 |
| SmolLM-135M | 484,254.03 PFLOPs | 25.74 |
| BananaMind-2-Medium | 14,867.33 PFLOPs | 15.37 |

### Base Bench Checkpoint Progression

![BananaMind 2 Pro Base Bench checkpoint progression](base_bench_progression.png)

The progression series is a consistent sweep over 24 exported checkpoints using CUDA, bfloat16, batch size 1, and the complete 350-item Base Bench 1.1 split. The 96K point in this sweep is 1105 Elo with 227/350 correct; the primary comparison and category tables use the separate CPU float32 result of 1106 Elo with the same 227/350 raw accuracy.

### Base Bench Category Results

| Category | Elo | Correct | Accuracy | Weighted accuracy |
|---|---:|---:|---:|---:|
| Language completion | 1570 | 50/50 | 100.00% | 100.00% |
| Commonsense | 1160 | 39/50 | 78.00% | 76.34% |
| World knowledge | 1142 | 39/50 | 78.00% | 74.29% |
| Context tracking | 897 | 19/50 | 38.00% | 36.31% |
| Quantitative | 967 | 19/50 | 38.00% | 38.80% |
| Logical reasoning | 1026 | 23/50 | 46.00% | 40.13% |
| Code Only (code completion) | 1295 | 38/50 | 76.00% | 75.48% |
| **Overall** | **1106** | **227/350** | **64.86%** | **61.44%** |

Evaluation results can vary with harness version, tokenizer handling, dtype, and scoring configuration. The published values are self-reported checkpoint evaluations.

## Tokenizer

BananaMind-2-Pro-Preview uses a custom 32,768-token byte-level BPE tokenizer trained on 75 GiB of representative FineWeb-Edu, DCLM, Cosmopedia-v2, FineMath-4+, and NPSet-2 Python educational data. It uses NFKC normalization and digit-aware pre-tokenization.

Digits are isolated before byte-level BPE so complete numbers are not merged into large number tokens.

| Token | ID |
|---|---:|
| `0` | 19 |
| `1` | 20 |
| `2` | 21 |
| `3` | 22 |
| `4` | 23 |
| `5` | 24 |
| `6` | 25 |
| `7` | 26 |
| `8` | 27 |
| `9` | 28 |

Special token IDs:

| Token | ID |
|---|---:|
| `<|pad|>` | 0 |
| `<|bos|>` | 1 |
| `<|eos|>` | 2 |
| `<|unk|>` | 3 |

## Training Data

The ongoing 100B-token curriculum combines educational web text, broad web text, synthetic textbook material, mathematics, and Python educational code. The table describes the full-run target allocation; this preview was exported after 51.904512B tokens.

| Dataset | Full-run target | Aggregate share |
|---|---:|---:|
| FineWeb-Edu | 50.166B | 50.17% |
| DCLM | 26.125B | 26.13% |
| Cosmopedia-v2 | 13.525B | 13.53% |
| FineMath-4+ | 7.875B | 7.88% |
| NPSet-2 Python Edu | 2.309B | 2.31% |
| **Total** | **100.000B** | **100.00%** |

The run uses a capacity-aware curriculum:

| Phase | Token range | Purpose |
|---|---:|---|
| Breadth foundation | 0B to 25B | Web-heavy language and knowledge foundation |
| Knowledge ramp | 25B to 40B | Gradual increase in synthetic, mathematics, and code data |
| Reasoning core | 40B to 75B | Sustained reasoning-oriented mixture |
| Synthesis ramp | 75B to 90B | Transition toward the finishing distribution |
| Quality finish | 90B to 100B | Final quality-focused mixture |

## Training Setup

| Field | Value |
|---|---:|
| Sequence length | 3,072 |
| Micro batch | 4 |
| Gradient accumulation | 44 |
| Effective batch | 176 sequences |
| Tokens per optimizer step | 540,672 |
| Preview optimizer steps | 96,000 |
| Planned optimizer steps | 184,954 |
| Scheduled training tokens | 99,999,449,088 |
| Optimizer | AdamW |
| Betas | 0.9, 0.95 |
| Peak learning rate | 1.5e-3 |
| Warmup steps | 2,000 |
| LR schedule | Warmup-stable-decay with cosine decay |
| Decay ratio | 0.15 |
| Weight decay | 0.1, then 0.01 after 40B tokens |
| Gradient clipping | 1.0 |
| Z-loss coefficient | 1e-4 until 40B tokens, then off |
| Compile | PyTorch compile enabled |
| Seed | 1337 |

## Energy and Carbon Estimate

The following is an engineering estimate for training through this 96K preview checkpoint, not a wall-meter measurement. Runtime is derived from 51,904,512,000 tokens at the observed run-average throughput of 52,438 tokens/s. **Only GPU and CPU package power were measured; all other component power, PSU loss, electricity-use, and emissions figures are estimates.**

| Item | Basis | Value |
|---|---|---:|
| Derived training time | 51.904512B tokens / 52,438 tokens/s | 274.95 hours (11.46 days) |
| GPU power | Measured: 12-second `nvidia-smi` average at 99-100% utilization | 262 W |
| CPU package power | Measured: two Intel RAPL samples of 42 W and 38 W | 40 W |
| MSI B760 motherboard, chipset, and VRM losses | Estimated | 25 W |
| 2x16 GiB Kingston DDR5-5600 memory | Estimated combined power | 8 W |
| Kingston NV3 NVMe SSD | Estimated | 3 W |
| Seagate 2 TB hard drive | Estimated idle/spinning | 4 W |
| Fans, controllers, and miscellaneous devices | Estimated | 10 W |
| Other components total | Estimated | 50 W |
| DC system load | Estimated | 352 W |
| PSU efficiency | Assumed | 90% |
| Wall power | Estimated | 391 W |
| Electricity use | Estimated | 108 kWh |
| Austrian grid intensity used | Recent daily estimate | 140 gCO2e/kWh |
| **Training emissions through 96K** | **Estimated** | **15.1 kg CO2e** |

The grid factor is a recent Austrian daily consumption-based estimate from [Electricity Maps](https://app.electricitymaps.com/map/zone/AT/3mo/daily). Applying its reported 2024 and 2025 flow-traced annual means of 125.5 and 169.4 gCO2e/kWh to the same energy estimate gives **13.5-18.2 kg CO2e**. This estimate excludes embodied hardware emissions, the display, and external networking or storage infrastructure.

## Usage

Install the runtime dependencies:

```bash
pip install -U torch transformers safetensors
```

Load the model with custom architecture code enabled:

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "BananaMind/BananaMind-2-Pro-Preview"

tokenizer = AutoTokenizer.from_pretrained(
    model_id,
    trust_remote_code=True,
)

device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = (
    torch.bfloat16
    if torch.cuda.is_available() and torch.cuda.is_bf16_supported()
    else torch.float32
)

model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    dtype=dtype,
).to(device).eval()

prompt = "The capital of France is"
inputs = tokenizer(prompt, return_tensors="pt").to(device)

with torch.no_grad():
    output = model.generate(
        **inputs,
        max_new_tokens=80,
        do_sample=True,
        temperature=0.7,
        top_p=0.9,
        repetition_penalty=1.1,
        pad_token_id=tokenizer.eos_token_id,
        eos_token_id=tokenizer.eos_token_id,
        use_cache=True,
    )

print(tokenizer.decode(output[0], skip_special_tokens=True))
```

For deterministic continuation scoring, use `do_sample=False`. For free-form sampling, a temperature of `0.6` to `0.8`, `top_p=0.9`, and `repetition_penalty=1.1` are reasonable starting points.

## Intended Use

BananaMind-2-Pro-Preview is intended for base-model research, local experimentation, text continuation, tokenizer research, arithmetic evaluation, checkpoint analysis, and small-language-model comparisons.

It is not instruction-tuned and does not use a chat template. It has not received dedicated safety alignment and may produce incorrect, biased, repetitive, or otherwise undesirable text. Do not rely on its output for high-stakes decisions.

## License

This repository is released under theBananaMind Community License 1.0. Commercial products or services exceeding either threshold in Section 1 require a separate commercial license from Banaxi-Tech.