GGMLGuy's picture
Duplicate from BananaMind/BananaMind-2-Pro-Preview
2847929
|
Raw
History Blame Contribute Delete
12.7 kB
---
license: other
license_name: bananamind-community-license-1.0
license_link: LICENSE
language:
- en
library_name: transformers
pipeline_tag: text-generation
datasets:
- HuggingFaceFW/fineweb-edu
- mlfoundations/dclm-baseline-1.0
- HuggingFaceTB/smollm-corpus
- HuggingFaceTB/finemath
tags:
- causal-lm
- language-model
- base-model
- small-language-model
- bananamind
- bananamind2
- bananamind2-pro
- preview-checkpoint
- digit-tokenizer
- pytorch
- safetensors
- custom-code
- trust-remote-code
- custom-architecture
---
![BananaMind 2 Pro](banner.png)
# BananaMind-2-Pro-Preview
BananaMind-2-Pro-Preview is the first public checkpoint preview of BananaMind 2 Pro, a decoder-only base causal language model trained from scratch by BananaMind. This checkpoint was captured after **96,000 completed optimizer steps** and **51,904,512,000 training tokens** in an ongoing 100B-token pretraining run.
The model has **138,971,520 parameters**, a **3,072-token context window**, and a custom **32,768-token digit-aware byte-level BPE tokenizer**. It uses grouped-query attention, QK normalization, RoPE, SwiGLU, RMSNorm, tied input/output embeddings, and a KV cache for generation.
This is a base model, not an instruction-tuned or chat model. Use continuation-style prompts and load the repository with `trust_remote_code=True`.
![BananaMind 2 Pro Preview benchmark comparison](benchmarks.png)
## Preview Status
| Field | Value |
|---|---:|
| Release type | First public preview checkpoint |
| Checkpoint step | 95,999 |
| Optimizer steps completed | 96,000 |
| Tokens seen | 51,904,512,000 |
| Full-run target | 100B tokens |
| Training phase | Reasoning core |
| Training status | Ongoing |
Benchmark scores describe this exact 96K preview checkpoint. They should not be treated as final BananaMind 2 Pro results.
## Model Details
| Field | Value |
|---|---:|
| Parameters | 138,971,520 |
| Architecture | BananaMind2Pro decoder-only Transformer |
| Layers | 24 |
| Hidden size | 640 |
| Intermediate size | 1,920 |
| Attention heads | 8 |
| KV heads | 4 |
| Head dimension | 80 |
| Attention style | Grouped-query attention with QK norm |
| MLP | SwiGLU |
| Position embeddings | RoPE |
| RoPE theta | 100,000 |
| Normalization | RMSNorm |
| RMSNorm epsilon | 1e-6 |
| Vocabulary size | 32,768 |
| Context length | 3,072 |
| Embeddings | Tied input/output embeddings |
| Generation cache | KV cache supported |
| Weight format | safetensors |
| HF architecture | `BananaMind2ProForCausalLM` |
| HF model type | `bananamind2_pro` |
## Evaluation
The BananaMind 2 Pro scores below were measured on the exported 96K checkpoint. ARC Easy, ARC Challenge, PIQA, and HellaSwag use `acc_norm,none`; ArithMark 3 uses length-normalized continuation accuracy; ArithMark 2 uses raw continuation accuracy. INT Index uses the Open SLM Leaderboard-style aggregate. Code Only is the Base Bench 1.1 code-completion category Elo, while Base Bench 1.1 reports overall fixed-item Elo.
| Model | Parameters | ARC Easy | ARC Challenge | PIQA | HellaSwag | ArithMark 3 | ArithMark 2 | INT Index | Code Only | Base Bench 1.1 |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| **BananaMind-2-Pro-Preview 96K** | **139M** | **51.01%** | **27.13%** | **66.76%** | **39.83%** | **38.90%** | 28.60% | 23.04 | **1295** | **1106** |
| [GPT-X2-125M](https://huggingface.co/AxiomicLabs/GPT-X2-125M) | 125M | 51.47% | 27.82% | 67.30% | 40.41% | 37.20% | **30.68%** | 23.36 | 1078 | 1062 |
| [GPT-X-125M](https://huggingface.co/AxiomicLabs/GPT-X-125M) | 125M | 50.76% | 26.62% | 64.96% | 36.57% | 35.60% | 30.24% | 19.94 | 916 | 1013 |
| [SmolLM-135M](https://huggingface.co/HuggingFaceTB/SmolLM-135M) | 135M | **56.31%** | **29.01%** | **68.28%** | **42.70%** | 36.80% | 28.84% | **25.74** | **1585** | **1125** |
| [BananaMind-2-Medium](https://huggingface.co/BananaMind/BananaMind-2-Medium) | 49.6M | 43.81% | 25.34% | 61.86% | 32.43% | 36.20% | 28.20% | 15.37 | 1269 | 1034 |
| [GPT-2](https://huggingface.co/openai-community/gpt2) | 124M | 39.35% | 22.35% | 62.08% | 31.26% | 35.70% | 26.48% | N/A | 1052 | 996 |
| [Pythia-160M](https://huggingface.co/EleutherAI/pythia-160m) | 160M | 39.81% | 24.23% | 61.75% | 30.05% | N/A | N/A | N/A | N/A | N/A |
The preview row is bold for emphasis; the strongest reported score in each metric is also bold. Base Bench comparison values for GPT-X2-125M, GPT-X-125M, SmolLM-135M, BananaMind-2-Medium, and GPT-2 are taken from the BananaMind Base Bench leaderboard. Metrics without a supplied or leaderboard result are marked N/A.
**Code Only (Base Bench 1.1): 1295 Elo | 38/50 correct (76.00%) | 75.48% weighted accuracy**
### INT Index vs Training Compute
![INT Index versus estimated training compute](int_index_vs_compute.png)
Comparison-model training compute is estimated as `6 x parameters x training tokens`, matching the referenced GPT-X2 chart methodology. The Pro Preview point uses the supplied run estimate rather than recomputing it with the comparison approximation. GPT-2 and Pythia are excluded.
| Model | Training compute | INT Index |
|---|---:|---:|
| BananaMind-2-Pro-Preview | 72,669.44 PFLOPs | 23.04 |
| GPT-X2-125M | 56,286.75 PFLOPs | 23.36 |
| GPT-X-125M | 11,210.56 PFLOPs | 19.94 |
| SmolLM-135M | 484,254.03 PFLOPs | 25.74 |
| BananaMind-2-Medium | 14,867.33 PFLOPs | 15.37 |
### Base Bench Checkpoint Progression
![BananaMind 2 Pro Base Bench checkpoint progression](base_bench_progression.png)
The progression series is a consistent sweep over 24 exported checkpoints using CUDA, bfloat16, batch size 1, and the complete 350-item Base Bench 1.1 split. The 96K point in this sweep is 1105 Elo with 227/350 correct; the primary comparison and category tables use the separate CPU float32 result of 1106 Elo with the same 227/350 raw accuracy.
### Base Bench Category Results
| Category | Elo | Correct | Accuracy | Weighted accuracy |
|---|---:|---:|---:|---:|
| Language completion | 1570 | 50/50 | 100.00% | 100.00% |
| Commonsense | 1160 | 39/50 | 78.00% | 76.34% |
| World knowledge | 1142 | 39/50 | 78.00% | 74.29% |
| Context tracking | 897 | 19/50 | 38.00% | 36.31% |
| Quantitative | 967 | 19/50 | 38.00% | 38.80% |
| Logical reasoning | 1026 | 23/50 | 46.00% | 40.13% |
| Code Only (code completion) | 1295 | 38/50 | 76.00% | 75.48% |
| **Overall** | **1106** | **227/350** | **64.86%** | **61.44%** |
Evaluation results can vary with harness version, tokenizer handling, dtype, and scoring configuration. The published values are self-reported checkpoint evaluations.
## Tokenizer
BananaMind-2-Pro-Preview uses a custom 32,768-token byte-level BPE tokenizer trained on 75 GiB of representative FineWeb-Edu, DCLM, Cosmopedia-v2, FineMath-4+, and NPSet-2 Python educational data. It uses NFKC normalization and digit-aware pre-tokenization.
Digits are isolated before byte-level BPE so complete numbers are not merged into large number tokens.
| Token | ID |
|---|---:|
| `0` | 19 |
| `1` | 20 |
| `2` | 21 |
| `3` | 22 |
| `4` | 23 |
| `5` | 24 |
| `6` | 25 |
| `7` | 26 |
| `8` | 27 |
| `9` | 28 |
Special token IDs:
| Token | ID |
|---|---:|
| `<|pad|>` | 0 |
| `<|bos|>` | 1 |
| `<|eos|>` | 2 |
| `<|unk|>` | 3 |
## Training Data
The ongoing 100B-token curriculum combines educational web text, broad web text, synthetic textbook material, mathematics, and Python educational code. The table describes the full-run target allocation; this preview was exported after 51.904512B tokens.
| Dataset | Full-run target | Aggregate share |
|---|---:|---:|
| FineWeb-Edu | 50.166B | 50.17% |
| DCLM | 26.125B | 26.13% |
| Cosmopedia-v2 | 13.525B | 13.53% |
| FineMath-4+ | 7.875B | 7.88% |
| NPSet-2 Python Edu | 2.309B | 2.31% |
| **Total** | **100.000B** | **100.00%** |
The run uses a capacity-aware curriculum:
| Phase | Token range | Purpose |
|---|---:|---|
| Breadth foundation | 0B to 25B | Web-heavy language and knowledge foundation |
| Knowledge ramp | 25B to 40B | Gradual increase in synthetic, mathematics, and code data |
| Reasoning core | 40B to 75B | Sustained reasoning-oriented mixture |
| Synthesis ramp | 75B to 90B | Transition toward the finishing distribution |
| Quality finish | 90B to 100B | Final quality-focused mixture |
## Training Setup
| Field | Value |
|---|---:|
| Sequence length | 3,072 |
| Micro batch | 4 |
| Gradient accumulation | 44 |
| Effective batch | 176 sequences |
| Tokens per optimizer step | 540,672 |
| Preview optimizer steps | 96,000 |
| Planned optimizer steps | 184,954 |
| Scheduled training tokens | 99,999,449,088 |
| Optimizer | AdamW |
| Betas | 0.9, 0.95 |
| Peak learning rate | 1.5e-3 |
| Warmup steps | 2,000 |
| LR schedule | Warmup-stable-decay with cosine decay |
| Decay ratio | 0.15 |
| Weight decay | 0.1, then 0.01 after 40B tokens |
| Gradient clipping | 1.0 |
| Z-loss coefficient | 1e-4 until 40B tokens, then off |
| Compile | PyTorch compile enabled |
| Seed | 1337 |
## Energy and Carbon Estimate
The following is an engineering estimate for training through this 96K preview checkpoint, not a wall-meter measurement. Runtime is derived from 51,904,512,000 tokens at the observed run-average throughput of 52,438 tokens/s. **Only GPU and CPU package power were measured; all other component power, PSU loss, electricity-use, and emissions figures are estimates.**
| Item | Basis | Value |
|---|---|---:|
| Derived training time | 51.904512B tokens / 52,438 tokens/s | 274.95 hours (11.46 days) |
| GPU power | Measured: 12-second `nvidia-smi` average at 99-100% utilization | 262 W |
| CPU package power | Measured: two Intel RAPL samples of 42 W and 38 W | 40 W |
| MSI B760 motherboard, chipset, and VRM losses | Estimated | 25 W |
| 2x16 GiB Kingston DDR5-5600 memory | Estimated combined power | 8 W |
| Kingston NV3 NVMe SSD | Estimated | 3 W |
| Seagate 2 TB hard drive | Estimated idle/spinning | 4 W |
| Fans, controllers, and miscellaneous devices | Estimated | 10 W |
| Other components total | Estimated | 50 W |
| DC system load | Estimated | 352 W |
| PSU efficiency | Assumed | 90% |
| Wall power | Estimated | 391 W |
| Electricity use | Estimated | 108 kWh |
| Austrian grid intensity used | Recent daily estimate | 140 gCO2e/kWh |
| **Training emissions through 96K** | **Estimated** | **15.1 kg CO2e** |
The grid factor is a recent Austrian daily consumption-based estimate from [Electricity Maps](https://app.electricitymaps.com/map/zone/AT/3mo/daily). Applying its reported 2024 and 2025 flow-traced annual means of 125.5 and 169.4 gCO2e/kWh to the same energy estimate gives **13.5-18.2 kg CO2e**. This estimate excludes embodied hardware emissions, the display, and external networking or storage infrastructure.
## Usage
Install the runtime dependencies:
```bash
pip install -U torch transformers safetensors
```
Load the model with custom architecture code enabled:
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "BananaMind/BananaMind-2-Pro-Preview"
tokenizer = AutoTokenizer.from_pretrained(
model_id,
trust_remote_code=True,
)
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = (
torch.bfloat16
if torch.cuda.is_available() and torch.cuda.is_bf16_supported()
else torch.float32
)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
dtype=dtype,
).to(device).eval()
prompt = "The capital of France is"
inputs = tokenizer(prompt, return_tensors="pt").to(device)
with torch.no_grad():
output = model.generate(
**inputs,
max_new_tokens=80,
do_sample=True,
temperature=0.7,
top_p=0.9,
repetition_penalty=1.1,
pad_token_id=tokenizer.eos_token_id,
eos_token_id=tokenizer.eos_token_id,
use_cache=True,
)
print(tokenizer.decode(output[0], skip_special_tokens=True))
```
For deterministic continuation scoring, use `do_sample=False`. For free-form sampling, a temperature of `0.6` to `0.8`, `top_p=0.9`, and `repetition_penalty=1.1` are reasonable starting points.
## Intended Use
BananaMind-2-Pro-Preview is intended for base-model research, local experimentation, text continuation, tokenizer research, arithmetic evaluation, checkpoint analysis, and small-language-model comparisons.
It is not instruction-tuned and does not use a chat template. It has not received dedicated safety alignment and may produce incorrect, biased, repetitive, or otherwise undesirable text. Do not rely on its output for high-stakes decisions.
## License
This repository is released under theBananaMind Community License 1.0. Commercial products or services exceeding either threshold in Section 1 require a separate commercial license from Banaxi-Tech.