BananaMind 2 Pro

BananaMind-2-Pro-Preview

BananaMind-2-Pro-Preview is the first public checkpoint preview of BananaMind 2 Pro, a decoder-only base causal language model trained from scratch by BananaMind. This checkpoint was captured after 96,000 completed optimizer steps and 51,904,512,000 training tokens in an ongoing 100B-token pretraining run.

The model has 138,971,520 parameters, a 3,072-token context window, and a custom 32,768-token digit-aware byte-level BPE tokenizer. It uses grouped-query attention, QK normalization, RoPE, SwiGLU, RMSNorm, tied input/output embeddings, and a KV cache for generation.

This is a base model, not an instruction-tuned or chat model. Use continuation-style prompts and load the repository with trust_remote_code=True.

BananaMind 2 Pro Preview benchmark comparison

Preview Status

Field Value
Release type First public preview checkpoint
Checkpoint step 95,999
Optimizer steps completed 96,000
Tokens seen 51,904,512,000
Full-run target 100B tokens
Training phase Reasoning core
Training status Ongoing

Benchmark scores describe this exact 96K preview checkpoint. They should not be treated as final BananaMind 2 Pro results.

Model Details

Field Value
Parameters 138,971,520
Architecture BananaMind2Pro decoder-only Transformer
Layers 24
Hidden size 640
Intermediate size 1,920
Attention heads 8
KV heads 4
Head dimension 80
Attention style Grouped-query attention with QK norm
MLP SwiGLU
Position embeddings RoPE
RoPE theta 100,000
Normalization RMSNorm
RMSNorm epsilon 1e-6
Vocabulary size 32,768
Context length 3,072
Embeddings Tied input/output embeddings
Generation cache KV cache supported
Weight format safetensors
HF architecture BananaMind2ProForCausalLM
HF model type bananamind2_pro

Evaluation

The BananaMind 2 Pro scores below were measured on the exported 96K checkpoint. ARC Easy, ARC Challenge, PIQA, and HellaSwag use acc_norm,none; ArithMark 3 uses length-normalized continuation accuracy; ArithMark 2 uses raw continuation accuracy. INT Index uses the Open SLM Leaderboard-style aggregate. Code Only is the Base Bench 1.1 code-completion category Elo, while Base Bench 1.1 reports overall fixed-item Elo.

Model Parameters ARC Easy ARC Challenge PIQA HellaSwag ArithMark 3 ArithMark 2 INT Index Code Only Base Bench 1.1
BananaMind-2-Pro-Preview 96K 139M 51.01% 27.13% 66.76% 39.83% 38.90% 28.60% 23.04 1295 1106
GPT-X2-125M 125M 51.47% 27.82% 67.30% 40.41% 37.20% 30.68% 23.36 1078 1062
GPT-X-125M 125M 50.76% 26.62% 64.96% 36.57% 35.60% 30.24% 19.94 916 1013
SmolLM-135M 135M 56.31% 29.01% 68.28% 42.70% 36.80% 28.84% 25.74 1585 1125
BananaMind-2-Medium 49.6M 43.81% 25.34% 61.86% 32.43% 36.20% 28.20% 15.37 1269 1034
GPT-2 124M 39.35% 22.35% 62.08% 31.26% 35.70% 26.48% N/A 1052 996
Pythia-160M 160M 39.81% 24.23% 61.75% 30.05% N/A N/A N/A N/A N/A

The preview row is bold for emphasis; the strongest reported score in each metric is also bold. Base Bench comparison values for GPT-X2-125M, GPT-X-125M, SmolLM-135M, BananaMind-2-Medium, and GPT-2 are taken from the BananaMind Base Bench leaderboard. Metrics without a supplied or leaderboard result are marked N/A.

Code Only (Base Bench 1.1): 1295 Elo | 38/50 correct (76.00%) | 75.48% weighted accuracy

INT Index vs Training Compute

INT Index versus estimated training compute

Comparison-model training compute is estimated as 6 x parameters x training tokens, matching the referenced GPT-X2 chart methodology. The Pro Preview point uses the supplied run estimate rather than recomputing it with the comparison approximation. GPT-2 and Pythia are excluded.

Model Training compute INT Index
BananaMind-2-Pro-Preview 72,669.44 PFLOPs 23.04
GPT-X2-125M 56,286.75 PFLOPs 23.36
GPT-X-125M 11,210.56 PFLOPs 19.94
SmolLM-135M 484,254.03 PFLOPs 25.74
BananaMind-2-Medium 14,867.33 PFLOPs 15.37

Base Bench Checkpoint Progression

BananaMind 2 Pro Base Bench checkpoint progression

The progression series is a consistent sweep over 24 exported checkpoints using CUDA, bfloat16, batch size 1, and the complete 350-item Base Bench 1.1 split. The 96K point in this sweep is 1105 Elo with 227/350 correct; the primary comparison and category tables use the separate CPU float32 result of 1106 Elo with the same 227/350 raw accuracy.

Base Bench Category Results

Category Elo Correct Accuracy Weighted accuracy
Language completion 1570 50/50 100.00% 100.00%
Commonsense 1160 39/50 78.00% 76.34%
World knowledge 1142 39/50 78.00% 74.29%
Context tracking 897 19/50 38.00% 36.31%
Quantitative 967 19/50 38.00% 38.80%
Logical reasoning 1026 23/50 46.00% 40.13%
Code Only (code completion) 1295 38/50 76.00% 75.48%
Overall 1106 227/350 64.86% 61.44%

Evaluation results can vary with harness version, tokenizer handling, dtype, and scoring configuration. The published values are self-reported checkpoint evaluations.

Tokenizer

BananaMind-2-Pro-Preview uses a custom 32,768-token byte-level BPE tokenizer trained on 75 GiB of representative FineWeb-Edu, DCLM, Cosmopedia-v2, FineMath-4+, and NPSet-2 Python educational data. It uses NFKC normalization and digit-aware pre-tokenization.

Digits are isolated before byte-level BPE so complete numbers are not merged into large number tokens.

Token ID
0 19
1 20
2 21
3 22
4 23
5 24
6 25
7 26
8 27
9 28

Special token IDs:

Token ID
`< pad
`< bos
`< eos
`< unk

Training Data

The ongoing 100B-token curriculum combines educational web text, broad web text, synthetic textbook material, mathematics, and Python educational code. The table describes the full-run target allocation; this preview was exported after 51.904512B tokens.

Dataset Full-run target Aggregate share
FineWeb-Edu 50.166B 50.17%
DCLM 26.125B 26.13%
Cosmopedia-v2 13.525B 13.53%
FineMath-4+ 7.875B 7.88%
NPSet-2 Python Edu 2.309B 2.31%
Total 100.000B 100.00%

The run uses a capacity-aware curriculum:

Phase Token range Purpose
Breadth foundation 0B to 25B Web-heavy language and knowledge foundation
Knowledge ramp 25B to 40B Gradual increase in synthetic, mathematics, and code data
Reasoning core 40B to 75B Sustained reasoning-oriented mixture
Synthesis ramp 75B to 90B Transition toward the finishing distribution
Quality finish 90B to 100B Final quality-focused mixture

Training Setup

Field Value
Sequence length 3,072
Micro batch 4
Gradient accumulation 44
Effective batch 176 sequences
Tokens per optimizer step 540,672
Preview optimizer steps 96,000
Planned optimizer steps 184,954
Scheduled training tokens 99,999,449,088
Optimizer AdamW
Betas 0.9, 0.95
Peak learning rate 1.5e-3
Warmup steps 2,000
LR schedule Warmup-stable-decay with cosine decay
Decay ratio 0.15
Weight decay 0.1, then 0.01 after 40B tokens
Gradient clipping 1.0
Z-loss coefficient 1e-4 until 40B tokens, then off
Compile PyTorch compile enabled
Seed 1337

Energy and Carbon Estimate

The following is an engineering estimate for training through this 96K preview checkpoint, not a wall-meter measurement. Runtime is derived from 51,904,512,000 tokens at the observed run-average throughput of 52,438 tokens/s. Only GPU and CPU package power were measured; all other component power, PSU loss, electricity-use, and emissions figures are estimates.

Item Basis Value
Derived training time 51.904512B tokens / 52,438 tokens/s 274.95 hours (11.46 days)
GPU power Measured: 12-second nvidia-smi average at 99-100% utilization 262 W
CPU package power Measured: two Intel RAPL samples of 42 W and 38 W 40 W
MSI B760 motherboard, chipset, and VRM losses Estimated 25 W
2x16 GiB Kingston DDR5-5600 memory Estimated combined power 8 W
Kingston NV3 NVMe SSD Estimated 3 W
Seagate 2 TB hard drive Estimated idle/spinning 4 W
Fans, controllers, and miscellaneous devices Estimated 10 W
Other components total Estimated 50 W
DC system load Estimated 352 W
PSU efficiency Assumed 90%
Wall power Estimated 391 W
Electricity use Estimated 108 kWh
Austrian grid intensity used Recent daily estimate 140 gCO2e/kWh
Training emissions through 96K Estimated 15.1 kg CO2e

The grid factor is a recent Austrian daily consumption-based estimate from Electricity Maps. Applying its reported 2024 and 2025 flow-traced annual means of 125.5 and 169.4 gCO2e/kWh to the same energy estimate gives 13.5-18.2 kg CO2e. This estimate excludes embodied hardware emissions, the display, and external networking or storage infrastructure.

Usage

Install the runtime dependencies:

pip install -U torch transformers safetensors

Load the model with custom architecture code enabled:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "BananaMind/BananaMind-2-Pro-Preview"

tokenizer = AutoTokenizer.from_pretrained(
    model_id,
    trust_remote_code=True,
)

device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = (
    torch.bfloat16
    if torch.cuda.is_available() and torch.cuda.is_bf16_supported()
    else torch.float32
)

model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    dtype=dtype,
).to(device).eval()

prompt = "The capital of France is"
inputs = tokenizer(prompt, return_tensors="pt").to(device)

with torch.no_grad():
    output = model.generate(
        **inputs,
        max_new_tokens=80,
        do_sample=True,
        temperature=0.7,
        top_p=0.9,
        repetition_penalty=1.1,
        pad_token_id=tokenizer.eos_token_id,
        eos_token_id=tokenizer.eos_token_id,
        use_cache=True,
    )

print(tokenizer.decode(output[0], skip_special_tokens=True))

For deterministic continuation scoring, use do_sample=False. For free-form sampling, a temperature of 0.6 to 0.8, top_p=0.9, and repetition_penalty=1.1 are reasonable starting points.

Intended Use

BananaMind-2-Pro-Preview is intended for base-model research, local experimentation, text continuation, tokenizer research, arithmetic evaluation, checkpoint analysis, and small-language-model comparisons.

It is not instruction-tuned and does not use a chat template. It has not received dedicated safety alignment and may produce incorrect, biased, repetitive, or otherwise undesirable text. Do not rely on its output for high-stakes decisions.

License

This repository is released under theBananaMind Community License 1.0. Commercial products or services exceeding either threshold in Section 1 require a separate commercial license from Banaxi-Tech.

Downloads last month
-
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 2 Ask for provider support

Model tree for BananaMind/BananaMind-2-Pro-Preview

Finetunes
1 model

Datasets used to train BananaMind/BananaMind-2-Pro-Preview