HuggingFaceFW/fineweb-edu
Viewer • Updated • 3.5B • 420k • 1.31k
BananaMind-Sundae is a compact pretrained language model built entirely from scratch using the LLaMA architecture.
The model was pretrained on approximately 2 billion tokens from the FineWeb-Edu dataset.
| Property | Value |
|---|---|
| Parameters | 20M |
| Context Length | 1024 tokens |
| Model Type | Base pretrained language model |
| Architecture | LLaMA |
| Training | Pretrained from scratch |
| Training Tokens | ~2B |
| Dataset | FineWeb-Edu |
Evaluated on BananaMind Base Bench 1.1, using 350 multiple-choice examples across 7 categories. Each item contains 4 candidate continuations and is scored using conditional mean log-probability.
| Metric | Result |
|---|---|
| Overall Elo | 896 |
| Random-chance Elo floor | 805 |
| Elo above random | +91 |
| Accuracy | 38.9% |
| 95% CI | [33.8%, 44.0%] |
| z vs. random | +5.99 |
| Parameters | 20,156,544 |
| Category | Accuracy | Elo |
|---|---|---|
| Language Completion | 82.0% | 1167 |
| World Knowledge | 44.0% | 874 |
| Context Tracking | 44.0% | 938 |
| Commonsense | 34.0% | 789 |
| Logical Reasoning | 28.0% | 920 |
| Quantitative | 22.0% | 815 |
| Code Completion | 18.0% | 831 |
| Difficulty | Accuracy |
|---|---|
| Easy | 49.6% |
| Medium | 34.2% |
| Hard | 32.8% |
The model performs significantly above the 25% random-choice baseline (z = +5.99).
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "BananaMind/BananaMind-Sundae"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)
prompt = "The future of artificial intelligence is"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(
**inputs,
max_new_tokens=100,
do_sample=True,
temperature=0.8,
top_p=0.95,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
BananaMind-Sundae is a base pretrained model, not an instruction-tuned or chat model.
Built by BananaMind Research Community