Our preview of BananaMind 2 Pro will release on August 3. Before we do that we want to hit a goal Lets get 150 followers on my account and 75 on BananaMind! Me: @Banaxi-Tech BananaMind:
BananaMind Full Release on August 10-14 Would really appreciate it!
A Small Model is All You Need. Meet palmer-006 (90M)
After 3 years of experiments, we are finally releasing our flagship tiny model: **palmer-006**.
If you are building for edge hardware, SBCs (Raspberry Pi, etc.), or low-power devices, this is for you. Inspired by Andrej Karpathy's idea of a self-contained "cognitive core," we wanted to see how much power we could pack into a sub-100M parameter footprint.
๐ง **How we "Palmerized" it:** We believe in starting our experiments with the absolute strongest baseline possible. 1. Light fine-tuning on highly curated data 2. Model merging 3. Another light fine-tuning round 4. Adjusted Mamba for maximum token speed โก๏ธ
โ ๏ธ *Note: This is a foundational language model. It has not been instruction-tuned yet!*
Also, since this needs instruction tuning next to become a chat assistantโ**what dataset would you recommend we use for the instruct tune?**
--- ๐ **Quick Links & Info:**
* **License:** Open for research, education, hobby, and modification! (For commercial use/hosted APIs, shoot an email to nosoyhackercodigo@gmail.com. *PS: Donators can claim a free commercial license!*)
* **Attribution:** Built using AI tech from the Technology Innovation Institute (TII).
Can't wait to see what you build at the edge. Let me know your prompt completions below! ๐
BananaMind 2 Pro is training! The current checkpoint (ONLY 20% DONE) GETS #6 On the entire Open SLM Leaderboard. We are going to release the first public preview on August 2-4 (estimated from speed)
We're excited to release BananaMindBench Leaderboard, our leaderboard for BananaMind Base Bench 1.1. It measures model performance on a variety of different tasks: Language Completion Common sense too World Knowledge Context Tracking Quantitative Logical Reasoning Code Completion Each has a different score and 1 overall score. Submit your own model: BananaMind/BananaMindBench-Leaderboard Check it out: BananaMind/BananaMindBench-Leaderboard
We're excited to release BananaMind Base Bench 1.1 A new benchmark for base language models with 350 text-completion examples across seven categories. Models are scored using continuation likelihood and receive an Overall Elo score. Initial results: BananaMind-2-Medium: 1034 BananaMind-2-Mini: 974 Supra-50M-Base: 973 Supra-1.5-50M-Base-exp: 948 BananaMind-2-Nano: 910 The official script downloads the gated dataset directly from Hugging Face. The dataset is for benchmarking only and may not be used for model training. BananaMind/BananaMind-Base-Bench-1.1
We're excited to announce BananaMind 2V, our small vision model series! These models are NOT released yet. We will release them in mid-august! BananaMind 2V will include: BananaMind 2V 256M, the flagship based on BananaMind 2 Pro (BananaMind 2 Pro is not released yet). BananaMind 2V 100M, our mid model, based on BananaMind 2 Medium. BananaMind 2V 50M, our smallest vision model, based on BananaMind 2 Mini. These are currently unreleased and will release in mid-august. Our training will start after BananaMind 2 Pro has finished training.
We're excited to release BananaMind 2 Medium and BananaMind 2 Medium Chat!
Theyโre both 50M parameter models trained on 50B tokens from FineWeb-Edu, DCLM, Cosmopedia v2, FineMath-4+ and NPSet-2 Python-Edu.
The base model reached 61.86% on PIQA, 43.81% on ARC Easy and 32.43% on HellaSwag. The Chat version was fine-tuned on Smol-SmolTalk and scored 38% overall on our internal instruction benchmark, with 56% on multi-turn, 60% on context recall and 80% on code.
Lets get 80 followers on my account and 30 on BananaMind. When we hit that we are going to release BananaMind 2 Medium tomorrow. Me: @Banaxi-Tech BananaMind:
If you make cool smol ๐ค๐ค models (below 0.5b parameters), leave a reply and I will follow you! I'm serious, you don't need to follow me at all just share something through the replies and I (and potentially more people) will follow you (if your models are decent ofc).
if you are a tinkerer of small language models and want to stay ahead of what small models can do, follow me!!! seriously, start following people that actually still makes small models
i've made one recently btw
also, i'm keeping an eye on AxiomicLabs leaderboard, looks like the only current alternative to check where the things are going to
though, between us, i think they should add agentic/tool use benchmarks there
A huge amount of large synthetic datasets on huggingface looks surprisingly like templates, that might be one of the main reasons open models might not be as good as other models, we need more people to create smaller, human-curated datasets instead of lazily sending millions of requests to large models for us to fulfill.
BananaMind 2 Nano is the smallest member of the BananaMind 2.0 family โ a 10M-parameter language model that shows how much you can squeeze out of a tiny footprint. It uses the family's digit-isolated tokenizer, so it keeps solid arithmetic despite its size, and it's small enough to run just about anywhere.
Trained on 30B tokens in about a day on a single RTX 5070 Ti (16GB), 4096-token context.
We're excited to release BananaMind 2 MoE, a new addition to the BananaMind 2 series! BananaMind 2 MoE is a sparse mixture-of-experts model with 25M total parameters but only 2M active per token, Like the rest of the series, it uses our custom digit-aware BPE tokenizer that keeps every digit isolated, fixing the core arithmetic weakness of our earlier models. It's trained on 30B tokens from FineWeb-Edu, DCLM, Cosmopedia-v2 and FineMath-4+, and outperforms Pythia-31M on average despite activating just 2M parameters. Check it out at BananaMind/BananaMind-2-MoE
BananaMind 2 Nano is Coming Next. Already training. Apache 2.0.
We are announcing 3 more models in our BananaMind 2 Family of models! BananaMind 2 Nano, a small 10M parameter model, fits on your Pentium 4 BananaMind 2 Medium, our medium model, 50M parameters BananaMind 2 MoE, 25M parameters, 2M active per tokens as fast as a 2M.
Because of this our release dates have changed a bit our currently estimates are: BananaMind 2 MoE July 16-18 BananaMind 2 Nano July 18-20 BananaMind 2 Medium July 24-28 BananaMind 2 Pro August 10-16 Keep in mind these dates are estimates and we don't have a speed number currently, we will post for details going forward!
small reasoning models are overrated, these little ones just doom loop a lot by default. good data will always be the moat when training or finetuning small models and latest sota models like fable 5 and gpt 5.6 are increasingly making this a lot easier to do.
We're excited to release BananaMind 2 Mini the first model in our BananaMind 2 series!
BananaMind 2 Mini features a custom digit-aware BPE tokenizer that keeps every digit isolated, fixing the core arithmetic weakness of our previous models. It's trained on 30B tokens from FineWeb-Edu, DCLM, Cosmopedia-v2 and FineMath-4+, and already outperforms Pythia-31M despite having fewer parameters. Check it out at BananaMind/BananaMind-2-Mini