Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
๐
Is AI conscious?
65.0
TFLOPS
Banaxi
PRO
Banaxi-Tech
173
16
144
Follow
feli8023's profile picture
SeaWolf-AI's profile picture
dipankarsarkar's profile picture
193 followers
ยท
75 following
Banaxi-Tech
AI & ML interests
SLMs, training from scratch, LoRA, TTS, Ternary models. AI Interpretability. BCI. Contact at banaxitech@gmail.com
Recent Activity
reacted
to
their
post
with ๐ฅ
about 1 hour ago
We're announcing our BananaMind 2.1 model series! The models will include: - BananaMind 2.1 Nano: 10M parameters with 60B tokens. - BananaMind 2.1 Lite: 25M parameters with 40B tokens. - BananaMind 2.1 Flash: 50M parameters with 55B tokens. - BananaMind 2.1 Pro: 135M-145M parameters (still deciding) with 100B tokens. These model will use a multi tower architecture (like https://huggingface.co/BananaMind/BananaMind-2.1-Unified) with some more architectural changes. BananaMind 2.1 Pro will probrably use 2 no output towers, instead of one! We're currently training some experimental models based on this architecture to see its scaling! Follow us: https://huggingface.co/BananaMind @Banaxi-Tech @vovaRL @DedeProGames https://huggingface.co/bananamind-research-community
posted
an
update
about 3 hours ago
We're announcing our BananaMind 2.1 model series! The models will include: - BananaMind 2.1 Nano: 10M parameters with 60B tokens. - BananaMind 2.1 Lite: 25M parameters with 40B tokens. - BananaMind 2.1 Flash: 50M parameters with 55B tokens. - BananaMind 2.1 Pro: 135M-145M parameters (still deciding) with 100B tokens. These model will use a multi tower architecture (like https://huggingface.co/BananaMind/BananaMind-2.1-Unified) with some more architectural changes. BananaMind 2.1 Pro will probrably use 2 no output towers, instead of one! We're currently training some experimental models based on this architecture to see its scaling! Follow us: https://huggingface.co/BananaMind @Banaxi-Tech @vovaRL @DedeProGames https://huggingface.co/bananamind-research-community
replied
to
their
post
about 3 hours ago
We're releasing BananaMind 2.1 Unified, a 35M three-tower relay model where the two output towers can only talk to each other through a silent middle tower that has no output head and no loss term. Tower B trains entirely on indirect gradient. It was never told what to predict. It woke up anyway. Jacobian lens shows it carrying the correct answer ("Paris", "oxygen", "blue") at its deepest layer. Its bridge gates grew 4-47x from init. Feed it from only one side and the representations collapse to junk โ it needs both outer towers to become semantic. The PIQA result is the cleanest demonstration: Tower A alone scores 50.11 (chance). Tower C alone 52.12. Full system 61.75. All physical reasoning lives in the integration. The 35M three-tower beats the 50M single-tower BananaMind 2 Medium on PIQA. Trained on 38B tokens in ~7.5 hours on 8x RTX PRO 6000. Ships with 7 ablation modes so you can surgically cut the model apart without retraining. Full training logs, J-lens fits, and eval outputs for every mode included. This is a research model. It is interesting because you can take it apart. https://huggingface.co/BananaMind/BananaMind-2.1-Unified Follow us for more: https://huggingface.co/BananaMind @Banaxi-Tech
View all activity
Organizations
Banaxi-Tech
's datasets
1
Sort:ย Recently updated
Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500
Viewer
โข
Updated
May 24
โข
2.56k
โข
552
โข
12