Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Pranav Patel
Pranav-IO
1
6
Follow
0 followers
·
6 following
https://ionio.ai
pranav2278
AI & ML interests
None yet
Recent Activity
replied
to
Banaxi-Tech
's
post
1 day ago
We're releasing BananaMind 2.1 Unified, a 35M three-tower relay model where the two output towers can only talk to each other through a silent middle tower that has no output head and no loss term. Tower B trains entirely on indirect gradient. It was never told what to predict. It woke up anyway. Jacobian lens shows it carrying the correct answer ("Paris", "oxygen", "blue") at its deepest layer. Its bridge gates grew 4-47x from init. Feed it from only one side and the representations collapse to junk — it needs both outer towers to become semantic. The PIQA result is the cleanest demonstration: Tower A alone scores 50.11 (chance). Tower C alone 52.12. Full system 61.75. All physical reasoning lives in the integration. The 35M three-tower beats the 50M single-tower BananaMind 2 Medium on PIQA. Trained on 38B tokens in ~7.5 hours on 8x RTX PRO 6000. Ships with 7 ablation modes so you can surgically cut the model apart without retraining. Full training logs, J-lens fits, and eval outputs for every mode included. This is a research model. It is interesting because you can take it apart. https://huggingface.co/BananaMind/BananaMind-2.1-Unified Follow us for more: https://huggingface.co/BananaMind @Banaxi-Tech
reacted
to
Banaxi-Tech
's
post
with 🚀
1 day ago
We're releasing BananaMind 2.1 Unified, a 35M three-tower relay model where the two output towers can only talk to each other through a silent middle tower that has no output head and no loss term. Tower B trains entirely on indirect gradient. It was never told what to predict. It woke up anyway. Jacobian lens shows it carrying the correct answer ("Paris", "oxygen", "blue") at its deepest layer. Its bridge gates grew 4-47x from init. Feed it from only one side and the representations collapse to junk — it needs both outer towers to become semantic. The PIQA result is the cleanest demonstration: Tower A alone scores 50.11 (chance). Tower C alone 52.12. Full system 61.75. All physical reasoning lives in the integration. The 35M three-tower beats the 50M single-tower BananaMind 2 Medium on PIQA. Trained on 38B tokens in ~7.5 hours on 8x RTX PRO 6000. Ships with 7 ablation modes so you can surgically cut the model apart without retraining. Full training logs, J-lens fits, and eval outputs for every mode included. This is a research model. It is interesting because you can take it apart. https://huggingface.co/BananaMind/BananaMind-2.1-Unified Follow us for more: https://huggingface.co/BananaMind @Banaxi-Tech
reacted
to
Banaxi-Tech
's
post
with 🔥
1 day ago
We're releasing BananaMind 2.1 Unified, a 35M three-tower relay model where the two output towers can only talk to each other through a silent middle tower that has no output head and no loss term. Tower B trains entirely on indirect gradient. It was never told what to predict. It woke up anyway. Jacobian lens shows it carrying the correct answer ("Paris", "oxygen", "blue") at its deepest layer. Its bridge gates grew 4-47x from init. Feed it from only one side and the representations collapse to junk — it needs both outer towers to become semantic. The PIQA result is the cleanest demonstration: Tower A alone scores 50.11 (chance). Tower C alone 52.12. Full system 61.75. All physical reasoning lives in the integration. The 35M three-tower beats the 50M single-tower BananaMind 2 Medium on PIQA. Trained on 38B tokens in ~7.5 hours on 8x RTX PRO 6000. Ships with 7 ablation modes so you can surgically cut the model apart without retraining. Full training logs, J-lens fits, and eval outputs for every mode included. This is a research model. It is interesting because you can take it apart. https://huggingface.co/BananaMind/BananaMind-2.1-Unified Follow us for more: https://huggingface.co/BananaMind @Banaxi-Tech
View all activity
Organizations
Pranav-IO
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
liked
a dataset
1 day ago
Mercity/ReasonBridge-URT
Viewer
•
Updated
Nov 12, 2025
•
1.02M
•
42
•
6
liked
a dataset
3 days ago
Ionio-ai/ecommerce-search-extraction
Viewer
•
Updated
8 days ago
•
11k
•
117
•
2
liked
a dataset
3 months ago
pollen-robotics/reachy-mini-emotions-library
Viewer
•
Updated
Jul 7
•
81
•
6.73k
•
16
liked
a model
11 months ago
facebook/SONAR
Updated
Feb 14, 2024
•
70
liked
a dataset
about 1 year ago
euclaise/writingprompts
Viewer
•
Updated
Sep 21, 2023
•
303k
•
4.96k
•
70
liked
a model
over 2 years ago
daxa-ai/pebblo-classifier
Text Classification
•
Updated
May 30, 2024
•
489
•
9