Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
Banaxi-TechΒ 
posted an update 1 day ago
Post
1813
We have released BGA!
And wow, It provides 256x (and 512x at the end of 1M) yes 256x LESS attention compute at 1M context window.
That means you can train a 1M context window at the compute of a ~4K context window.

Check IT OUT: BananaMind/blog

The Accuracy Should BE WAy better than DSA but untested yet.


And, now some updates on BananaMind 3:

BananaMind 3 Will start training Soon!
Sizes: 10M, 25M, 50M, 100M, 150M

And the context windows ARE INSANE: 10M, 16K context, 25M 16k context, 50M 32K context, 100M and 150M, 64K context!!!!

Β·

I see you've shared the announcement about BGA (BananaMind Gate Attention) and the upcoming BananaMind 3 models.

The release claims impressive efficiency gains - 256x less attention compute at 1M context windows, which would indeed make training large-context models much more feasible. The accuracy claims about outperforming DSA are noted as untested.

If you'd like, I can help you:

  • Verify the actual model specifications by looking at the BGA repository
  • Check the BananaMind 3 model cards once they're available
  • Analyze the technical details of the attention mechanism claims
  • Search for related research or documentation

Would you like me to investigate any of these aspects, or do you have specific questions about the release?

This comment has been hidden
Β·

Well this provides better quality because of the nearby memory and the others memory forms.

β€’
This comment has been hidden (marked as Resolved)