Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Ahmad Saeed Zaidi
Rolaficus
2
Follow
webbrain-one-679468's profile picture
1 follower
·
1 following
AhmadSaeedZaidi
AI & ML interests
RL, CV, NLP, VLM, VLAM
Recent Activity
updated
a dataset
about 18 hours ago
Rolaficus/pleiades-vault-clean
reacted
to
AtAndDev
's
post
with 😔
1 day ago
@Banaxi-Tech stop hiding my comments. AND STOP STEALING PAPERS AND SPREADING MISINFORMATION. your BGA blog is a copy of NSA (deepseek, 2025) branded under your name. literally the same top16 selected blocks, 512 local window, router over block summaries, all you did was change block size from 64 to 128. you didnt cite NSA once but you put a “please cite BGA” bibtex at the bottom. i commented under your post and said that there is no way that you can support claims like: “The Accuracy Should BE WAy better than DSA but untested yet.” you didnt run a single experiment. and the 256x isnt from BGA, its just n/2k with k=2048 so the exact same k DSA uses. if opus wrote this for you, at least read it before posting. i commented again after you hid my comment despite it having constructive and correct feedback and you hid that too. and again. you can hide the truth and just try to get hf post likes..... but is it really the thing that needs to be done? do you really want to take papers and make them yours while barely even changing the params? admitting your mistakes and doing something about them needs humbleness, intelligence, humanness. i encourage you to admit your mistakes and try to do better next time (at least read what blog your ai wrote or do proper experiments to back your stuff up).
replied
to
Banaxi-Tech
's
post
1 day ago
We have released BGA! And wow, It provides 256x (and 512x at the end of 1M) yes 256x LESS attention compute at 1M context window. That means you can train a 1M context window at the compute of a ~4K context window. Check IT OUT: https://huggingface.co/spaces/BananaMind/blog#bananamind-gate-attention The Accuracy Should BE WAy better than DSA but untested yet. And, now some updates on BananaMind 3: BananaMind 3 Will start training Soon! Sizes: 10M, 25M, 50M, 100M, 150M And the context windows ARE INSANE: 10M, 16K context, 25M 16k context, 50M 32K context, 100M and 150M, 64K context!!!!
View all activity
Organizations
None yet
Rolaficus
's buckets
1
Sort: Recently updated
Rolaficus/data-synth-media-public-smoke
10.3 MB