AI & ML interests

Community of researchers interested in OpenBuddy Early Access. Please note that this is not an official group for OpenBuddy, and the members have no affiliation with the OpenBuddy team. Every independent researchers can apply to join by submitting a Form available in our GitHub.

AtAndDev 
posted an update 1 day ago
view post
Post
2627
@Banaxi-Tech stop hiding my comments. AND STOP STEALING PAPERS AND SPREADING MISINFORMATION.
your BGA blog is a copy of NSA (deepseek, 2025) branded under your name. literally the same top16 selected blocks, 512 local window, router over block summaries, all you did was change block size from 64 to 128.
you didnt cite NSA once but you put a “please cite BGA” bibtex at the bottom.
i commented under your post and said that there is no way that you can support claims like: “The Accuracy Should BE WAy better than DSA but untested yet.” you didnt run a single experiment. and the 256x isnt from BGA, its just n/2k with k=2048 so the exact same k DSA uses. if opus wrote this for you, at least read it before posting.
i commented again after you hid my comment despite it having constructive and correct feedback and you hid that too. and again.
you can hide the truth and just try to get hf post likes..... but is it really the thing that needs to be done? do you really want to take papers and make them yours while barely even changing the params?

admitting your mistakes and doing something about them needs humbleness, intelligence, humanness.
i encourage you to admit your mistakes and try to do better next time (at least read what blog your ai wrote or do proper experiments to back your stuff up).
  • 41 replies
·
raincandy-u 
posted an update 5 days ago
view post
Post
3735
20K parameters can tell a story. 🚀

🤗 We trained a ~20k-parameter Transformer that can actually write stories!

raincandy-u/MacroStories

→ ~50× smaller than the 1M-parameter TinyStories model
→ ~3,000× smaller than AlexNet
→ 81 KB in FP32

yayyy the whole model. ૮ ˶ᵔ ᵕ ᵔ˶ ა

She has a 32-dimensional hidden state, a 378-token vocabulary, and just one decoder block — recurrently applied 4 times with shared weights.

Despite having only 19,969 parameters, she can maintain a simple narrative across 100–300 words: establish a goal, encounter a problem, take relevant actions, and reach an outcome.

She runs extremely fast on CPU — no GPU required. The entire model is tiny enough to load almost instantly! ☺️
  • 3 replies
·
AtAndDev 
posted an update about 1 month ago
view post
Post
2921
SPECK 2 IS ALREADY OUT: specklabs/Speck2-140M

Pretrained on 4x more tokens than the previous releases (20b vs 5b).
Instruct tuned versions are coming soon.
Very interesting models are coming soon too (hint: super long context).

Thanks for everyone supporting!
  • 3 replies
·
AtAndDev 
posted an update about 1 month ago
view post
Post
211
SPECK1.5 IS COMING SOON!
Same 5B token budget but much better corpus quality.

Also getting a ton of downloads, thanks for everyone downloading and liking <3

specklabs
AtAndDev 
posted an update about 2 months ago
view post
Post
160
NEW SPECK UPDATES:

Just hit #14 and #15 with out FIRST models on Open SLM Leaderboard. The models were trained on 5B tokens, while competing with similarly sized models trained on more than 6-20x the data.

A new base model Speck1.5-140M being trained right now on a higher quality corpus and will be released soon.
SpeckChat3 is coming very soon with 1 million samples, specifically designed to post train small base models.

Also, just to clarify stuff, we will NOT release anything that is NOT MIT licensed EVER. Openness is needed in small language research.

Thanks to everyone supporting the project, and stay tuned for new releases!
AtAndDev 
posted an update about 2 months ago
view post
Post
1894
SPECK UPDATES:
1 New instruct model tuned on top of Speck1-140M: specklabs/Speck1-140M-Instruct
2 Instruction tuning datasets
2 GGUFs

Much more coming soon:
Speck1.1-140M-Instruct that is post trained on SpeckChat2 will be coming very soon
New base model Speck1.5-140M is coming with a much higher quality corpus

Thanks to everyone who is already supporting the project, and stay tuned for new releases!
  • 3 replies
·
AtAndDev 
posted an update about 2 months ago
view post
Post
2130
FIRST SPECK MODEL RELEASED:
specklabs/Speck1-140M

new models coming very soon (both instruct and much better models), with much much higher training scale as i am getting marenostrum5 access soon!
we will be looking at 100b-2t token budgets :)
  • 4 replies
·
raincandy-u 
posted an update 8 months ago
view post
Post
3274
Introducing Rain-v2: Democratizing LLM training on gaming GPUs! ⚡

​Following Rain-100M, we’re scaling up. Rain-v2 features a larger training dataset.

We’ve published a comprehensive blog covering the end-to-end journey—from raw data collection to rigorous evaluation and safety testing.

​HF Repo: 🤗 raincandy-u/Rain-v2

​Blog: 📚
https://angelkawaii.xyz/2026/01/29/rain-v2/

​Special thanks to the open-source community and the SmolLM2 team for their foundational work! 🚀

HuggingFaceTB

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model (2502.02737)
  • 2 replies
·
raincandy-u 
posted an update 9 months ago
view post
Post
5875
🤗 Just released Rain-100M, an experimental ~97M-parameter Qwen3-style language model trained from random initialization.

Repo: raincandy-u/Rain-100M

Data: HuggingFaceFW/fineweb-edu, ~3B tokens, English only

Tokenizer: custom 16k BPE, context length 4096

Architecture: 12 Transformer layers, hidden size 768, 12 heads, MLP 2048, SiLU, bf16


Rain-100M is a raw base model (not instruction-tuned or safety-aligned), aimed at small-scale research, debugging training pipelines, and CPU/edge experiments. If you run evaluations, finetunes, or visualizations with it, I would be very interested in your results!
  • 3 replies
·
AtAndDev 
posted an update about 1 year ago
view post
Post
786
Qwen 3 Coder is a personal attack to k2, and I love it.
It achieves near SOTA on LCB while not having reasoning.
Finally people are understanding that reasoning isnt necessary for high benches...

Qwen ftw!

DECENTRALIZE DECENTRALIZE DECENTRALIZE
AtAndDev 
posted an update over 1 year ago
view post
Post
3258
deepseek-ai/DeepSeek-R1-0528

This is the end
  • 2 replies
·
AtAndDev 
posted an update over 1 year ago
view post
Post
3164
Llama 4 is out...
  • 3 replies
·
AtAndDev 
posted an update over 1 year ago
view post
Post
4400
There seems to multiple paid apps shared here that are based on models on hf, but some ppl sell their wrappers as "products" and promote them here. For a long time, hf was the best and only platform to do oss model stuff but with the recent AI website builders anyone can create a product (really crappy ones btw) and try to sell it with no contribution to oss stuff. Please dont do this, or try finetuning the models you use...
Sorry for filling yall feed with this bs but yk...
  • 6 replies
·
AtAndDev 
posted an update over 1 year ago
view post
Post
1688
Gemma 3 seems to be really good at human preference. Just waiting for ppl to see it.
AtAndDev 
posted an update over 1 year ago
view post
Post
2513
@nroggendorff is that you sama?
  • 2 replies
·
AtAndDev 
posted an update over 1 year ago
view post
Post
1962
everywhere i go i see his face
AtAndDev 
posted an update over 1 year ago
view post
Post
597
Deepseek gang on fire fr fr
AtAndDev 
posted an update over 1 year ago
view post
Post
1673
R1 is out! And with a lot of other R1 releated models...
AtAndDev 
posted an update almost 2 years ago
view post
Post
510
@s3nh Hey man check your discord! Got some news.
  • 4 replies
·