Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
h s
HubertSchmitt
2
5
Follow
0 followers
·
3 following
AI & ML interests
None yet
Recent Activity
replied
to
Banaxi-Tech
's
post
14 days ago
Today we wanted to release BananaMind 2 Pico, our smallest model yet at ~0.9M parameters. Instead, we accidentally ran a very expensive experiment on what happens when you push a tiny model way past its useful token budget. Short version: we trained on 200B tokens (~222K:1 tokens-per-parameter). The model peaked at 20B tokens with an INT Index of 4.55, then degraded monotonically over the next 160B to 3.31 — a 27% regression. Three of four Open SLM benchmarks were worse at the end of training than they were at 10% through. The useful compute-optimal range for Pico-tier models looks like ~22K–30K tokens per parameter. Ratios like 7K:1, 15K:1, and 22K:1 all work fine — TinyStories and most sub-3M community models sit in this range. Push much further and benchmarks start rotting. Follow us for more: https://huggingface.co/BananaMind @vovaRL @Banaxi-Tech Full writeup with all checkpoints, the Chinchilla-ratio control run, and the schedule-vs-overtraining analysis: https://huggingface.co/blog/Banaxi-Tech/ovdadadadd And if anyone, i dont know the reason why you would, wants the 20B token checkpoint reply and ill upload it as BananaMind 2.1 Pico EXP
liked
a model
14 days ago
Qwen/Qwen3.8-27B
commented
on
a paper
19 days ago
GPT-Red: Automated Red Teaming via Self-Play at Scale
View all activity
Organizations
None yet
HubertSchmitt
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
commented
a paper
19 days ago
GPT-Red: Automated Red Teaming via Self-Play at Scale
Paper
•
2607.26115
•
Published
Jul 28
•
9
•
4
New activity in
nvidia/Nemotron-Post-Training-Dataset-v2
4 months ago
difference between nvidia/Nemotron-Post-Training-Dataset-v1 and nvidia/Nemotron-Post-Training-Dataset-v2
2
#4 opened 12 months ago by
cmathx