Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
CodeSoft 
posted an update 3 days ago

@CodeSoft Very nice first model, you still got a lot to improve on (since 26% on Hellaswag) but you could definitely get competitive!

·

My tips would be
Use more tokens even 10B tokens could get competitive!
Replace Infiniwebmath to FineMAth 4 PLus
Add some Fineweb HQ

But thats just my opinion.

@CodeSoft Congrats on the release! A clean 25M Qwen2 model with native GGUF and Transformers support trained in under 4 hours on a 5060 Ti is awesome work.

If you are planning v2, scaling to around 2.5B to 5B tokens (~100 to 200 tokens/param) with a careful mix of datasets will give you huge quality jumps.

Just be careful of extreme overtraining past that range. Pushing a 25M model past 10B+ tokens hits sharp diminishing returns, locks weights into brittle minima, and degrades fine-tuning plasticity and quantization stability.

Focusing on clean data curation and curriculum over raw token volume is the best path. Excited to see what you do with the v2 chat version!

·

What data curation is he going to use? Make his own dataset?

Then show me you making your own dataset.
And extreme overtraining, that only appears around 22K:1, that would be 550B, not 10B