🔄 In a Training Loop

Akshay Mhaskar

celestialcreator

AI & ML interests

None yet

Recent Activity

liked a model 26 days ago

celestialcreator/Llama-3.2-1B-MTP-k8

updated a model 3 months ago

celestialcreator/axon-smollm2-360m

published a model 3 months ago

celestialcreator/axon-smollm2-360m

View all activity

Organizations

liked a model 26 days ago

celestialcreator/Llama-3.2-1B-MTP-k8

Text Generation • Updated Mar 5 • 552 • • 3

updated a model 3 months ago

celestialcreator/axon-smollm2-360m

Text Generation • 0.4B • Updated Apr 10 • 4

published a model 3 months ago

celestialcreator/axon-smollm2-360m

Text Generation • 0.4B • Updated Apr 10 • 4

reacted to SeaWolf-AI's post with 🤯 4 months ago

Post

11147

🏟️ Smol AI WorldCup: A 4B Model Just Beat 8B — Here's the Data

We evaluated 18 small language models from 12 makers on 125 questions across 7 languages. The results challenge the assumption that bigger is always better.

Community Article: https://huggingface.co/blog/FINAL-Bench/smol-worldcup
Live Leaderboard: ginigen-ai/smol-worldcup
Dataset: ginigen-ai/smol-worldcup

What we found:

→ Gemma-3n-E4B (4B, 2GB RAM) outscores Qwen3-8B (8B, 5.5GB). Doubling parameters gained only 0.4 points. RAM cost: 2.75x more.

→ GPT-OSS-20B fits in 1.5GB yet matches Champions-league dense models requiring 8.5GB. MoE architecture is the edge AI game-changer.

→ Thinking models hurt structured output. DeepSeek-R1-7B scores 8.7 points below same-size Qwen3-8B and runs 2.7x slower.

→ A 1.3B model fabricates confident fake content 80% of the time when prompted with nonexistent entities. Qwen3 family hits 100% trap detection across all sizes.

→ Qwen3-1.7B (1.2GB) outscores Mistral-7B, Llama-3.1-8B, and DeepSeek-R1-14B. Latest architecture at 1.7B beats older architecture at 14B.

What makes this benchmark different?

Most benchmarks ask "how smart?" — we measure five axes simultaneously: Size, Honesty, Intelligence, Fast, Thrift (SHIFT). Our ranking metric WCS = sqrt(SHIFT x PIR_norm) rewards models that are both high-quality AND efficient. Smart but massive? Low rank. Tiny but poor? Also low.

Top 5 by WCS:
1. GPT-OSS-20B — WCS 82.6 — 1.5GB — Raspberry Pi tier
2. Gemma-3n-E4B — WCS 81.8 — 2.0GB — Smartphone tier
3. Llama-4-Scout — WCS 79.3 — 240 tok/s — Fastest model
4. Qwen3-4B — WCS 76.6 — 2.8GB — Smartphone tier
5. Qwen3-1.7B — WCS 76.1 — 1.2GB — IoT tier

Built in collaboration with the FINAL Bench research team. Interoperable with ALL Bench Leaderboard for full small-to-large model comparison.

Dataset is open under Apache 2.0 (125 questions, 7 languages). We welcome new model submissions.

1 reply

updated a model 4 months ago

zosmaai/Qwen3.5-0.8B-GRPO-Math

Text Generation • 0.8B • Updated Mar 10 • 6 • 1

published a model 4 months ago

zosmaai/Qwen3.5-0.8B-GRPO-Math

Text Generation • 0.8B • Updated Mar 10 • 6 • 1

updated a dataset 4 months ago

zosmaai/Qwen3.5-0.8B-GRPO-Math-Dataset

Viewer • Updated Mar 10 • 1k • 12 • 1

published a dataset 4 months ago

zosmaai/Qwen3.5-0.8B-GRPO-Math-Dataset

Viewer • Updated Mar 10 • 1k • 12 • 1

updated a model 4 months ago

celestialcreator/Qwen3.5-0.8B-GRPO-Math

Text Generation • 0.8B • Updated Mar 9 • 3

published a model 4 months ago

celestialcreator/Qwen3.5-0.8B-GRPO-Math

Text Generation • 0.8B • Updated Mar 9 • 3

updated a dataset 4 months ago

celestialcreator/Qwen3.5-0.8B-GRPO-Math-Dataset

Viewer • Updated Mar 9 • 1k • 25 • 1

published a dataset 4 months ago

celestialcreator/Qwen3.5-0.8B-GRPO-Math-Dataset

Viewer • Updated Mar 9 • 1k • 25 • 1

reacted to kostakoff's post with 👍 4 months ago

Post

2257

Mining GPU Nvidia CMP 170HX - let's run some models!

To satisfy my curiosity, I investigated different GPUs and found this: a mining version of the A100 — the CMP 170HX.

It is a very interesting GPU. Based on public documentation, it has hardware similar to the datacenter A100. If you open it up and look at the board, you will see that it's very similar to an A100 board; it even has NVLink connectors.

Online, I found almost no information about how to run it, whether it works with LLMs, or if it's supported by default Nvidia drivers and CUDA. So, I decided to test it myself.
I installed it in my lab (see previous post https://huggingface.co/posts/kostakoff/584269728210158) and found that the default nvidia-driver-570 works with it out of the box. After that, I checked if CUDA was available, and it worked too.

The next step was to try running some models:
- Stable Diffusion XL with BNB4 quantization: It took around two minutes to generate an image, but it works!
- Compiled llama.cpp for CUDA (https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md#compilation): I run Mistral 7B Q4_K_M, and this actually worked even better. It was able to generate 33 tokens per second and read 400 tokens per second.

There are some limitations related to power utilization:
- When running PyTorch, it doesn't utilize more than 80 watts.
- When running llama.cpp, utilization is a bit better but still limited to 113 watts.

I found this GitHub thread about the Nvidia CMP https://github.com/dartraiden/NVIDIA-patcher/issues/73, and it looks like this mining GPU has an internal rate limiter based on FMA compute calls. I haven't found a solution to bypass it yet.