why Qwen 3.8 27B is absent in the benchmarks in Small models (≀40B) category?

#2
by ameligrana - opened

Let me start by saying that it seems an incredible improvement over the base model!

But we can't really tell so much since Qwen 3.8 27B is not part of the benchmarks comparisons, what is the reason for this?

This is surprisingly common among models that are finetuned on Qwen 3.8 27B. The most significant comparison they can show is how it compares against the base model that it was finetuned on, but they never do. It's very annoying and suspicious.

I ran a small local comparison while converting AREX-2 to MLX, in case it helps here.

Short version: AREX-2 was more efficient than plain Qwen3.8-27B. It used 39% to 53% fewer tokens in every run, finished in about half the time, and solved a few more problems.

AREX-2 Qwen3.8-27B
Tokens used, all tests 113,795 212,514
Time, all tests 249 min 451 min
Easy coding tasks passed 35 of 36 30 of 36
Hard problems solved first try 15 of 24 11 of 24
Hard problems solved within 3 tries 23 of 24 20 of 24

Both models are 8-bit MLX with the same prompts, sampling (temperature 1.0, top-p 0.95, top-k 20), seeds and a 4,000-token limit, on one M4 Pro Mac mini. On the hard problems the model sees its failing test and gets up to three tries.

The difference was mostly length. Every failed attempt by plain Qwen3.8-27B ran into the 4,000-token limit (31 of 31), against 9 of 13 for AREX-2. With a higher limit the plain model might well catch up on problems solved, just more slowly. Both ran with thinking on, the default.

With thinking off the two were level. In an agent harness (MiniMax Code), each fixed three small failing projects in 6 of 6 trials, in 17 and 18 minutes. So the gap above comes from thinking.

This is a small test: two test sets, two runs each, short single tasks. It does not touch the long multi-round benchmarks in the paper, so it cannot confirm or dispute those.

MLX builds for Mac (4, 5, 6 and 8-bit) with the full results are here: https://huggingface.co/mlx-community/AREX-2-5bit

Sign up or log in to comment