Benchmark Scores

#5
by Austriani - opened

Empero-AI, are benchmarks provided for Qwen3.5-9B made by using Qwen3.5-9B-Base or Qwen3.5-9B? I wonder why its written 32.8 for Qwen3.5-9B, while search shows that Qwen3.5-9B scores more than 80% on that benchmark.
I read that you used your own types of benchmarks, but how much do they differ from original benchmarks, and what was the model judging benchmark answers?

Didn't understood at first, but now I know why scores differ.

Austriani changed discussion status to closed

Sign up or log in to comment