Instructions to use OrionLLM/GRM-2.6-Plus-0628 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Inference
Some thoughts
Hey!
I've seen your post in DavidAU's repo about benches.
I fully agree with you since I am interested in agentic performance. Their benches say little about that.
Still the hype is fantastically extreme.
What I believe you should do: run ARC-Challenge and post results if they are better than theirs. Just another one bench.
They specifically tune to that bench, so that is not honest comparison. That's why post only if yours is better.
Yes, people generally look for benchmarks like Terminal Bench 2.1, which specifically shows the model's Agentic Performance. Even if the hype is huge and it might even beat Fable 5 on the ARC-Challenge, that says nothing about whether it will outperform Fable 5 in Agentic Performance.
Regarding benchmarking my model, I'm not sure if I should benchmark it on ARC, because nobody really cares about that. There are many other benchmarks that are much better at showing whether a model has a high capacity for organized logical reasoning. If it lacks modern benchmarks used by state-of-the-art models (and instead relies on classic benchmarks), there's no way to confidently tell the public if that model truly has the incredible reasoning performance that Fable demonstrates. This just leaves people more in doubt about whether the model has actually exceeded the 'OpenAI, Claude and Gemini zone of intelligence' or if it doesn't even know how to use tools.
Furthermore, DavidAU never takes his models seriously. He presents a model 'as good as Fable-5', but it comes with an informal, poorly organized README.md with random GIFs.
In contrast, my models feature organized releases, a clean and easy-to-read README.md, demonstrate solid performance, and use the same benchmarks as frontier models.
Also check out this i opened => https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF/discussions/83
Agreed. But somehow people download his models much more. That's why I've proposed to beat him on his bench instead of making him doing right benches (which he refuses from what I've seen). Just for marketing.
Most people don't understand all these benches and stuff. They install LM Studio (or Ollama, or whatever), download most popular gguf and live with it for the next few months. Some go little further: most popular has arc=711, let's check other guys, oh, they don't have arc at all and they don't have millions of downloads, obviously 711 is better. I would like to be wrong about people, but everything points in that direction.
Anyway I clearly understand why you don't do classic benchmarks. And I appreciate a lot your benchmark selection and effort you put in benchmarking (I've tried to do some benches myself and it was real pain). That's why I use GRM: quality is proven. And I see that improvement over original Qwen in my sessions.
BTW your post is deleted.
I know! My post was deleted BY David himself! He wants to censor me; here is an article with all the information from the post => https://huggingface.co/blog/DedeProGames/davidau-deleted-my-post-and-banned-me-for-asking-a
Also, I'm now banned from posting any Discussions on his models.
If I could ask you for one thing: could you help me spread the word about this?
Upvoted and left a comment.
I've properly evaluated his quants of 711. Some are good, most are not: https://huggingface.co/NikiKrutan/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-GGUF
Now I begin to understand this "market". Everything as usual: louder = better. Sadly.