Model Requests & Suggestions

#1
by IsValorum - opened

Suggest a model for a future handcrafted MiniPlus or NanoPlus release.

Please include:

  • Upstream model link
  • Intended use or strengths
  • Why it would be valuable at a MiniPlus or NanoPlus footprint

I will review requests based on demand, architecture, local-inference practicality, and whether the model would benefit from a custom tensor-by-tensor quantization.

IsValorum pinned discussion

CyberTiel

  1. Link: peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-GGUF-MTP (https://huggingface.co/peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-GGUF-MTP)
  2. Just like Tiel , agentic coding tasks, but stronger, less thinking(in my opinion) and uncensored
  3. Miniplus like what you did on Tiel

The MiniPlus V2.1 and NanoPlus quantizations are already underway; they will be available within an hour at most—follow me to get notified.

Wow
Thanks for your hard works and very fast reply, really impressive, this is the fastest model request i've seen on hf (or maybe you had already worked on this model from before :))

I will definitely check and test for if the Apex Miniplus CyberTiel work good too

Request for next model.
Reason: I've personally used k2 and it's the closest thing to match or exceed Qwen intelligence:
IFM/K2-Horizon-MoVA-36B-A4B

I personally tried the K2-Horizon, the smaller versions and that MoVA, and they didn't convince me because of the cost of the KV Cache.

I could barely get it to 8K context in Q4_0, not enough to properly test it, is it really any good?

For me, the fact that it had such an absurdly expensive KV cache was enough to completely rule it out. Did you guys test it in real-world use?

You were using Q4_0 KV?
How much VRAM on the rig?
I assume your models were tested on 24gb ram?

I tried it on an RTX 3090 and I simply didn't find that series of models appealing; the KV cache was too expensive.

Sign up or log in to comment