Request: Qwen3.8-27B-NVIDIA-NVFP4-no-MTP-GGUF and -uncensored/ABLITERATED
Request: Qwen3.8-27B-NVIDIA-NVFP4-no-MTP-GGUF and -uncensored/ABLITERATED
Hi mradermacher team. First, thank you for your many models produced over the last few years.
My hardware isn't the best but it goes - 5060+3050, 14gb total. MTP just isn't reasonable in this setup since the OS uses 2-3gb of the 3050's 6gb, leaving a max of ~12.1gb. It's a tight space to work in.
Looks like the new qwen 3.8 is koth. Been steering towards any optimizations I can, nvfp4 now being a go-to. Stripping out MTP I think is useful though llm neural surgeons like yourselves may refute real gains. I'd like to request a q4km and/or a q6k of nvfp4, no-mtp build of qwen3.8, and if possible an uncensored/ABLITERATED version.
Speaking of uncensored, this producer just dropped a build with a technique that seems interesting to consider. Thoughts on if this is useful to integrate?
https://huggingface.co/jaromer/0bserverx-Qwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF
Have a great day!
you provided me with a gguf, so cant queue automatically, and he provided what seems like most quants, so I dont think I should really do anything manually right now
Thoughts on if this is useful to integrate?
integrate what? we only quantize, we do not heretic, we do not create mtp, we do not use custom builds of llamacpp. If you want heretic, heretic-org is your best friend in this case, people like davidAU, muxodious, coder3101 and others are very nice at this, you should consult with them. Quants? You ask us, sure =)