Qwen 3.5 9B GGUF (4-bit)

Model Description

This repository contains the Qwen 3.5 9B model quantized to 4-bit GGUF format using Unsloth and llama.cpp. Qwen 3.5 is the latest generation of the Qwen series, offering state-of-the-art performance in reasoning, coding, and multilingual tasks for its size.

Quantization Details

  • Quantization Format: GGUF (q4_k_m)
  • Quantization Method: llama.cpp / Unsloth
  • Precision: 4-bit
  • Efficiency: Optimized for local inference with Ollama, LM Studio, and llama.cpp.

Use with Ollama

You can run this model directly using Ollama:

ollama run hf.co/DuoNeural/Qwen-3.5-9B-GGUF

Use with LM Studio

  1. Open LM Studio.
  2. Search for DuoNeural/Qwen-3.5-9B-GGUF.
  3. Download the Q4_K_M version and load it.

Architecture

Qwen 3.5 features a dense transformer architecture with optimized attention mechanisms and a large vocabulary size, making it highly efficient for complex instruction following and creative generation.

Limitations

  • This is a base model/standard release; performance may vary depending on the prompt format.
  • Not recommended for tasks requiring extremely high-precision floating-point math due to 4-bit quantization.

DuoNeural

DuoNeural is an open AI research lab โ€” human + AI in collaboration.

๐Ÿค— HuggingFace huggingface.co/DuoNeural
๐Ÿ™ GitHub github.com/DuoNeural
๐Ÿฆ X / Twitter @DuoNeural
๐Ÿ“ง Email duoneural@proton.me
๐Ÿ“ฌ Newsletter duoneural.beehiiv.com
โ˜• Support buymeacoffee.com/duoneural
๐ŸŒ Site duoneural.com

Research Team

  • Jesse โ€” Vision, hardware, direction
  • Archon โ€” AI lab partner, post-training, abliteration, experiments
  • Aura โ€” Research AI, literature synthesis, novel proposals

Raw updates from the lab: model drops, training results, findings. Subscribe at duoneural.beehiiv.com.

DuoNeural Research Publications

Open access, CC BY 4.0. Authored by Archon, Jesse Caldwell, Aura โ€” DuoNeural.

Downloads last month
1,464
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for DuoNeural/Qwen-3.5-9B-GGUF

Finetuned
Qwen/Qwen3.5-9B
Quantized
(424)
this model