Configuration Parsing Warning:In config.json: "quantization_config.bits" must be an integer
AREX-Base EXL3 4.1 bpw
Quality-oriented EXL3 conversion of BAAI/AREX-Base, a Qwen3.5-122B-A10B multimodal MoE research-agent model.
Quantization
- EXL3 decoder target: 4.10 bpw
- Effective HQ allocation: 4.14 bpw in the generated configuration
- Sensitive MoE and attention tensors: higher-bit allocation via
--hq - LM head: 6 bpw
- Codebook: MCG
- Calibration: 250 rows × 2,048 tokens
- Output-channel scales: always enabled
- Vision tower: preserved in BF16
- MTP head: official Qwen3.5-122B-A10B BF16 donor, quantized at Q8
The conversion used ExLlamaV3 1.2.1. Use a recent ExLlamaV3-compatible runtime with Qwen3.5 MoE and multimodal support, such as a compatible TabbyAPI build.
Notes
This repository contains quantized weights, not benchmark results. Validate the model on your own prompts and serving configuration. Refer to the original model card for intended use, architecture details, limitations, and citation information.
The published AREX checkpoint advertises one MTP layer in its configuration
but contains no mtp.* tensors. This conversion restores the architecture-
compatible MTP head from the official Qwen3.5-122B-A10B base. Since AREX
fine-tuned the target model without publishing a matching fine-tuned draft
head, MTP acceptance may differ from the base model and should be measured
on representative prompts.
License
The source model is released under Apache-2.0. See the included LICENSE
and the original model repository.
- Downloads last month
- 17