Configuration Parsing Warning:In config.json: "quantization_config.bits" must be an integer

AREX-Base EXL3 4.1 bpw

Quality-oriented EXL3 conversion of BAAI/AREX-Base, a Qwen3.5-122B-A10B multimodal MoE research-agent model.

Quantization

  • EXL3 decoder target: 4.10 bpw
  • Effective HQ allocation: 4.14 bpw in the generated configuration
  • Sensitive MoE and attention tensors: higher-bit allocation via --hq
  • LM head: 6 bpw
  • Codebook: MCG
  • Calibration: 250 rows × 2,048 tokens
  • Output-channel scales: always enabled
  • Vision tower: preserved in BF16
  • MTP head: official Qwen3.5-122B-A10B BF16 donor, quantized at Q8

The conversion used ExLlamaV3 1.2.1. Use a recent ExLlamaV3-compatible runtime with Qwen3.5 MoE and multimodal support, such as a compatible TabbyAPI build.

Notes

This repository contains quantized weights, not benchmark results. Validate the model on your own prompts and serving configuration. Refer to the original model card for intended use, architecture details, limitations, and citation information.

The published AREX checkpoint advertises one MTP layer in its configuration but contains no mtp.* tensors. This conversion restores the architecture- compatible MTP head from the official Qwen3.5-122B-A10B base. Since AREX fine-tuned the target model without publishing a matching fine-tuned draft head, MTP acceptance may differ from the base model and should be measured on representative prompts.

License

The source model is released under Apache-2.0. See the included LICENSE and the original model repository.

Downloads last month
17
Safetensors
Model size
34B params
Tensor type
BF16
·
F16
·
I16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Dampish/AREX-Base-exl3-4.1bpw

Finetuned
BAAI/AREX-Base
Quantized
(3)
this model