AREX

AREX-2 27B · INT4 W4A16

Symmetric group-128 weight-only quantization for vLLM.

Base model · Paper · Project · vLLM

Format W4A16 Weights INT4 Activations BF16 License Apache 2.0

This is a numerical quantization of BAAI/AREX-2, not a fine-tune. It was exported from official revision d4e3502f92d9e889031c2d04e387ab2eb3520268. Model, architecture, training-data, and upstream evaluation credit belongs to BAAI and the upstream contributors. Results from the upstream BF16 model do not measure this INT4 checkpoint.

Model summary

AREX-2 is a 27B-parameter, Qwen3.8-compatible multimodal model for long-horizon agent tasks. The upstream checkpoint declares a native context length of 262,144 tokens. Refer to the upstream model card for architecture, capabilities, evaluation protocols, and usage guidance.

Dual RTX 3090 deployment

Validation scope: the checkpoint loaded and passed a short text-chat smoke test in the pinned vLLM 0.29.0 deployment image on 2×RTX 3090, and completed the synthetic serving matrix below. No task-quality or multimodal evaluation was run; the native 262,144-token limit was not tested.

Runtime property Tested profile
GPUs 2×RTX 3090 (Ampere sm_86)
Tensor parallelism 2
Activation dtype BF16
KV cache FP8 E4M3
CPU KV offload 28 GiB, using the deployment's vLLM backport
Maximum model length 262,144 configured; tested up to 128,003 prompt tokens + 1,024 generated tokens
Maximum active sequences 1
Optional speculation External DFlash2 W4A16 drafter; not included in this repository

Memory and capacity

Metric Result
Indexed tensor payload 18.60 GB / 17.33 GiB
vLLM GPU memory after short probes 16,784 MiB used and 7,394 MiB free per GPU
Long-context capacity 128K prompt requests benchmarked; the configured 262,144-token limit was not tested

GPU memory is the observed full serving-process allocation, not model weights alone. The longest benchmark request had 128,003 prompt tokens and 1,024 generated tokens; this does not establish capacity at the configured 262,144-token limit.

Measured serving performance

The matrix used synthetic, salted prompts at approximately 1K, 8K, 32K, 64K, and 128K tokens, with exact 512- or 1,024-token greedy completions. There were two cold-prefix runs per cell; TTFT is time to first streamed token, and prompt usage comes from the server. The table reports medians; the decode column also shows the two-run range. DFlash2 remained enabled (7 draft tokens, 15-token verify ceiling, lookup and chains on), with the production KV-cache and CPU-offload profile.

Prompt tokens New tokens Median TTFT (s) Prefill (tok/s) Decode median (tok/s; min–max)
1,027 512 0.69 1479 225.6 (191.3–260.0)
1,028 1,024 0.71 1450 149.0 (88.0–209.9)
8,194 512 5.73 1430 109.8 (75.9–143.7)
8,196 1,024 5.79 1416 266.2 (253.9–278.4)
32,002 512 25.17 1273 66.4 (65.2–67.7)
32,002 1,024 28.71 1115 125.0 (63.3–186.8)
64,002 512 74.38 861 78.5 (56.7–100.3)
64,002 1,024 77.99 821 89.8 (57.6–122.0)
128,002 512 180.24 710 42.6 (39.3–45.9)
128,003 1,024 183.37 698 84.0 (81.6–86.4)

Quantization fidelity

This export uses data-free, symmetric round-to-nearest (RTN) quantization. The aggregate relative L2 reconstruction error over the quantized weights is 0.12160. This is a weight-space diagnostic only; it is not perplexity, logit/KLD agreement, or a task-quality score. No calibration dataset, behavioral evaluation, or quality comparison against the BF16 source has been run.

Checkpoint profile

Property Value
Base checkpoint BAAI/AREX-2, revision d4e3502f92d9e889031c2d04e387ab2eb3520268
Quantization Symmetric INT4 W4A16, group size 128, RTN; no calibration
Quantized modules 400 two-dimensional Linear weights
Runtime format compressed-tensors / pack-quantized
Kernel dispatch CompressedTensorsWNA16 → MarlinLinearKernel in the tested runtime
Scale storage FP16
Preserved source precision Vision tower, token embeddings, output head, recurrent GDN gates, and other excluded/non-Linear tensors
Runtime vLLM with compressed-tensors support; this is not a GGUF checkpoint

The quantization ignore rules are inherited from the validated W8A16 profile. The AREX-2 source checkpoint contains no MTP tensors; any speculative-decoding drafter must be supplied separately.

Quantization design

Component Treatment
400 eligible two-dimensional Linear weights Symmetric INT4, group size 128, round-to-nearest
Vision tower, embeddings, output head, recurrent GDN gates Retained at source precision
Other non-Linear and excluded tensors Retained at source precision

The export uses the same exact eligible-module and ignore profile as the local W8A16 reference. It does not fine-tune or calibrate the model.

Why W4A16

W4A16 stores the selected Linear weights in 4-bit integers with group-wise scales while keeping activations at 16-bit precision. The smaller tensor payload does not guarantee higher inference speed or equivalent quality. Benchmark on the intended hardware and workload before choosing this variant for production.

Operational notes

Question Guidance
Which runtime is validated? The pinned vLLM 0.29.0 deployment image used for the profile above. Other versions and kernels need separate validation.
Can llama.cpp load this checkpoint? No. It uses the vLLM compressed-tensors format, not GGUF.
Is vision validated? The source-precision vision weights are retained, but the smoke test was text-only; multimodal quality and memory use are unmeasured.
Is 262K serving validated? No. The source model's context setting is retained, but long-context capacity and quality were not tested.
Is DFlash2 included? No. The deployment smoke used an external drafter; this model repository contains only AREX-2 checkpoint assets.

Files and provenance

The repository contains the sharded SafeTensors checkpoint, model.safetensors.index.json, model/tokenizer configuration, and the upstream Apache-2.0 LICENSE. Quantized tensors are represented by packed weights, per-group scales, and original weight shapes in the index.

Acknowledgements and license

This repository repackages numerical weights derived from BAAI/AREX-2. It does not claim authorship of the base model, architecture, training data, or upstream evaluations. The upstream Apache-2.0 license is included; retain the source attribution and follow the upstream terms.

Downloads last month
399
Safetensors
Model size
27B params
Tensor type
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for numsu/AREX-2-27B-INT4-W4A16

Base model

Qwen/Qwen3.8-27B
Finetuned
BAAI/AREX-2
Quantized
(14)
this model

Paper for numsu/AREX-2-27B-INT4-W4A16