AREX-2-FP8

AREX-2-FP8 is an FP8 (W8A8) dynamic-quantized variant of BAAI/AREX-2, a 27B long-horizon, self-improving agent model built on a Qwen3.8-compatible multimodal dense architecture. Compression was performed with LLM Compressor using the FP8_DYNAMIC scheme, which quantizes weights to FP8 per-channel and activations to FP8 per-token at runtime, so no calibration data is required. The result reduces memory footprint and improves inference throughput on FP8-capable GPUs, with the checkpoint stored in the compressed-tensors format for direct loading in vLLM. For model architecture, training details, intended use, evaluation results, and the citation, refer to the original model card: BAAI/AREX-2.

Model Overview

Attribute Value
Base model BAAI/AREX-2 (derived from Qwen/Qwen3.8-27B)
Quantized model prithivMLmods/AREX-2-FP8
Parameters 27B
Context length 262,144 tokens
Modality Image-Text-to-Text
Quantization scheme FP8_DYNAMIC (W8A8)
Weights FP8, per-channel
Activations FP8, per-token, dynamic
Calibration data Not required
Format compressed-tensors (safetensors)
Tooling LLM Compressor
License Apache 2.0

Quantization Details

Only Linear layers are quantized. The following modules are excluded and remain in their original precision to preserve output quality: the language-model head, the token embeddings, the entire vision tower, and all linear-attention layers.

recipe.yaml

default_stage:
  default_modifiers:
    QuantizationModifier:
      targets: [Linear]
      ignore: ['re:.*lm_head', 're:.*embed_tokens$', 're:.*visual.*', 're:.*model.visual.*',
        're:.*linear_attn.*']
      scheme: FP8_DYNAMIC
      bypass_divisibility_checks: false
      requires_calibration_data: false
Setting Value
targets Linear
ignore lm_head, embed_tokens, visual, model.visual, linear_attn
scheme FP8_DYNAMIC
bypass_divisibility_checks false
requires_calibration_data false

Inference with vLLM

Install a recent vLLM release with Qwen3.8 support.

pip install -U vllm

Serve

vllm serve prithivMLmods/AREX-2-FP8 \
  --max-model-len 32768 \
  --trust-remote-code

Raise --max-model-len toward the native 262,144 tokens as GPU memory allows. Add --tensor-parallel-size N for multi-GPU setups.

Query the OpenAI-compatible endpoint

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

response = client.chat.completions.create(
    model="prithivMLmods/AREX-2-FP8",
    messages=[
        {
            "role": "user",
            "content": "Propose a solution and explain how you would improve it over several rounds.",
        }
    ],
    max_tokens=1024,
    seed=42,
)
print(response.choices[0].message.content)

Offline inference

from vllm import LLM, SamplingParams

llm = LLM(
    model="prithivMLmods/AREX-2-FP8",
    max_model_len=32768,
    trust_remote_code=True,
)

params = SamplingParams(max_tokens=1024, seed=42)
messages = [{
    "role": "user",
    "content": "Propose a solution and explain how you would improve it over several rounds.",
}]

outputs = llm.chat(messages, params)
print(outputs[0].outputs[0].text)

Reproducing the Quantization

pip install -U llmcompressor transformers accelerate torch
import torch
from transformers import AutoModelForMultimodalLM, AutoProcessor
from llmcompressor import oneshot

model_id = "BAAI/AREX-2"
save_dir = "AREX-2-FP8"

model = AutoModelForMultimodalLM.from_pretrained(
    model_id, dtype=torch.bfloat16, device_map="auto"
)
processor = AutoProcessor.from_pretrained(model_id)

oneshot(model=model, recipe="recipe.yaml")

model.save_pretrained(save_dir, save_compressed=True)
processor.save_pretrained(save_dir)

Notes

  • FP8 hardware support is required for full speed benefits (NVIDIA Hopper, Ada Lovelace, or newer). Other GPUs may fall back to weight-only behavior with reduced gains.
  • Quantization can introduce small deviations from the BF16 reference. Benchmark scores reported on the BAAI/AREX-2 card apply to the original weights and were not re-measured for this checkpoint.
  • Follow the license terms and notices of the original AREX-2 release and the Qwen base model.

Acknowledgements

Downloads last month
-
Safetensors
Model size
27B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/AREX-2-FP8

Base model

Qwen/Qwen3.8-27B
Finetuned
BAAI/AREX-2
Quantized
(3)
this model

Collection including prithivMLmods/AREX-2-FP8