Instructions to use BAAI/AREX-2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use BAAI/AREX-2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="BAAI/AREX-2") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("BAAI/AREX-2") model = AutoModelForMultimodalLM.from_pretrained("BAAI/AREX-2", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use BAAI/AREX-2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "BAAI/AREX-2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BAAI/AREX-2", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/BAAI/AREX-2
- SGLang
How to use BAAI/AREX-2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "BAAI/AREX-2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BAAI/AREX-2", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "BAAI/AREX-2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BAAI/AREX-2", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use BAAI/AREX-2 with Docker Model Runner:
docker model run hf.co/BAAI/AREX-2
why Qwen 3.8 27B is absent in the benchmarks in Small models (β€40B) category?
Let me start by saying that it seems an incredible improvement over the base model!
But we can't really tell so much since Qwen 3.8 27B is not part of the benchmarks comparisons, what is the reason for this?
This is surprisingly common among models that are finetuned on Qwen 3.8 27B. The most significant comparison they can show is how it compares against the base model that it was finetuned on, but they never do. It's very annoying and suspicious.
I ran a small local comparison while converting AREX-2 to MLX, in case it helps here.
Short version: AREX-2 was more efficient than plain Qwen3.8-27B. It used 39% to 53% fewer tokens in every run, finished in about half the time, and solved a few more problems.
| AREX-2 | Qwen3.8-27B | |
|---|---|---|
| Tokens used, all tests | 113,795 | 212,514 |
| Time, all tests | 249 min | 451 min |
| Easy coding tasks passed | 35 of 36 | 30 of 36 |
| Hard problems solved first try | 15 of 24 | 11 of 24 |
| Hard problems solved within 3 tries | 23 of 24 | 20 of 24 |
Both models are 8-bit MLX with the same prompts, sampling (temperature 1.0, top-p 0.95, top-k 20), seeds and a 4,000-token limit, on one M4 Pro Mac mini. On the hard problems the model sees its failing test and gets up to three tries.
The difference was mostly length. Every failed attempt by plain Qwen3.8-27B ran into the 4,000-token limit (31 of 31), against 9 of 13 for AREX-2. With a higher limit the plain model might well catch up on problems solved, just more slowly. Both ran with thinking on, the default.
With thinking off the two were level. In an agent harness (MiniMax Code), each fixed three small failing projects in 6 of 6 trials, in 17 and 18 minutes. So the gap above comes from thinking.
This is a small test: two test sets, two runs each, short single tasks. It does not touch the long multi-round benchmarks in the paper, so it cannot confirm or dispute those.
MLX builds for Mac (4, 5, 6 and 8-bit) with the full results are here: https://huggingface.co/mlx-community/AREX-2-5bit