Instructions to use BAAI/AREX-Base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use BAAI/AREX-Base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="BAAI/AREX-Base") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("BAAI/AREX-Base") model = AutoModelForMultimodalLM.from_pretrained("BAAI/AREX-Base", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use BAAI/AREX-Base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "BAAI/AREX-Base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BAAI/AREX-Base", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/BAAI/AREX-Base
- SGLang
How to use BAAI/AREX-Base with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "BAAI/AREX-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BAAI/AREX-Base", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "BAAI/AREX-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BAAI/AREX-Base", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use BAAI/AREX-Base with Docker Model Runner:
docker model run hf.co/BAAI/AREX-Base
AREX-Base Inference
This folder provides a minimal one-turn inference example and the complete BrowseComp prompts. It follows the XML tool-call protocol used by the public AREX evaluation code.
Serve the model
Run the following commands from the model repository root. Recent versions of vLLM, SGLang, or another OpenAI-compatible server with Qwen3.5 support can be used. For a text-only vLLM deployment:
vllm serve . \
--served-model-name AREX-Base \
--tensor-parallel-size 8 \
--max-model-len 262144 \
--reasoning-parser qwen3 \
--language-model-only
The example uses eight-way tensor parallelism as a starting point. Adjust the parallelism and maximum context length for your hardware.
Run one generation
Install the client:
pip install -U openai
Then send a BrowseComp-style question:
export AREX_BASE_URL="http://127.0.0.1:8000/v1"
export AREX_API_KEY="EMPTY"
export AREX_MODEL="AREX-Base"
python inference/inference.py \
--question "Your BrowseComp question"
The script returns the model's next action. When it emits an XML <tool_call>, execute that tool, append the assistant output to the message history, and add the real tool result as:
<tool_response>
actual tool result
</tool_response>
Continue until the model calls finish. The example intentionally leaves tool execution to the caller.
Use the prompts directly
prompts.py exports the BrowseComp system and user prompt constants. Tool descriptions are already embedded in the system prompt, so only the question needs formatting:
from inference.prompts import (
BROWSECOMP_SYSTEM_PROMPT,
BROWSECOMP_USER_PROMPT,
)
question = "Your BrowseComp question"
messages = [
{"role": "system", "content": BROWSECOMP_SYSTEM_PROMPT},
{
"role": "user",
"content": BROWSECOMP_USER_PROMPT.format(question=question),
},
]
build_messages(question) is a convenience wrapper for the same formatting.
BrowseComp exposes search, google_scholar, visit, update_context, and finish.