How to use from
Docker Model Runner
docker model run hf.co/dots-studio/dots3-note-prev-fp8
Quick Links

中文 | English


dots logo

dots3-note Preview

GitHub: studio-dots-ai Transformers: dots3-note SGLang: dots3-note vLLM: dots3-note ModelScope: dots-studio

Dots Studio Discord X: dotsstudioai License: Apache 2.0

🌐 Tech Blog  |   📄 Full Report (coming soon)


Table of Contents


Model Introduction

dots3-note preview is the first open-weight model in the dots3 family. It is a Mixture-of-Experts model with 280B total parameters, 16B activated parameters, and support for a context length of up to 512K tokens. The model can understand text, images, video, and audio, and produces text outputs.

dots3-note preview is optimized for a broad range of tasks, including:

  • General knowledge and instruction following;
  • Mathematical and logical reasoning;
  • Tool use and multi-step agent workflows;
  • Interactive tasks that require exploration, memory updates, and adaptation;
  • Code generation and code-based problem solving;
  • Image, document, chart, audio, and video understanding;
  • Long-context information processing.

The dots3 family is designed to include models with different trade-offs among capability, latency, and inference cost. dots3-note preview is the most lightweight member of the family.

Model Overview

Property Value
Architecture Multimodal MoE
Total Parameters 280B
Activated Parameters 16B
MTP 1 shared layer, 1.13B
Number of Layers 1 dense + 45 MoE
Hidden Size 5120
FFN Hidden Size 13824 (dense), 1536 (per expert)
Experts 256 routed + 1 shared, top-8
Attention 13 DSA + 33 SWA (~1:3)
DSA Top-2048
Context Length 512K
Vocabulary Size 152K
Vision Encoder MoE ViT, 7B total, 1.2B activated
Audio Encoder Dense, 800M
Supported Precision BF16, FP8
Input Text, image, video, audio
Output Text

Evaluation Results

General Reasoning and Agent

General Reasoning and Agent evaluation results

Multimodal Understanding

Multimodal Understanding evaluation results

Model Links

Model Name Description HuggingFace ModelScope
dots3-note-prev Preview multimodal model 🤗 Model ModelScope Model
dots3-note-prev-fp8 FP8-quantized preview multimodal model 🤗 Model ModelScope Model

Quickstart

Recommended: serve the FP8 checkpoint on one 8-GPU node with SGLang or vLLM.

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="EMPTY")

response = client.chat.completions.create(
    model="dots3-note-prev",
    messages=[
        {"role": "user", "content": "Hello! Can you briefly introduce yourself?"},
    ],
    temperature=1.0,
    top_p=0.95,
    max_tokens=256,
    # Set enable_thinking=True for reasoning; False returns a direct response.
    extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
print(response.choices[0].message.content)

For a multimodal request, replace messages with one of these public examples:

examples = {
    "image": [
        {"type": "image_url", "image_url": {"url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/cats.png"}},
        {"type": "text", "text": "How many cats are in this image?"},
    ],
    "audio": [
        {"type": "audio_url", "audio_url": {"url": "https://huggingface.co/datasets/hf-internal-testing/dummy-audio-samples/resolve/main/mary_had_lamb.mp3"}},
        {"type": "text", "text": "Transcribe this nursery rhyme."},
    ],
    "video": [
        {"type": "video_url", "video_url": {"url": "https://huggingface.co/datasets/merve/vlm_test_images/resolve/main/concert.mp4"}},
        {"type": "text", "text": "Describe the performance and what can be heard."},
    ],
}
messages = [{"role": "user", "content": examples["image"]}]

Video inputs include their audio track when available.

Deployment

The commands below target FP8 on one 8-GPU node. BF16 requires more memory. Tune the context length to available memory, concurrency, and input modalities.

Native support is available on vLLM main. Transformers #47844 and SGLang #33829 are still under review; until they are merged, use the PR revisions below.

Transformers

First install mutually compatible PyTorch and torchvision builds supported by your NVIDIA driver. For audio and video, also install a PyTorch-compatible torchcodec (included below) and FFmpeg with your system package manager. Then install Transformers #47844:

pip install accelerate pillow torchcodec kernels==0.16.0 "transformers @ git+https://github.com/huggingface/transformers.git@refs/pull/47844/head"

Run a minimal local inference:

from transformers import AutoModelForMultimodalLM, AutoProcessor

model_id = "dots-studio/dots3-note-prev-fp8"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForMultimodalLM.from_pretrained(model_id, dtype="auto", device_map="auto")

messages = [
    {"role": "user", "content": "Hello! Please briefly introduce yourself."},
]
inputs = processor.tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    return_tensors="pt",
    return_dict=True,
    enable_thinking=False,
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=128)
print(processor.decode(outputs[0, inputs.input_ids.shape[1] :], skip_special_tokens=True))

Use SGLang or vLLM for multi-GPU OpenAI-compatible serving.

SGLang

Recommended: use the release image lmsysorg/sglang:dev-dots3-note. Full one-node recipes and tuning notes are in the Dots3-Note cookbook. Source support is tracked in SGLang #33829.

Docker (the image downloads the checkpoint from Hugging Face on first run):

docker run --gpus all --ipc=host -p 8000:8000 \
  lmsysorg/sglang:dev-dots3-note \
  sglang serve \
    --model-path dots-studio/dots3-note-prev-fp8 \
    --served-model-name dots3-note-prev \
    --host 0.0.0.0 \
    --port 8000 \
    --context-length 524288 \
    --enable-dp-attention \
    --dp-size 8 \
    --tp-size 8 \
    --ep-size 8 \
    --moe-dense-tp-size 1 \
    --page-size 64 \
    --trust-remote-code \
    --attention-backend fa3 \
    --moe-a2a-backend deepep \
    --enable-multimodal \
    --speculative-algorithm NEXTN \
    --speculative-num-steps 3 \
    --speculative-eagle-topk 1 \
    --speculative-num-draft-tokens 4 \
    --speculative-draft-model-path dots-studio/dots3-note-prev-fp8

Or install from source / the PR and run the same sglang serve arguments locally. --attention-backend fa3 sets prefill, decode, and (when speculative decoding is enabled) draft attention. MTP/NEXTN (--speculative-algorithm NEXTN and the related flags) is optional and can reduce TPOT by more than 50%. Prefill CUDA graph is not supported yet.

Optional features:

# Load only the language model
--language-only

# Enable OpenAI-compatible tool calling
--tool-call-parser dots

vLLM

Native dots3-note preview support is available on vLLM main. Use a recent nightly build until it is included in a stable release.

The following example deploys the FP8 checkpoint on eight NVIDIA H100 GPUs with TP=8 and EP=8:

vllm serve dots-studio/dots3-note-prev-fp8 \
  --served-model-name dots3-note-prev \
  --host 0.0.0.0 \
  --tensor-parallel-size 8 \
  --enable-expert-parallel \
  --moe-backend deep_gemm \
  --max-model-len 262144

Optional features:

# Load only the language model
--language-model-only

# Enable three-token MTP speculative decoding
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'

# Enable OpenAI-compatible automatic tool calling
--enable-auto-tool-choice --tool-call-parser dots

Benchmark Appendix

General Reasoning and Agent benchmark appendix

Multimodal benchmark appendix

License

Copyright (c) 2026 Xiaohongshu.

Developed and released by dots studio.

The dots3-note preview model weights and modeling code in this repository are licensed under the Apache License, Version 2.0.

See the LICENSE file for details.

Transformers, SGLang, vLLM, and other third-party software are subject to their respective licenses.

Contact Us

For questions and feedback, please contact us through:


dots3-note preview is developed and released by dots studio.

Downloads last month
34
Safetensors
Model size
288B params
Tensor type
BF16
·
F8_E4M3
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including dots-studio/dots3-note-prev-fp8