| # Recipe |
|
|
| Everything needed to serve, tune, and benchmark this model — in order. Each step |
| says what to run, what you should see, and what to do when it goes wrong. |
|
|
| **Time:** ~20 minutes to a working server, plus ~1 hour for the full benchmark suite. |
|
|
| --- |
|
|
| ## What you need |
|
|
| | | Minimum used here | |
| |---|---| |
| | GPUs | 2× NVIDIA RTX PRO 6000 Blackwell (96 GB, SM120) | |
| | VRAM | ~193 GB total across both cards | |
| | Disk | ~400 GB for the checkpoint | |
| | Software | vLLM with DeepSeek-V4 sparse-MLA support, CUDA 13.x | |
|
|
| Less VRAM? Lower `--max-model-len` from 262144 — that is the main memory lever, |
| since the KV dtype cannot be reduced on this architecture (see Step 6). |
|
|
| --- |
|
|
| ## Step 1 — Serve the model |
|
|
| ```bash |
| vllm serve /path/to/model \ |
| --served-model-name dsv4-hybrid-vision dsv4-hybrid \ |
| --host 127.0.0.1 --port 8000 \ |
| --tensor-parallel-size 2 \ |
| --tokenizer-mode deepseek_v4 \ |
| --kv-cache-dtype fp8_ds_mla \ |
| --block-size 256 \ |
| --max-model-len 262144 \ |
| --max-num-seqs 16 \ |
| --max-num-batched-tokens 8192 \ |
| --gpu-memory-utilization 0.968 \ |
| --speculative-config '{"method":"dspark","model":"/path/to/model","num_speculative_tokens":5}' \ |
| --compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY","max_cudagraph_capture_size":8}' \ |
| --limit-mm-per-prompt '{"vision_chunk":8}' \ |
| --enable-auto-tool-choice \ |
| --tool-call-parser deepseek_v4 \ |
| --generation-config vllm \ |
| --default-chat-template-kwargs '{"thinking": false, "enable_thinking": false}' \ |
| --override-generation-config '{"max_new_tokens": 1024, "temperature": 0.6, "top_p": 0.95}' |
| ``` |
|
|
| **Expect:** ~3 minutes to load. Success looks like: |
|
|
| ``` |
| GPU KV cache size: 263,310 tokens |
| Application startup complete |
| ``` |
|
|
| **Check it:** |
|
|
| ```bash |
| curl -s http://127.0.0.1:8000/v1/models | python3 -m json.tool |
| ``` |
|
|
| ### Why these flags |
|
|
| | Flag | Reason | |
| |---|---| |
| | `--tokenizer-mode deepseek_v4` | Required. This checkpoint ships **no** `chat_template`; turn formatting is built in Python by the `deepseek_v4` tokenizer, which never evaluates Jinja. | |
| | `--kv-cache-dtype fp8_ds_mla` | The only working option on SM120 (Step 6). | |
| | `--max-num-seqs 16` | 4 capped throughput at 227 tok/s with 3.3 s tail latency; 16 gives 486 tok/s at 1.74 s. KV cache usage stayed under 1%, so the cap was a scheduler limit, not memory. | |
| | `--override-generation-config` | With `--generation-config vllm`, the checkpoint's own `generation_config.json` is **ignored** and you inherit vLLM's `temperature=1.0`. This sets sane defaults (Step 3). | |
| | `--default-chat-template-kwargs` | Pins chain-of-thought off so a stray client kwarg cannot enable it server-wide. | |
|
|
| --- |
|
|
| ## Step 2 — Warm up the JIT kernels |
|
|
| vLLM compiles Triton/TileLang kernels the first time it sees each tensor shape, |
| **inside whichever request happens to arrive**. Its own log says so: |
|
|
| ``` |
| WARNING [jit_monitor] Triton kernel JIT compilation during inference: |
| _compute_prefill_metadata_kernel. This causes a latency spike; |
| consider extending warmup to cover this shape/config. |
| ``` |
|
|
| Measured: **12 events in the first 14 minutes**, each ~6 s against an 0.086 s |
| median. Issue one request per shape class right after startup instead: |
|
|
| ```bash |
| #!/usr/bin/env bash |
| # warmup.sh -- run once after the server is ready |
| URL=http://127.0.0.1:8000/v1/chat/completions |
| warm() { # $1 = filler word count, $2 = max_tokens, $3 = label |
| local filler; filler=$(python3 -c "print('word '*$1)") |
| curl -s -m 600 "$URL" -H 'Content-Type: application/json' \ |
| -d "$(python3 -c " |
| import json |
| print(json.dumps({'model':'dsv4-hybrid-vision','max_tokens':$2, |
| 'messages':[{'role':'user','content':'Reply ok. '+'''$filler'''}]}))")" -o /dev/null |
| echo " warmed $3" |
| } |
| warm 5 16 tiny |
| warm 100 16 short |
| warm 800 16 medium |
| warm 3000 16 long |
| warm 10000 16 xlong |
| warm 20 256 decode-deep |
| ``` |
|
|
| Also send one 448 px image and one tool-calling request — those are separate code |
| paths with their own kernels. |
|
|
| **Expect:** ~2 s once the cache is populated; longer the first time. Compiled |
| kernels persist in vLLM's JIT cache directory, so mount it on durable storage to |
| keep the benefit across restarts. |
|
|
| **Verify it worked:** every `jit_monitor` line in the server log should fall |
| inside your warmup window, not during later user traffic. |
|
|
| --- |
|
|
| ## Step 3 — Fix verbose output |
|
|
| **Symptom:** ask for one bash one-liner, get five variants plus a bullet-point |
| explainer — 347 tokens for a one-line question. |
|
|
| **Cause, in two parts:** |
|
|
| 1. **No system prompt.** The checkpoint ships no chat template, so nothing steers |
| response length. |
| 2. **Sampling.** `--generation-config vllm` ignores the checkpoint's |
| `generation_config.json`, leaving `temperature=1.0`, `top_p=1.0`. |
|
|
| **Fix part 1** — sampling, already in the Step 1 command: |
|
|
| ``` |
| --override-generation-config '{"max_new_tokens": 1024, "temperature": 0.6, "top_p": 0.95}' |
| ``` |
|
|
| **Fix part 2** — a default system prompt. vLLM has **no flag** for this, and the |
| usual `--chat-template` trick does not apply because `--tokenizer-mode deepseek_v4` |
| never evaluates Jinja. Send one per request: |
|
|
| ```python |
| SYSTEM = """Be concise and direct. Answer what was asked, then stop. |
| |
| - Give ONE best answer, not a menu of alternatives. Compare options only if asked. |
| - No preamble, no restating the question, no closing summary or recap. |
| - Do not explain self-evident code. Note only genuinely surprising behaviour. |
| - Prefer a sentence over a paragraph, a paragraph over a list. |
| - Match length to the question: a one-line question gets a one-line answer.""" |
| |
| messages = [{"role": "system", "content": SYSTEM}] + user_messages |
| ``` |
|
|
| For a fleet, put a small proxy in front that injects this when the caller supplies |
| no system message of their own. |
|
|
| **Result:** |
|
|
| | Prompt | Before | After | |
| |---|---|---| |
| | "bash one-liner, 10 largest files under /var/log" | 347 tokens, 5 variants | **29 tokens**, 1 command | |
| | "What port does SSH use?" | 54 tokens | **3 tokens** (`22.`) | |
|
|
| --- |
|
|
| ## Step 4 — Cap image resolution at 448 px |
|
|
| **Do this before sending any image.** Above ~450 px the model drifts toward |
| answering "yes" to everything and invents text that is not in the picture. |
|
|
| ```python |
| from PIL import Image |
| import base64, io |
| |
| def encode_image(path, max_side=448): |
| im = Image.open(path).convert("RGB") |
| im.thumbnail((max_side, max_side)) # preserves aspect ratio |
| buf = io.BytesIO(); im.save(buf, "PNG") |
| return base64.b64encode(buf.getvalue()).decode() |
| ``` |
|
|
| **Measured effect** (balanced yes/no questions, equal true-yes and true-no): |
|
|
| | Image | Full resolution | Capped at 448 px | |
| |---|---|---| |
| | Desert landscape | 5/7 | **7/7** | |
| | Mountain at sunset | 5/7 | **6/7** | |
|
|
| It also cuts prompt tokens **12×** — a 4736×2656 photo drops from ~3,100 to ~260. |
|
|
| --- |
|
|
| ## Step 5 — Benchmark it |
|
|
| Install the harness: |
|
|
| ```bash |
| pip install lm-eval==0.4.12 transformers langdetect immutabledict |
| ``` |
|
|
| ### Multiple choice (loglikelihood — needs `/v1/completions`) |
|
|
| ```bash |
| lm_eval --model local-completions \ |
| --model_args "model=dsv4-hybrid,base_url=http://127.0.0.1:8000/v1/completions,\ |
| num_concurrent=16,max_retries=3,tokenized_requests=False,\ |
| tokenizer=/path/to/model,tokenizer_backend=huggingface" \ |
| --tasks arc_challenge,hellaswag,winogrande,piqa,openbookqa,truthfulqa_mc2 \ |
| --limit 300 --batch_size 16 --output_path eval_mc |
| ``` |
|
|
| ### Generative (chat endpoint) |
|
|
| ```bash |
| lm_eval --model local-chat-completions \ |
| --model_args "model=dsv4-hybrid,base_url=http://127.0.0.1:8000/v1/chat/completions,\ |
| num_concurrent=16,max_retries=3,tokenized_requests=False,\ |
| tokenizer=/path/to/model,tokenizer_backend=huggingface" \ |
| --tasks gsm8k,ifeval --limit 200 --apply_chat_template --output_path eval_gen |
| ``` |
|
|
| ### Code — use the custom script |
|
|
| ```bash |
| python3 benchmarks/humaneval_chat.py |
| ``` |
|
|
| **Do not trust `lm_eval --tasks humaneval` on a chat endpoint.** It reports |
| **0.0** because the task expects a raw completion while the model returns prose |
| plus a ```` ```python ```` fence. The script above extracts the fenced code and |
| runs the official tests: **78.0% pass@1**. |
| |
| ### Gotchas |
| |
| | Error | Fix | |
| |---|---| |
| | `got multiple values for keyword argument 'batch_size'` | `batch_size` is a CLI flag, not a `--model_args` key. | |
| | `No module named 'transformers'` | Install into the **same** environment as `lm_eval`. | |
| | `dsv4-hybrid is not a local folder...` | Pass `tokenizer=/path/to/model` — it defaults to resolving the served name on the Hub. | |
| | `No module named 'langdetect'` | `pip install langdetect immutabledict` (IFEval). | |
| | HumanEval scores 0.0 | Expected on a chat endpoint. Use `humaneval_chat.py`. | |
|
|
| --- |
|
|
| ## Step 6 — Do not bother with nvfp4 KV cache |
|
|
| `nvfp4_ds_mla` would cut the KV record from 656 to 432 B/token (~1.52× more |
| cache). Both the vLLM CLI and the b12x backend advertise it. **It does not work on |
| SM120** — the engine fails at load: |
|
|
| ``` |
| AssertionError: DeepseekV4 fp8_ds_mla layout only supports fp8 kv-cache, |
| got nvfp4_ds_mla |
| ``` |
|
|
| Three independent blockers, each verified: |
|
|
| | Layer | Blocker | |
| |---|---| |
| | vLLM Python | Every DSv4 attention class reachable on SM120 sets `use_fp8_ds_mla_layout = True`, which asserts `dtype.startswith("fp8")` | |
| | FlashInfer API | Only `trtllm_batch_decode_sparse_mla_dsv4` exists; no nvfp4 variant | |
| | CUDA kernel | Prebuilt cubin, documented as **584 B/token**, BF16 or FP8 E4M3 only | |
|
|
| Patching only the Python assertion would feed 432-byte records to a kernel reading |
| 584-byte records: **silent numerical corruption**, not a clean error. Do not do it. |
|
|
| Of the 17 dtypes the CLI advertises, six are accepted (`fp8`, `fp8_ds_mla`, |
| `fp8_e4m3`, `fp8_e5m2`, `fp8_inc`, `fp8_per_token_head`) and **all six resolve to |
| the same 656 B record** — none changes capacity. This needs upstream support. |
|
|
| --- |
|
|
| ## Step 7 — Verify before you ship |
|
|
| Quick checks that catch the common regressions: |
|
|
| ```bash |
| # 1. Terse output (expect ~3 tokens: "22.") |
| curl -s http://127.0.0.1:8000/v1/chat/completions -H 'Content-Type: application/json' \ |
| -d '{"model":"dsv4-hybrid","messages":[{"role":"user","content":"What port does SSH use?"}]}' \ |
| | python3 -c "import json,sys;d=json.load(sys.stdin);print(d['usage']['completion_tokens'],'tok:',d['choices'][0]['message']['content'])" |
| |
| # 2. Tool calling (expect finish_reason: tool_calls) |
| curl -s http://127.0.0.1:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{ |
| "model":"dsv4-hybrid","messages":[{"role":"user","content":"Weather in Paris?"}], |
| "tools":[{"type":"function","function":{"name":"get_weather","description":"Get weather", |
| "parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}}}]}' \ |
| | python3 -c "import json,sys;d=json.load(sys.stdin);print(d['choices'][0]['finish_reason'])" |
| |
| # 3. Streaming (expect several chunks then [DONE]) |
| curl -sN http://127.0.0.1:8000/v1/chat/completions -H 'Content-Type: application/json' \ |
| -d '{"model":"dsv4-hybrid","stream":true,"messages":[{"role":"user","content":"Count 1 to 5."}]}' \ |
| | grep -c '^data:' |
| ``` |
|
|
| **Benchmark tip:** when comparing runs, give every prompt a **unique prefix**. |
| vLLM's prefix cache is on by default (~86% hit rate in normal use); re-sending an |
| identical prompt returns in 0.11 s instead of 2.68 s. Measuring that reports |
| impossible numbers — 276,000 tok/s prefill in one early run here. A nonce in the |
| prompt prefix invalidates the cached sequence and restores honest timings. |
|
|