| --- |
| license: other |
| license_name: deepseek |
| license_link: https://github.com/deepseek-ai/DeepSeek-V2/blob/main/LICENSE-MODEL |
| language: |
| - en |
| pipeline_tag: image-text-to-text |
| library_name: vllm |
| tags: |
| - deepseek |
| - multimodal |
| - vision-language |
| - fp8 |
| - speculative-decoding |
| - vllm |
| - blackwell |
| --- |
| |
| # DeepSeek V4 Flash Hybrid Vision Reseek |
|
|
| A vision-enabled assembly of **DeepSeek-V4-Flash-0731**: the DeepSeek-V4 Flash text |
| model (48 shards, FP8) joined to a MoonViT-style vision tower and multimodal |
| projector, served through vLLM with **DSpark speculative decoding**. |
|
|
| Built and benchmarked on 2× NVIDIA RTX PRO 6000 Blackwell Workstation Edition |
| (96 GB each, SM120), tensor-parallel 2. |
|
|
| | | | |
| |---|---| |
| | Architecture | `DeepSeekV4VisionForConditionalGeneration` | |
| | Context length | 262,144 tokens | |
| | KV cache | `fp8_ds_mla` (656 B/token) | |
| | Vision encoder | MoonViT-style, patch 14, 2×2 merge | |
| | Speculative decoding | DSpark, 5 draft tokens | |
| | Tool calling | Yes (`deepseek_v4` parser) | |
|
|
| --- |
|
|
| ## Benchmarks |
|
|
| All numbers produced with **EleutherAI lm-evaluation-harness v0.4.12** against a |
| live vLLM endpoint. Raw result JSON is in [`benchmarks/`](./benchmarks). |
| Reproduction commands are in [`RECIPE.md`](./RECIPE.md). |
|
|
| ### Reasoning and instruction following |
|
|
| | Benchmark | Metric | Score | n | |
| |---|---|---|---| |
| | **GSM8K** (5-shot) | exact_match, flexible | **96.5%** | 200 | |
| | GSM8K (5-shot) | exact_match, strict | 69.0% | 200 | |
| | **IFEval** | inst-level loose | **84.0%** | 200 | |
| | IFEval | inst-level strict | 81.5% | 200 | |
| | IFEval | prompt-level loose | 76.5% | 200 | |
| | IFEval | prompt-level strict | 73.5% | 200 | |
|
|
| ### Knowledge and commonsense (0-shot, loglikelihood) |
|
|
| | Benchmark | Metric | Score | n | |
| |---|---|---|---| |
| | **PIQA** | acc_norm | **87.7%** | 300 | |
| | PIQA | acc | 83.0% | 300 | |
| | **Winogrande** | acc | **79.7%** | 300 | |
| | **HellaSwag** | acc_norm | **71.3%** | 300 | |
| | HellaSwag | acc | 57.0% | 300 | |
| | **ARC-Challenge** | acc_norm | **63.3%** | 300 | |
| | ARC-Challenge | acc | 62.3% | 300 | |
| | TruthfulQA MC2 | acc | 54.9% | 300 | |
| | OpenBookQA | acc_norm | 51.3% | 300 | |
|
|
| ### Code |
|
|
| | Benchmark | Metric | Score | n | |
| |---|---|---|---| |
| | **HumanEval** | pass@1 | **78.0%** | 164 | |
|
|
| > **Note on HumanEval.** The stock `lm-eval` HumanEval task scores this model at |
| > **0.0** — an artifact, not a result. That task expects a raw completion, but a |
| > chat endpoint returns prose plus a ```` ```python ```` block, so the built-in |
| > filter extracts nothing. The 78.0% above comes from |
| > [`benchmarks/humaneval_chat.py`](./benchmarks/humaneval_chat.py), which pulls the |
| > fenced code and runs the official unit tests in a sandboxed subprocess. If you |
| > benchmark any chat model on HumanEval, check for this first. |
|
|
| ### Throughput |
|
|
| Measured on 2× RTX PRO 6000 Blackwell, TP=2, `max_num_seqs=16`: |
|
|
| | Load | Aggregate throughput | Tail latency | |
| |---|---|---| |
| | 1 concurrent | 305 tok/s | 0.39 s | |
| | 8 concurrent | 581 tok/s | 1.64 s | |
| | 16 concurrent | 1150 tok/s | 1.66 s | |
|
|
| Time to first token: **0.085 s** median. Prefill: ~8,900 tok/s at 32k tokens. |
| DSpark speculative decoding accepts 3.2–5.9 of 5 draft tokens in practice. |
|
|
| --- |
|
|
| ## Vision: architecture, capabilities, limits |
|
|
| ### The adapter |
|
|
| **457M parameters total** — roughly 0.1% of the assembled model: |
|
|
| | Component | Params | Shape | |
| |---|---|---| |
| | Vision tower (MoonViT-style) | 416.9M | 27 layers, width 1152, patch 14 | |
| | Multimodal projector | 40.1M | `pre_norm → 4608×4608 → GELU → 4608×4096` | |
|
|
| Dimensions chain exactly: the tower emits 1152-wide patch embeddings, a 2×2 merge |
| concatenates them to 4608, and the projector maps that to 4096 — the language |
| model's hidden size. A 448×448 image becomes 256 patches → **64 tokens** after |
| merging, which is frugal next to tile-based encoders that spend thousands. |
|
|
| Preprocessing was verified numerically: the normalization LUT produces exactly |
| ±1.0 at grey 0 and 255, with correct 16×16 spatial ordering. |
|
|
| ### Strengths |
|
|
| - **Cheap.** ~0.8 GB VRAM, no measurable cost to text quality — all text |
| benchmarks score identically with the vision tower loaded. |
| - **Accurate within its design point.** 9/10 on balanced yes/no probes at 448 px; |
| 4/5 on 4-way forced choice with the correct answer listed last. It correctly |
| rejects absent objects (cat, car, people, text, bicycle), so it is not simply |
| answering "yes". |
| - **Token-efficient**, per the merge arithmetic above. |
|
|
| ### Weaknesses |
|
|
| - **Fabricates above ~450 px.** The dominant failure. Balanced accuracy falls from |
| 6/7 to 4/7 while "yes" answers rise from 4/7 to 6/7 — it drifts toward agreeable |
| confabulation rather than degrading gracefully. |
| - **Weak fine-grained recognition.** Category right, identity wrong: a numbat read |
| as "giraffe", a rainbow lightbulb as "camera lens". Scene structure survives; |
| species and object identity do not. |
| - **Blind to synthetic images.** Solid colour fills score 1/4 — chance. Expected |
| for a patch-based ViT, since a flat 14×14 patch carries no edges or texture, but |
| it means **synthetic images cannot be used to smoke-test the pipeline**. A flat |
| test image will look like total failure on a perfectly healthy deployment. |
| - **Thin projector.** A 2-layer MLP is the minimum viable bridge from a 1152-wide |
| tower into a 4096-wide LM, and is the most likely cause of the identity |
| confusions above. |
|
|
| ### Where this adapter could improve |
|
|
| 1. **Deeper or attention-based projector.** Replacing the 2-layer MLP with a |
| cross-attention resampler (Q-Former / Perceiver style) targets the identity |
| errors directly, and at 40M params it is by far the cheapest component to |
| retrain — the tower need not be touched. |
| 2. **Higher-resolution alignment training.** The >450 px collapse points at the |
| projector rather than the tower: token counts scale correctly with image area, |
| so the encoder handles more patches; it is the mapping that degrades. |
| 3. **Negative-example alignment.** The failure mode is confident invention, not |
| refusal. The model answers "is there text in this image?" correctly at 448 px, |
| so the capability exists but is not robust to resolution. |
| 4. **OCR evaluation.** Every fabrication observed involved invented *text* |
| ("Airbus", "The Great Forest"). This is a hypothesis, not a measurement — it |
| was not tested against an OCR benchmark. |
|
|
| > **Evaluation gap.** The vision figures below come from hand-built probes over |
| > roughly ten images, not a standard benchmark. The text scores in this card are |
| > `lm-eval` reproducible; **the vision claims are not**. Run MMMU, MMBench, or |
| > TextVQA before relying on them. |
|
|
| ### Measured behaviour |
|
|
| **What works** — images at or below **448 px** on the longest side: |
|
|
| | Test | Result | |
| |---|---| |
| | Balanced yes/no on real photos | 9/10 | |
| | 4-way forced choice (correct answer listed last) | 4/5 (chance 25%) | |
| | Detail questions (water present? snow-capped? time of day?) | correct | |
|
|
| **What fails:** |
|
|
| - **Open-ended captioning above ~450 px.** The model confabulates. Asked to |
| describe a desert landscape at 896 px it reported *"a bird flying over a sunset |
| with the words 'the great outdoors'"* — text that does not exist in the image. |
| Feeding a full-resolution photo produced an invented *"Airbus"* logo. |
| - **Synthetic flat-colour images.** Solid fills score at chance (1/4). Flat colour |
| is out of distribution for a patch-based ViT — no edges or texture in any patch. |
| This is expected behaviour, not a defect. |
|
|
| **Recommendations:** |
|
|
| 1. **Downscale images to 448 px** before sending. Measured effect: one test image |
| went from 5/7 to **7/7** correct. It also cuts a 4736×2656 photo from ~3,100 |
| prompt tokens to ~260 (**12× fewer**). |
| 2. **Ask specific questions** ("Is there water in this image?") rather than |
| "describe this image." Factual queries stay accurate where captions drift. |
|
|
| --- |
|
|
| ## Quick start |
|
|
| ### Serve with vLLM |
|
|
| ```bash |
| vllm serve /path/to/DeepSeek-V4-Flash-Hybrid-Vision-Reseek \ |
| --served-model-name dsv4-hybrid-vision \ |
| --tensor-parallel-size 2 \ |
| --tokenizer-mode deepseek_v4 \ |
| --kv-cache-dtype fp8_ds_mla \ |
| --block-size 256 \ |
| --max-model-len 262144 \ |
| --max-num-seqs 16 \ |
| --gpu-memory-utilization 0.968 \ |
| --speculative-config '{"method":"dspark","model":"/path/to/model","num_speculative_tokens":5}' \ |
| --enable-auto-tool-choice \ |
| --tool-call-parser deepseek_v4 |
| ``` |
|
|
| ### Query it |
|
|
| ```python |
| import base64, io, json, urllib.request |
| from PIL import Image |
| |
| URL = "http://127.0.0.1:8000/v1/chat/completions" |
| |
| def ask(messages, max_tokens=512): |
| body = {"model": "dsv4-hybrid-vision", "messages": messages, |
| "max_tokens": max_tokens, "temperature": 0.6, "top_p": 0.95} |
| req = urllib.request.Request(URL, data=json.dumps(body).encode(), |
| headers={"Content-Type": "application/json"}) |
| return json.loads(urllib.request.urlopen(req).read())["choices"][0]["message"]["content"] |
| |
| # Text |
| print(ask([{"role": "user", "content": "What port does SSH use?"}])) |
| |
| # Vision -- downscale to 448 px first (see the vision notes above) |
| im = Image.open("photo.jpg").convert("RGB") |
| im.thumbnail((448, 448)) |
| buf = io.BytesIO(); im.save(buf, "PNG") |
| b64 = base64.b64encode(buf.getvalue()).decode() |
| |
| print(ask([{"role": "user", "content": [ |
| {"type": "text", "text": "Is there a mountain in this image? Answer yes or no."}, |
| {"type": "image_url", "image_url": {"url": "data:image/png;base64," + b64}}, |
| ]}])) |
| ``` |
|
|
| --- |
|
|
| ## Recommended sampling |
|
|
| The checkpoint's `generation_config.json` defaults to `temperature=1.0`, |
| `top_p=1.0`. At those settings the model **pads its answers heavily** — asked for |
| a single bash one-liner it returned five variants plus a bullet-point explainer |
| (347 tokens for a one-line question). |
|
|
| These settings were measured to fix that: |
|
|
| ```json |
| {"temperature": 0.6, "top_p": 0.95, "max_tokens": 1024} |
| ``` |
|
|
| Pair them with a concise system prompt for short answers. With both applied, the |
| same one-liner question returns **29 tokens** — one command, no menu. |
|
|
| Note that `--generation-config vllm` makes vLLM **ignore** the checkpoint's |
| `generation_config.json`; pass `--override-generation-config` to set defaults |
| server-side. |
|
|
| --- |
|
|
| ## Reasoning mode |
|
|
| Chain-of-thought is **off by default** — the `deepseek_v4` tokenizer sets |
| `thinking=False` unless a caller opts in with `enable_thinking` or |
| `reasoning_effort`. To pin it off server-side regardless of what clients send: |
|
|
| ```bash |
| --default-chat-template-kwargs '{"thinking": false, "enable_thinking": false}' |
| ``` |
|
|
| --- |
|
|
| ## Known constraints |
|
|
| - **KV cache dtype is fixed at `fp8_ds_mla` on SM120.** `nvfp4_ds_mla` would cut |
| the record from 656 to 432 B/token (~1.52× more cache), and both the CLI and the |
| b12x backend advertise it — but every DeepSeek-V4 attention class reachable on |
| SM120 sets `use_fp8_ds_mla_layout = True`, which asserts the dtype starts with |
| `fp8`. Attempting it fails at load: |
| `AssertionError: DeepseekV4 fp8_ds_mla layout only supports fp8 kv-cache`. |
| The one class with the flag off is gated to compute capability 9/10 (Hopper). |
| The underlying FlashInfer kernel is a prebuilt cubin documented as 584 B/token, |
| BF16 or FP8 E4M3 only, so this needs upstream support — not a config change. |
| - **JIT warmup.** vLLM compiles Triton/TileLang kernels on first sight of each |
| tensor shape and logs *"causes a latency spike"*. Measured: 12 such events in the |
| first 14 minutes after boot, each ~6 s against an 0.086 s median. Issue one |
| request per shape class at startup (short/medium/long prompt, vision, tools) to |
| absorb these before real traffic. Compiled kernels persist in the JIT cache. |
| - **Full-context concurrency is 1×.** At 262,144 `max_model_len` the KV cache holds |
| 263,310 tokens — one maximum-length request. Ordinary workloads are unaffected |
| (16k tokens × 4 concurrent measured fine, KV usage under 1%), but simultaneous |
| full-context requests will queue. |
|
|
| --- |
|
|
| ## Files |
|
|
| ``` |
| config.json model configuration |
| generation_config.json default sampling parameters |
| preprocessor_config.json vision preprocessing (MoonViT) |
| tokenizer.json tokenizer |
| tokenizer_config.json tokenizer configuration |
| model.safetensors.index.json weight shard index |
| provenance.json component checksums |
| benchmarks/ raw lm-eval results + HumanEval script |
| RECIPE.md step-by-step build and evaluation guide |
| ``` |
|
|
| ## Licence |
|
|
| Inherits the DeepSeek model licence from the base checkpoint. The vision tower and |
| projector derive from their respective upstream sources. Verify licence |
| compatibility for your use case before deploying commercially. |
|
|
| ## Acknowledgements |
|
|
| - DeepSeek AI — DeepSeek-V4-Flash base model |
| - vLLM — serving stack, sparse-MLA and DSpark support |
| - EleutherAI — lm-evaluation-harness |
|
|