--- license: other license_name: deepseek license_link: https://github.com/deepseek-ai/DeepSeek-V2/blob/main/LICENSE-MODEL language: - en pipeline_tag: image-text-to-text library_name: vllm tags: - deepseek - multimodal - vision-language - fp8 - speculative-decoding - vllm - blackwell --- # DeepSeek V4 Flash Hybrid Vision Reseek A vision-enabled assembly of **DeepSeek-V4-Flash-0731**: the DeepSeek-V4 Flash text model (48 shards, FP8) joined to a MoonViT-style vision tower and multimodal projector, served through vLLM with **DSpark speculative decoding**. Built and benchmarked on 2× NVIDIA RTX PRO 6000 Blackwell Workstation Edition (96 GB each, SM120), tensor-parallel 2. | | | |---|---| | Architecture | `DeepSeekV4VisionForConditionalGeneration` | | Context length | 262,144 tokens | | KV cache | `fp8_ds_mla` (656 B/token) | | Vision encoder | MoonViT-style, patch 14, 2×2 merge | | Speculative decoding | DSpark, 5 draft tokens | | Tool calling | Yes (`deepseek_v4` parser) | --- ## Benchmarks All numbers produced with **EleutherAI lm-evaluation-harness v0.4.12** against a live vLLM endpoint. Raw result JSON is in [`benchmarks/`](./benchmarks). Reproduction commands are in [`RECIPE.md`](./RECIPE.md). ### Reasoning and instruction following | Benchmark | Metric | Score | n | |---|---|---|---| | **GSM8K** (5-shot) | exact_match, flexible | **96.5%** | 200 | | GSM8K (5-shot) | exact_match, strict | 69.0% | 200 | | **IFEval** | inst-level loose | **84.0%** | 200 | | IFEval | inst-level strict | 81.5% | 200 | | IFEval | prompt-level loose | 76.5% | 200 | | IFEval | prompt-level strict | 73.5% | 200 | ### Knowledge and commonsense (0-shot, loglikelihood) | Benchmark | Metric | Score | n | |---|---|---|---| | **PIQA** | acc_norm | **87.7%** | 300 | | PIQA | acc | 83.0% | 300 | | **Winogrande** | acc | **79.7%** | 300 | | **HellaSwag** | acc_norm | **71.3%** | 300 | | HellaSwag | acc | 57.0% | 300 | | **ARC-Challenge** | acc_norm | **63.3%** | 300 | | ARC-Challenge | acc | 62.3% | 300 | | TruthfulQA MC2 | acc | 54.9% | 300 | | OpenBookQA | acc_norm | 51.3% | 300 | ### Code | Benchmark | Metric | Score | n | |---|---|---|---| | **HumanEval** | pass@1 | **78.0%** | 164 | > **Note on HumanEval.** The stock `lm-eval` HumanEval task scores this model at > **0.0** — an artifact, not a result. That task expects a raw completion, but a > chat endpoint returns prose plus a ```` ```python ```` block, so the built-in > filter extracts nothing. The 78.0% above comes from > [`benchmarks/humaneval_chat.py`](./benchmarks/humaneval_chat.py), which pulls the > fenced code and runs the official unit tests in a sandboxed subprocess. If you > benchmark any chat model on HumanEval, check for this first. ### Throughput Measured on 2× RTX PRO 6000 Blackwell, TP=2, `max_num_seqs=16`: | Load | Aggregate throughput | Tail latency | |---|---|---| | 1 concurrent | 305 tok/s | 0.39 s | | 8 concurrent | 581 tok/s | 1.64 s | | 16 concurrent | 1150 tok/s | 1.66 s | Time to first token: **0.085 s** median. Prefill: ~8,900 tok/s at 32k tokens. DSpark speculative decoding accepts 3.2–5.9 of 5 draft tokens in practice. --- ## Vision: architecture, capabilities, limits ### The adapter **457M parameters total** — roughly 0.1% of the assembled model: | Component | Params | Shape | |---|---|---| | Vision tower (MoonViT-style) | 416.9M | 27 layers, width 1152, patch 14 | | Multimodal projector | 40.1M | `pre_norm → 4608×4608 → GELU → 4608×4096` | Dimensions chain exactly: the tower emits 1152-wide patch embeddings, a 2×2 merge concatenates them to 4608, and the projector maps that to 4096 — the language model's hidden size. A 448×448 image becomes 256 patches → **64 tokens** after merging, which is frugal next to tile-based encoders that spend thousands. Preprocessing was verified numerically: the normalization LUT produces exactly ±1.0 at grey 0 and 255, with correct 16×16 spatial ordering. ### Strengths - **Cheap.** ~0.8 GB VRAM, no measurable cost to text quality — all text benchmarks score identically with the vision tower loaded. - **Accurate within its design point.** 9/10 on balanced yes/no probes at 448 px; 4/5 on 4-way forced choice with the correct answer listed last. It correctly rejects absent objects (cat, car, people, text, bicycle), so it is not simply answering "yes". - **Token-efficient**, per the merge arithmetic above. ### Weaknesses - **Fabricates above ~450 px.** The dominant failure. Balanced accuracy falls from 6/7 to 4/7 while "yes" answers rise from 4/7 to 6/7 — it drifts toward agreeable confabulation rather than degrading gracefully. - **Weak fine-grained recognition.** Category right, identity wrong: a numbat read as "giraffe", a rainbow lightbulb as "camera lens". Scene structure survives; species and object identity do not. - **Blind to synthetic images.** Solid colour fills score 1/4 — chance. Expected for a patch-based ViT, since a flat 14×14 patch carries no edges or texture, but it means **synthetic images cannot be used to smoke-test the pipeline**. A flat test image will look like total failure on a perfectly healthy deployment. - **Thin projector.** A 2-layer MLP is the minimum viable bridge from a 1152-wide tower into a 4096-wide LM, and is the most likely cause of the identity confusions above. ### Where this adapter could improve 1. **Deeper or attention-based projector.** Replacing the 2-layer MLP with a cross-attention resampler (Q-Former / Perceiver style) targets the identity errors directly, and at 40M params it is by far the cheapest component to retrain — the tower need not be touched. 2. **Higher-resolution alignment training.** The >450 px collapse points at the projector rather than the tower: token counts scale correctly with image area, so the encoder handles more patches; it is the mapping that degrades. 3. **Negative-example alignment.** The failure mode is confident invention, not refusal. The model answers "is there text in this image?" correctly at 448 px, so the capability exists but is not robust to resolution. 4. **OCR evaluation.** Every fabrication observed involved invented *text* ("Airbus", "The Great Forest"). This is a hypothesis, not a measurement — it was not tested against an OCR benchmark. > **Evaluation gap.** The vision figures below come from hand-built probes over > roughly ten images, not a standard benchmark. The text scores in this card are > `lm-eval` reproducible; **the vision claims are not**. Run MMMU, MMBench, or > TextVQA before relying on them. ### Measured behaviour **What works** — images at or below **448 px** on the longest side: | Test | Result | |---|---| | Balanced yes/no on real photos | 9/10 | | 4-way forced choice (correct answer listed last) | 4/5 (chance 25%) | | Detail questions (water present? snow-capped? time of day?) | correct | **What fails:** - **Open-ended captioning above ~450 px.** The model confabulates. Asked to describe a desert landscape at 896 px it reported *"a bird flying over a sunset with the words 'the great outdoors'"* — text that does not exist in the image. Feeding a full-resolution photo produced an invented *"Airbus"* logo. - **Synthetic flat-colour images.** Solid fills score at chance (1/4). Flat colour is out of distribution for a patch-based ViT — no edges or texture in any patch. This is expected behaviour, not a defect. **Recommendations:** 1. **Downscale images to 448 px** before sending. Measured effect: one test image went from 5/7 to **7/7** correct. It also cuts a 4736×2656 photo from ~3,100 prompt tokens to ~260 (**12× fewer**). 2. **Ask specific questions** ("Is there water in this image?") rather than "describe this image." Factual queries stay accurate where captions drift. --- ## Quick start ### Serve with vLLM ```bash vllm serve /path/to/DeepSeek-V4-Flash-Hybrid-Vision-Reseek \ --served-model-name dsv4-hybrid-vision \ --tensor-parallel-size 2 \ --tokenizer-mode deepseek_v4 \ --kv-cache-dtype fp8_ds_mla \ --block-size 256 \ --max-model-len 262144 \ --max-num-seqs 16 \ --gpu-memory-utilization 0.968 \ --speculative-config '{"method":"dspark","model":"/path/to/model","num_speculative_tokens":5}' \ --enable-auto-tool-choice \ --tool-call-parser deepseek_v4 ``` ### Query it ```python import base64, io, json, urllib.request from PIL import Image URL = "http://127.0.0.1:8000/v1/chat/completions" def ask(messages, max_tokens=512): body = {"model": "dsv4-hybrid-vision", "messages": messages, "max_tokens": max_tokens, "temperature": 0.6, "top_p": 0.95} req = urllib.request.Request(URL, data=json.dumps(body).encode(), headers={"Content-Type": "application/json"}) return json.loads(urllib.request.urlopen(req).read())["choices"][0]["message"]["content"] # Text print(ask([{"role": "user", "content": "What port does SSH use?"}])) # Vision -- downscale to 448 px first (see the vision notes above) im = Image.open("photo.jpg").convert("RGB") im.thumbnail((448, 448)) buf = io.BytesIO(); im.save(buf, "PNG") b64 = base64.b64encode(buf.getvalue()).decode() print(ask([{"role": "user", "content": [ {"type": "text", "text": "Is there a mountain in this image? Answer yes or no."}, {"type": "image_url", "image_url": {"url": "data:image/png;base64," + b64}}, ]}])) ``` --- ## Recommended sampling The checkpoint's `generation_config.json` defaults to `temperature=1.0`, `top_p=1.0`. At those settings the model **pads its answers heavily** — asked for a single bash one-liner it returned five variants plus a bullet-point explainer (347 tokens for a one-line question). These settings were measured to fix that: ```json {"temperature": 0.6, "top_p": 0.95, "max_tokens": 1024} ``` Pair them with a concise system prompt for short answers. With both applied, the same one-liner question returns **29 tokens** — one command, no menu. Note that `--generation-config vllm` makes vLLM **ignore** the checkpoint's `generation_config.json`; pass `--override-generation-config` to set defaults server-side. --- ## Reasoning mode Chain-of-thought is **off by default** — the `deepseek_v4` tokenizer sets `thinking=False` unless a caller opts in with `enable_thinking` or `reasoning_effort`. To pin it off server-side regardless of what clients send: ```bash --default-chat-template-kwargs '{"thinking": false, "enable_thinking": false}' ``` --- ## Known constraints - **KV cache dtype is fixed at `fp8_ds_mla` on SM120.** `nvfp4_ds_mla` would cut the record from 656 to 432 B/token (~1.52× more cache), and both the CLI and the b12x backend advertise it — but every DeepSeek-V4 attention class reachable on SM120 sets `use_fp8_ds_mla_layout = True`, which asserts the dtype starts with `fp8`. Attempting it fails at load: `AssertionError: DeepseekV4 fp8_ds_mla layout only supports fp8 kv-cache`. The one class with the flag off is gated to compute capability 9/10 (Hopper). The underlying FlashInfer kernel is a prebuilt cubin documented as 584 B/token, BF16 or FP8 E4M3 only, so this needs upstream support — not a config change. - **JIT warmup.** vLLM compiles Triton/TileLang kernels on first sight of each tensor shape and logs *"causes a latency spike"*. Measured: 12 such events in the first 14 minutes after boot, each ~6 s against an 0.086 s median. Issue one request per shape class at startup (short/medium/long prompt, vision, tools) to absorb these before real traffic. Compiled kernels persist in the JIT cache. - **Full-context concurrency is 1×.** At 262,144 `max_model_len` the KV cache holds 263,310 tokens — one maximum-length request. Ordinary workloads are unaffected (16k tokens × 4 concurrent measured fine, KV usage under 1%), but simultaneous full-context requests will queue. --- ## Files ``` config.json model configuration generation_config.json default sampling parameters preprocessor_config.json vision preprocessing (MoonViT) tokenizer.json tokenizer tokenizer_config.json tokenizer configuration model.safetensors.index.json weight shard index provenance.json component checksums benchmarks/ raw lm-eval results + HumanEval script RECIPE.md step-by-step build and evaluation guide ``` ## Licence Inherits the DeepSeek model licence from the base checkpoint. The vision tower and projector derive from their respective upstream sources. Verify licence compatibility for your use case before deploying commercially. ## Acknowledgements - DeepSeek AI — DeepSeek-V4-Flash base model - vLLM — serving stack, sparse-MLA and DSpark support - EleutherAI — lm-evaluation-harness