TensorFold

DeepSeek-V4-Flash DSpark for MLX

This is DeepSeek's DSpark speculative-decoding module for DeepSeek-V4-Flash, converted for Apple Silicon. It drafts tokens for mlx-community/DeepSeek-V4-Flash-4bit on TensorFold's lane engine, which verifies every draft against the target.

  • Source: deepseek-ai/DeepSeek-V4-Flash-DSpark at revision 62af8fffb2f7030cac4de2f0169f5b8d1101b646. DSpark's three blocks are mtp.0.*, mtp.1.* and mtp.2.* in model-00046-of-00048.safetensors, model-00047-of-00048.safetensors and model-00048-of-00048.safetensors. Its block fields come from that revision's config.json.
  • License: MIT, DeepSeek's (see LICENSE, copyright 2023 DeepSeek). The conversion adds no terms.
  • Files: model.safetensors (10.67 GB), config.json, LICENSE, SHA256SUMS.

Use

TensorFold 0.3.6.4 or later on a Mac with 256 GB:

tensorfold pull mlx-community/DeepSeek-V4-Flash-4bit TensorFold/DeepSeek-V4-Flash-DSpark-MLX
tensorfold serve mlx-community/DeepSeek-V4-Flash-4bit

Once this repo has been pulled, serve drafts with it by default; --drafter TensorFold/DeepSeek-V4-Flash-DSpark-MLX names it explicitly. A round drafts up to five tokens in one pass of the three blocks, which read the target's streams after its layers 40, 41 and 42. A drafted reply equals the same request sent with "draft": false. Drafts change speed, never tokens.

What the conversion does

  • DSpark's dense projections are FP8 (E4M3) with an E8M0 scale per 128 x 128 block. They are dequantized to bf16, which is exact, then quantized to MLX affine 4-bit in groups of 64. That step is lossy, and it is the format the mlx-community target stores for its own dense projections.
  • The routed experts keep DeepSeek's mxfp4 bytes and E8M0 scales unchanged.
  • Norms, attention sinks, router weights and biases, hyper-connection tensors, main_norm and the Markov tables are copied as stored.
  • Block mtp.<i> becomes dspark.<i>, and its tensors take the names the mlx-community target uses for its own layers (attn.wq_a, ffn.switch_mlp.gate_proj and so on). config.json names the head ("model_type": "deepseek_v4_dspark") and carries DSpark's fields from the source: block size 5, noise token 128799, target layers 40-42, Markov rank 256.

TensorFold's converter reproduces model.safetensors byte for byte. Run it with the source's config.json beside the shards:

python -m tensorfold.families.deepseek_v4.convert dspark model-00046-of-00048.safetensors \
    model-00047-of-00048.safetensors model-00048-of-00048.safetensors DeepSeek-V4-Flash-DSpark-MLX

Measured

On an M3 Ultra (60-core GPU, 256 GB) with MLX 0.32.2, TensorFold's lane engine with this drafter against mlx-lm PR #1797's server on the same machine and the same 4-bit weights, with that server's defaults. Decode tok/s at 64 / 256 generated tokens, median over each cell's prompts and seeds, thinking off:

Cell TensorFold with DSpark mlx-lm PR #1797
Chat, greedy 51.4 / 49.3 23.4 / 23.1
Chat, sampled 48.4 / 48.0 29.3 / 28.7
Code, greedy 69.4 / 70.8 23.5 / 23.0
Code, sampled 65.5 / 67.2 29.4 / 28.9

Without drafts TensorFold decodes 44-46 tok/s on the same machine. DeepSeek's MTP layer, converted the same way, is TensorFold/DeepSeek-V4-Flash-MTP-MLX.

Downloads last month
85
Safetensors
Model size
3B params
Tensor type
F32
路
BF16
路
U32
路
U8
路
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for TensorFold/DeepSeek-V4-Flash-DSpark-MLX

Quantized
(13)
this model