Instructions to use TensorFold/DeepSeek-V4-Flash-DSpark-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use TensorFold/DeepSeek-V4-Flash-DSpark-MLX with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir DeepSeek-V4-Flash-DSpark-MLX TensorFold/DeepSeek-V4-Flash-DSpark-MLX
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
DeepSeek-V4-Flash DSpark for MLX
This is DeepSeek's DSpark speculative-decoding module for DeepSeek-V4-Flash, converted for Apple Silicon. It drafts tokens for mlx-community/DeepSeek-V4-Flash-4bit on TensorFold's lane engine, which verifies every draft against the target.
- Source: deepseek-ai/DeepSeek-V4-Flash-DSpark at
revision
62af8fffb2f7030cac4de2f0169f5b8d1101b646. DSpark's three blocks aremtp.0.*,mtp.1.*andmtp.2.*inmodel-00046-of-00048.safetensors,model-00047-of-00048.safetensorsandmodel-00048-of-00048.safetensors. Its block fields come from that revision'sconfig.json. - License: MIT, DeepSeek's (see
LICENSE, copyright 2023 DeepSeek). The conversion adds no terms. - Files:
model.safetensors(10.67 GB),config.json,LICENSE,SHA256SUMS.
Use
TensorFold 0.3.6.4 or later on a Mac with 256 GB:
tensorfold pull mlx-community/DeepSeek-V4-Flash-4bit TensorFold/DeepSeek-V4-Flash-DSpark-MLX
tensorfold serve mlx-community/DeepSeek-V4-Flash-4bit
Once this repo has been pulled, serve drafts with it by default; --drafter TensorFold/DeepSeek-V4-Flash-DSpark-MLX
names it explicitly. A round drafts up to five tokens in one pass of the three blocks, which read the target's
streams after its layers 40, 41 and 42. A drafted reply equals the same request sent with "draft": false. Drafts
change speed, never tokens.
What the conversion does
- DSpark's dense projections are FP8 (E4M3) with an E8M0 scale per 128 x 128 block. They are dequantized to bf16, which is exact, then quantized to MLX affine 4-bit in groups of 64. That step is lossy, and it is the format the mlx-community target stores for its own dense projections.
- The routed experts keep DeepSeek's mxfp4 bytes and E8M0 scales unchanged.
- Norms, attention sinks, router weights and biases, hyper-connection tensors,
main_normand the Markov tables are copied as stored. - Block
mtp.<i>becomesdspark.<i>, and its tensors take the names the mlx-community target uses for its own layers (attn.wq_a,ffn.switch_mlp.gate_projand so on).config.jsonnames the head ("model_type": "deepseek_v4_dspark") and carries DSpark's fields from the source: block size 5, noise token 128799, target layers 40-42, Markov rank 256.
TensorFold's converter reproduces model.safetensors byte for byte. Run it with the source's config.json beside
the shards:
python -m tensorfold.families.deepseek_v4.convert dspark model-00046-of-00048.safetensors \
model-00047-of-00048.safetensors model-00048-of-00048.safetensors DeepSeek-V4-Flash-DSpark-MLX
Measured
On an M3 Ultra (60-core GPU, 256 GB) with MLX 0.32.2, TensorFold's lane engine with this drafter against mlx-lm PR #1797's server on the same machine and the same 4-bit weights, with that server's defaults. Decode tok/s at 64 / 256 generated tokens, median over each cell's prompts and seeds, thinking off:
| Cell | TensorFold with DSpark | mlx-lm PR #1797 |
|---|---|---|
| Chat, greedy | 51.4 / 49.3 | 23.4 / 23.1 |
| Chat, sampled | 48.4 / 48.0 | 29.3 / 28.7 |
| Code, greedy | 69.4 / 70.8 | 23.5 / 23.0 |
| Code, sampled | 65.5 / 67.2 | 29.4 / 28.9 |
Without drafts TensorFold decodes 44-46 tok/s on the same machine. DeepSeek's MTP layer, converted the same way, is TensorFold/DeepSeek-V4-Flash-MTP-MLX.
- Downloads last month
- 85
Quantized
Model tree for TensorFold/DeepSeek-V4-Flash-DSpark-MLX
Base model
deepseek-ai/DeepSeek-V4-Flash-DSpark