DeepSeek V4 Flash — Hebrus ExpertMajor v2

This repository contains DeepSeek V4 Flash artifacts for andreaborio/hebrus, a source-built Apple Metal inference engine with bounded SSD expert streaming.

These files are Hebrus-native GGUF extensions. Routed MoE weights are stored once in an embedded, checksummed ds4.expert_major.v2 store. The ds4 token in existing filenames and in that tensor identifier is a stable compatibility field; the runtime and Hugging Face library name are Hebrus.

Hebrus is not a general GGUF runner. Use only the exact artifact and revision listed below. Generic llama.cpp, MLX, Ollama, and hosted Hugging Face inference cannot execute the embedded ExpertMajor store.

Supported artifacts

Artifact Support status Hugging Face revision Bytes SHA-256 Minimum unified memory
DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-DS4-ExpertMajor-v2.gguf Hebrus 0.3.0 default · Qualified pinned download; selected by deepseek-v2 ds4-v0.2.0, resolving to c0740cd4830f142068a6fbbe074f062f2c7c3d19 86,720,114,272 8378080263eb9224f7228d72e2afa4ac3cf74a116023fdec2c596ff228a33e3f 64 GiB
DeepSeek-V4-Flash-0731-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-DS4-ExpertMajor-v2.gguf Current · Opt in; public and covered by dated Hebrus evidence, but not selected by deepseek-v2 63e5f2a8fedeeb77add33e13da5f946c043698bc 86,720,114,240 d89dd628ed786ecf14285cb886459eec01df89fd7e7bf3cbff1416a551bcd966 64 GiB

The first row is the reproducible release path in the current Hebrus downloader and runtime contract. The second row is intentionally not substituted into that command: its retained results are scoped to the exact 0731 artifact, hardware, prompts, and modes in the 2026-08-05 evidence record.

The Hebrus 0.3.0 release records that a new exact-artifact DeepSeek smoke was skipped by maintainer decision. No result from the byte-different 0731 artifact is attributed to the pinned default.

The canonical GGUF without DS4-ExpertMajor-v2 is an offline converter input, not an inference fallback. Hebrus rejects it during model admission.

Quick start: qualified pinned artifact

Requirements:

  • an Apple Silicon Mac with at least 64 GiB unified memory;
  • Xcode Command Line Tools;
  • enough local SSD space for the 86,720,114,272-byte artifact;
  • the official Hugging Face CLI.
git clone https://github.com/andreaborio/hebrus.git
cd hebrus
git checkout v0.3.0

xcode-select --install
python3 -m pip install -U huggingface_hub hf_xet
make -j

./download_model.sh deepseek-v2

./hebrus \
  -m gguf/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-DS4-ExpertMajor-v2.gguf \
  --ctx 8192 \
  -p "Explain bounded SSD expert streaming in one paragraph."

The downloader pins ds4-v0.2.0 and verifies both the complete byte count and SHA-256 before reporting success. The model repositories are public and ungated, so downloading does not require a Hugging Face token.

Normal startup uses AUTO. On qualified hardware it may resolve to resident Metal execution or bounded SSD streaming. Do not add a sidecar, model-admission bypass, manual expert-cache size, or backend flag to the release command.

Optional 0731 artifact

To reproduce the separate opt-in record, download the exact immutable path and verify it independently:

hf download andreaborio/DeepSeek-V4-Flash-Hebrus-GGUF \
  DeepSeek-V4-Flash-0731-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-DS4-ExpertMajor-v2.gguf \
  --repo-type model \
  --revision 63e5f2a8fedeeb77add33e13da5f946c043698bc \
  --local-dir gguf

shasum -a 256 \
  gguf/DeepSeek-V4-Flash-0731-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-DS4-ExpertMajor-v2.gguf

Expected digest: d89dd628ed786ecf14285cb886459eec01df89fd7e7bf3cbff1416a551bcd966. This manual path does not change what deepseek-v2 downloads.

Compatibility and limits

  • Apple Silicon and Metal are the only qualified production backend.
  • The minimum qualified memory is 64 GiB; it is not a promise that every context or concurrent workload fits.
  • CPU is reference/debug only. CUDA and ROCm sources are absent, and distributed inference is retired.
  • ExpertMajor v1, sidecars, canonical routed-weight GGUF inference, and model admission bypasses are rejected.
  • Performance depends on the exact artifact, runtime revision, context, prompt, storage state, memory pressure, routing locality, and thermals.
  • inference: false is deliberate: these custom files are not served by the Hugging Face hosted inference widget.

The current authority for supported modes and hardware floors is the Hebrus runtime support contract.

License and attribution

The upstream deepseek-ai/DeepSeek-V4-Flash repository declares the MIT license. DeepSeek retains the base-model copyright. Quantization credit belongs to antirez and the DS4 project; the embedded ExpertMajor v2 layout and current runtime are maintained by Hebrus.

Downloads last month
2,081
GGUF
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for andreaborio/DeepSeek-V4-Flash-Hebrus-GGUF

Quantized
(126)
this model

Collection including andreaborio/DeepSeek-V4-Flash-Hebrus-GGUF