ARustyCoder93 / README.md
ayourtch's picture
Upload folder using huggingface_hub
fa8d379 verified
|
Raw
History Blame Contribute Delete
3.56 kB
metadata
license: mit
base_model: deepseek-ai/DeepSeek-V4-Flash
tags:
  - gguf
  - llama.cpp
  - moe
  - expert-pruning
  - deepseek4
  - coding
library_name: llama.cpp

ARustyCoder93

A 93-of-256 routed-expert cut of DeepSeek-V4-Flash in GGUF, made for local agentic coding on one 96 GB GPU. This is not a DeepSeek release. It is a derived model produced by expert pruning; DeepSeek did not make or endorse it. It inherits DeepSeek-V4-Flash's MIT license.

Full recipe, scripts, llama.cpp patch and every checksum: https://github.com/apchat-agent/arustycoder

file arustycoder93.gguf — 62,213,283,424 bytes
sha256 90a4c8a8a32fc664d822f1df093d5e228111c94eebf9398ae5c5737391ecb058
source antirez/deepseek-v4-ggufDeepSeek-V4-Flash-MXFP4Experts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-mxfp4-0731.gguf (MXFP4 experts, Q8 attention/shared/output, F16 compressor/indexer); kept tensors are byte-identical to it
experts 93 per layer kept out of 256 (top-6 routing unchanged); 43 MoE layers
runs on stock llama.cpp at or after commit 22b8e31 (2026-08-15). Our serving build for the 2026-08-16 smoke tests was that commit plus the census patch from the recipe repo, which is inert for this file (no full-router tensors). llama-server -m arustycoder93.gguf -ngl 99 -c 262144 -fa on
footprint ~67 GB resident at 256K context on an RTX PRO 6000; ~46 tok/s decode, ~470 tok/s prompt

How the 93 experts were chosen

Two "grow-from-nothing" censuses (start from a random 8 experts per layer with the full 256-wide router kept as side tensors, record which experts the router wants, rebuild at 16, 32, 64) — one on a code/systems corpus, one on a general corpus. ARustyCoder93 is the union of the two keep-64 sets: 78..93 experts per layer (mean 87.3), padded to 93 in the smaller layers by merged rank because a GGUF needs a single expert count. Each keep-64 alone predicts held-out text well but cannot finish generating; the union does. Method, numbers and the merge rule are in the repository README.

What it does and does not do

  • 6/6 coding asks (write Rust, fix Python, explain C, refactor JS, a tool call, a multi-file plan) finished cleanly, 86-1993 completion tokens, with the server settings above. Prompts and raw responses are in the recipe repo under results/.
  • Known limit: a harder single-shot prompt (a generic LRU cache) fell into a verbatim repetition loop in its reasoning to a 12k-token cap. Expect this on some hard reasoning; a DRY sampler stops the loop but corrupts identifiers in code.
  • No standard benchmark has been run on this cut. It is an experiment that turned out useful for interactive coding. Use the unpruned model when correctness matters more than footprint.

Files

  • arustycoder93.gguf — the model
  • config.txt, prune.log — build record of this exact file
  • merged-wants.tsv, coder-census-64.wants.tsv, general-census-64.wants.tsv — the expert rankings it was pruned from
  • PROVENANCE.md — checksums, versions, commits
  • LICENSE-DeepSeek-V4-Flash — the upstream MIT notice (DeepSeek); LICENSE — MIT for the recipe

HF metadata note: base_model points at DeepSeek-V4-Flash; there is no "pruned" relation tag on the Hub, so none is set.

Provenance / credit

Base model: DeepSeek-V4-Flash by DeepSeek (MIT). GGUF source and its quantization: antirez/deepseek-v4-gguf (tagged MIT on the Hub). Pruning method, patch and scripts: this project (MIT), 2026-08-15/17.