ARustyCoder93 / PROVENANCE.md
ayourtch's picture
Upload folder using huggingface_hub
fa8d379 verified
|
Raw
History Blame Contribute Delete
8.61 kB

Provenance — ARustyCoder93

Everything below was recorded on the build host at build time or checked against the files afterwards. Dates are UTC unless marked.

Output

file bytes sha256
arustycoder93.gguf (built as deepseek4-keep93.gguf, 2026-08-16 19:25 UTC) 62,213,283,424 90a4c8a8a32fc664d822f1df093d5e228111c94eebf9398ae5c5737391ecb058

GGUF header: arch deepseek4, expert_count = 93, expert_used_count = 6, 43 MoE layers (3 hash-routed + 40 router layers), all other tensors and metadata copied byte-for-byte from the source. Coarse ftype label as shown by llama.cpp is not a quant family — the file is mixed (MXFP4 experts, Q8/F16 elsewhere), exactly as in the source.

Source model

item value
repo https://huggingface.co/antirez/deepseek-v4-gguf (MIT), base deepseek-ai/DeepSeek-V4-Flash (MIT)
file DeepSeek-V4-Flash-MXFP4Experts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-mxfp4-0731.gguf
bytes 155,976,458,848 (matches the HF tree listing)
sha256 0e3a161b670f686128ec5f92a601dfde616a37bf5e7e48999fa2d32471b57ec6
arch deepseek4, 43 layers × 256 routed experts, top-6 + 1 shared; layers 0-2 hash-routed via ffn_gate_tid2eid

Build host

item value
GPU NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 96 GB (97887 MiB), power limit 350 W; one GPU used (CUDA_VISIBLE_DEVICES=0)
driver / CUDA 610.57.04 / CUDA 13.3 (nvcc V13.3.73)
OS Ubuntu 24.04.4 LTS, kernel 6.8.0-137-generic
llama.cpp ggml-org/llama.cpp commit 22b8e310b921d568e013e4533002be5a8fe53f17 (2026-08-15, "server: re-design yield_to_queue thread model (#27133)") + llama.cpp-patch/wants-census.patch; built Release, GGML_CUDA=ON, GGML_NATIVE=ON, gcc/g++ from Ubuntu 24.04
python 3.12.3 (venv), gguf==0.19.0, numpy==2.5.2, tqdm==4.70.0, PyYAML==6.0.3

Scripts (as run; identical to the copies in scripts/)

file sha256
prune_experts.py f1ea4f50c49c354da32a64e232b47e2a8e41ec178483824a6ef1765ca0bf5962
compare_sets.py 0d6992654b3cd2a08c0bbba643e83c638583ff0c53c888dba2c0be6db40bb76c
apply_wants_patch.py 592f3f34a76a0786fbbb512af9120252cbab59d93359aca102b8c3e7916b4406
grow.sh 1d49a05c619be943552660312a1e5391cdb1664c9c3ba1e5964f10f88458ff3c
grow2.sh 259ed5e0c78a15b1b85f67dc9cf5dd3d79b48eba84fce2df7bc29df1912602f1
grow-coder.sh 9dd532ca77f56b5f1eb1c84dab4ec0f28839f6042a3fed8f737d36ed6c4b32fd
build_corpus.py a516e863acf76136334e417dfe3590ef2d4fce3a8bb9ff73c5a17d9745b7b3c4
build_coder_corpus.py 36a762ae06ba5ae5a9f26327b5e74d0276b765d9ed0f95203d3895bf8cc8ec6b
merge_wants.py 62427369e07ccdb03afc7d0bce3d58c133f770b4275b91a884e52eead0158a4e — written after the fact (2026-08-17) as the re-implementation of the inline merge; verified top-93-identical per layer against merged-wants.tsv
grow-ppl-paired.sh c2272bcc1edb5d1d427542cdc23fe1deb5268baf9df1edc2834952a54c97cafc
grow-ppl-ci.sh e0350ea7d3490d3b98c626366e37f37f0acb29f6f5f2a3553e9f7e5eaafaf00f
ppl-paired.py 349be3a0af9f7c53e79d1bf620d919b396a65bdcf84f063eaa27046b5addee57 (the paired-test tool; results/RESULTS.md)
code-smoke.py cd2fb11c48f43fb02887ef717ece92ac632723ad4c128249c8164154ee093620

grow-coder.sh as published differs from the as-run copy in two comment/ provenance lines (an internal repo path and a wiki reference were removed); the sha256 above is the as-run file. Likewise configs/grow-coder/config.txt had one appended correction paragraph shortened and the hostname redacted in all three config.txt; nothing numeric changed.

Census / ranking files (in wants/)

file sha256 role
coder-census-64.wants.tsv (was grow-coder/census-64.imatrix.wants.tsv) bead82b2976d9c48ae4339c1f2ae2c1d32e72cc1386b530078258e572658cd0a final code-grow census, input A
general-census-64.wants.tsv (was grow/census-64.imatrix.wants.tsv) f20041fac879f2c4553f416fc03383d4bd30ea0435b38144668d97ffb665621b final general-grow census, input B
merged-wants.tsv 8152442949c297b0dc159533f71f883165a709df5b3deaa4debf07b3b02b263b the ranking the model was pruned with (PRUNE_WANTS)

Final prune command (from configs/grow-union/config.txt):

PRUNE_WANTS=merged-wants.tsv python prune_experts.py SRC.gguf deepseek4-keep93.gguf 93 6

Corpora (not redistributed; rebuild with the builders + these sources)

file bytes sha256 built from
coder-train.txt 4,197,328 f1698b26cc486134027b9207eaf0109a1b716e2c0643c74049a7681685466122 build_coder_corpus.py: ripgrep crates/**/*.rs, llama.cpp src/ggml/tools C/C++ + gguf-py/scripts Python + headers, Rust book src/*.md, OWASP CheatSheetSeries, Linux man pages; 4000-char chunks, random.seed(7), every 10th chunk to heldout
coder-heldout.txt 468,010 18314cd45eccf14cda9fee14542a6ee0b7bf9c9022bf27776bef2a6b5e13f7b0 same
domain-train.txt a3e3fd1a326e6568f7ea944d186cc8826a198bdec0fcd14a92d1c2e1f1b6e045 build_corpus.py: Rust book, ripgrep, OWASP, man pages, GSM8K (first 1800 rows), plus ~0.7 MB of our own model reasoning traces (agentic.txt, sha256 ae4f8f386dd9fdff1cd72385561dfaae2eb84ee82d2b958c51c6d7a5b4723e05, not published)
domain-heldout.txt 78e379ccaa5b5dfcb0bbb2f106c78aa10940069e484d238afb499ec683f111aa same
gsm8k.jsonl 4,166,206 17f347dc51477c50d4efb83959dbb7c56297aba886e5544ee2aaed3024813465 GSM8K train split, one JSON object per line (question, answer)
man.txt 1,602,349 a95d16db24d8ae616ef582c657c900f522fe77b26be6313fa0a644b3a740d213 concatenated man output from the build host (starts with SYSTEMD(1))
ds4flash-domain-200.imatrix 26ceeec5a0f1c6f29809c4154c4eb1872f5518eddfd9125732dc3a354f4b701e full-model "truth" census over domain-train.txt (205 chunks); used only for comparison
wiki.test.raw 173c87a53759e0201f33e0ccf978e510c2042d7f2cb78229d9a50d79b9e7dd08 wikitext-2 test, PPL reference only

Corpus source checkouts on the build host:

repo commit
github.com/BurntSushi/ripgrep 3fce3b5bb0236da2df6d99672afb8a719642eca7
github.com/OWASP/CheatSheetSeries 07111ee754e832e335377ac64fd0f8f848d9029c
github.com/rust-lang/book 917544888a55e4da7109bdba8c88c893c0da70f4
github.com/ggml-org/llama.cpp (code corpus source) 22b8e310b921d568e013e4533002be5a8fe53f17 (working tree with the wants patch applied)

Grow runs

General grow (grow.sh, then grow2.sh after a pruner parser crash at the keep-16 stage; 2026-08-15 21:57 → 23:28 UTC per configs/grow-general/grow.log and file mtimes): random keep-8 (PRUNE_SEED=1, PRUNE_FULL_ROUTER=1) → census 100 chunks of domain-train.txt at -c 2048 -b 2048 -ub 2048 → keep-16 → keep-32 → keep-64; final grown-keep64.gguf (45,531,490,912 bytes, from ls; no sha256 recorded) built from census-64.imatrix.wants.tsv with a normal pruned router. Per-stage logs in configs/grow-general/. This run predates the config-record habit, so it has no config.txt; the driver and logs are the record. The keep-8 stage ran with the pre-fix pruner (the bug was in PRUNE_WANTS parsing, which the random keep-8 stage does not use); every later stage used the pruner at the sha256 above.

Code grow (grow-coder.sh, 2026-08-15 23:44 → 2026-08-16 01:16 UTC): same schedule on coder-train.txt; full record in configs/grow-coder/config.txt (final deepseek4-coder-keep64.gguf 45,531,490,912 bytes, sha256 d9630be3be692823db1460d544035dda737f30d72bb84c8da2c9d0071e5fb186).

Union build (2026-08-16 19:25 UTC): configs/grow-union/config.txt + prune.log. Bytes in 145.3 GiB → out 57.9 GiB (tensor payload); file 62,213,283,424 bytes.

merge-128 (configs/grow-merge/, 2026-08-16 07:43 UTC): the earlier union + ~40 filler experts at keep-128. Generated cleanly but did badly on an internal 92-question reasoning eval (most cases hit the token cap; of the 17 it finished within budget it got 9 right where the unpruned base got 16) and was deleted; kept here because it is where the +1e7 union-bonus rule was found and verified.

Serving (as run for the smoke tests)

llama-server -m arustycoder93.gguf -ngl 99 -c 262144 -fa on

~67 GB resident on the PRO 6000 at 256K context; ~46 tok/s decode, ~470 tok/s prompt processing (from the server log, single request).