Field notes from a 2x DGX Spark (GB10) deployment of this quant β€” wedge recovery, firewall deadlock, MTP acceptance data

#4
by randomllama - opened

Ran this quant for two weeks as the production lane on 2x DGX Spark (SGLang TP=2, 262K context, MTP). Collected the operational findings that weren't in any published recipe: GB10 unified-memory exhaustion wedges the whole box instead of erroring (recovery is unplug/replug β€” the power button is dead), a multi-node deadlock that only appears behind a default-deny firewall, first MTP acceptance measurements we know of for this model, and three optimization dead ends documented as dead ends.

All of it: https://huggingface.co/randomllama/Qwen3.8-Flash-Next-DGX-Spark-field-notes (mirror of the GitHub repo)

Also from the same rig: a small controlled study on eliminating this model's hallucinated tool-call failure mode via prompt format pinning β€” https://huggingface.co/datasets/randomllama/qwen-format-pin-study

Happy to answer questions from anyone running this on Sparks.

Sign up or log in to comment