qwen-dflash2-spark / docs /README-agent.md
cfontes's picture
Qwen3.8-27B DFlash2 on 2x DGX Spark: one-stop stack (135 tok/s C1, lm_head BF16 fix)
5be2256 verified
|
Raw
History Blame Contribute Delete
1.6 kB

Agent runbook

For AI agents operating this stack. Humans: see README.md first.

Fast paths

Is it up?

curl -s -o /dev/null -w '%{http_code}\n' http://NODE1:8004/health   # want 200

Bring the pair up

bash scripts/04-launch-pair.sh    # r1 first, then r0; health ≤ 7.5 min

Bring it down

ssh NODE2 'docker rm -f qwen38-dflash2-tp2-r1' ; ssh NODE1 'docker rm -f qwen38-dflash2-tp2-r0'

Quick health of both ranks

ssh NODE1 'docker logs --tail 3 qwen38-dflash2-tp2-r0 2>&1' | tail -3
ssh NODE2 'docker logs --tail 3 qwen38-dflash2-tp2-r1 2>&1' | tail -3

Bench once

python3 bench/edit_bench.py http://NODE1:8004/v1 qwen3.8-27b-dflash2

Failure signatures → root causes

Symptom Root cause Fix
boot dies: unquantized LM head gate lm_head quantized (wrong checkpoint) run 02 + 03, relaunch
health 000 after >8 min see logs; often wrong checkpoint or NCCL docker logs qwen38-dflash2-tp2-r0
accept ~99.6%, tps low K=7 (1 block) set K=16 in launch
accept ~66%, tps below peak K=24 set K=16
r0 up, r1 dead r1 started after r0 relaunch both, r1 first

Hard rules

  • K must be a multiple of 8.
  • rank1 before rank0, always.
  • Kill stale background launcher jobs before relaunching (zombie once rm -f'd a healthy pair and relaunched the wrong checkpoint).
  • Never edit config.json in place on hardlink copies — temp+os.replace.
  • The original unsloth-nvfp4/ checkpoint must stay pristine (it is the rollback). If corrupted, re-download via 00.