K3-Stuff / scripts /TESTONE_RESULTS.md
TessaCoil's picture
Upload folder using huggingface_hub
ddf8c5b verified
|
Raw
History Blame Contribute Delete
3.98 kB
# Test One results β€” vast.ai instance 48699333 (2026-08-25 evening)
Box: dual E5-2660 v3 (40 threads), 503 GB RAM, **2Γ— RTX 3090 24 GB**, 32 GB disk,
~930/845 Mbps. Image: vast `llama.cpp` derivative. **Purpose: SSH + scripts + build
pipeline validation on the cheap box. All passed.**
## SSH connectivity (the thing Test One answered)
- Works end-to-end: `ssh -p 50990 root@173.75.47.72` with the account ed25519 key.
- Key is passphrase-protected β†’ must be in ssh-agent (`ssh-add ~/.ssh/id_ed25519`) for
scripted access; plain `ssh -i` in non-interactive shells fails silently as
"Permission denied (publickey)". This was the only connectivity false alarm.
- rsync/scp both directions verified (scripts pushed, logs pulled).
## Build & run
- Prebuilt `/opt/llama.cpp/cuda-12.8/llama-cli` is **v1 β€” useless**; built pinned
master `d222767c7` (10630) successfully. One extra step needed: vast's llama.cpp
image lacks cuBLAS dev headers β†’ `apt install libcublas-dev-12-9` (now baked into
`setup.sh`, version auto-detected from nvcc).
- Smoke gate **PASSED** on Qwen3-4B Q4_K_M stand-in (disk only had 20 GB free;
the real 30B smoke gate still runs on the primary box).
- GPU decode ~185 tok/s; prompt eval ~3.3–3.4K tok/s; llama-bench pp512 8.4K,
pp2048 12.5K, tg32 214 t/s. CPU-only decode 8.2 t/s.
- New metrics plumbing works: `logs/metrics.tsv` now records per-run exit, wall,
pp_tps, tg_tps, **disk_read_mb** (cgroup io.stat β€” real storage reads, cache hits
excluded), and model page-cache residency via mincore. llama-cli's compact
`[ Prompt: X t/s | Generation: Y t/s ]` format is parsed.
## Vast container constraints discovered (affect the K3 runbook!)
- **Unprivileged Docker**, no kernel control:
- ❌ drop_caches (`/proc/sys` read-only) β†’ "cold" expert-streaming states can't be
forced repeatedly. E1a-cold is only truly cold ONCE, right after model load.
`--load-mode dio` (E3) is the controllable proxy; scripts now log
CONTAINER_LIMITED instead of failing.
- ❌ numactl mempolicy (S4b skipped: `set_mempolicy: Operation not permitted`) and
/sys readahead writes (E2) β€” both now guarded with CONTAINER_LIMITED fallback.
- βœ… /dev/shm is tmpfs sized to half RAM (251 GB here) β†’ E4 RAM-filler trick works.
- βœ… cgroup v2 io.stat gives per-run disk-read bytes (the "how often do experts
come from disk" metric).
- If true VM instances exist at similar prices for the primary rental, they'd remove
the drop_caches/numactl limits β€” worth checking "VM" instances on vast.ai; the
chemistry of E1/E2 is better there. Not required: E1-warm/E3-dio still answer the
cache questions.
## Numbers recorded (home-build recon, this box)
- sysbench memory: ~52.5 GB/s (load-store, underestimates stream BW).
- fio (file-based, container overlay): randread-1M β‰ˆ 1.06 GB/s, seqread-4M β‰ˆ 2.9 GB/s.
Host SSD was advertised ~3.2 GB/s β€” overlay + shared host trims it.
- dmidecode unavailable in container (no /dev/mem) β€” DIMM recon impossible on vast
docker instances. (lscpu works: sockets visible.)
## Machine state at teardown-decision time
- Pinned llama.cpp built at /root/llama.cpp; k3-test at /root/k3-test; Qwen3-4B
Q4_K_M (2.4 GB) at /root/models/qwen4b. Disk ~95% full (32 GB alloc).
- Nothing on the box persists through recycle/destroy (no volume).
- Logs archived locally: `k3-test/logs-testone-48699333/`.
## Fixes made during test one (all committed)
1. `setup.sh`: env-assignment-before-$SUDO bug; libcublas-dev install.
2. `lib/common.sh`: drop_caches CONTAINER_LIMITED; io.stat+mincore run metrics;
compact t/s parsing; metrics.tsv.
3. `10_smoke.sh`: S4b container-limit skip; S6 pp-scaling bench.
4. `run_exp` name-passing bug; `-st` replaces unregistered conversation flags.
5. `00_recon.sh`: file-based fio with free-space sizing.
6. New: `15_pp_scaling.sh`, `prompts/gen_prompts.sh` (512/2K/8K graded prompts),
E9 (K3 pp scaling), E1 length-graded decode runs.