Piko-9b / reports /final_summary.md
Dexy2's picture
Rewrite model card around verified evidence; correct misattributed benchmarks and config path leak
0810902 verified
|
Raw
History Blame Contribute Delete
5.98 kB
# Final Summary — Piko-9b Release Work
2026-07-29.
## What Piko-9b turned out to be
A **tensor-level composition of two existing open checkpoints**, not a trained model:
* **Language backbone** (9,197,093,888 params, 427 tensors) — bitwise identical to
`wraithfast-phase15-150k-full-ft`, a fine-tune of `deepreinforce-ai/Ornith-1.0-9B` produced by
five merged QLoRA stages.
* **Vision tower and merger** (456,010,480 params, 348 tensors) — bitwise identical to
`Qwen/Qwen3.5-9B`, copied without modification.
Zero of the 775 tensors are unaccounted for. The six QLoRA adapters trained under the "Piko" name
exist on disk with non-zero deltas and are **absent from the published weights**: the published
tensors match the pre-adapter checkpoint exactly.
## The three findings that changed the release
**1. The published benchmark numbers belonged to a different model.** All nine headline scores
matched two local result files to the digit, both recording
`base_model: wraithfast-phase14-100k-full-ft, adapter: wraithfast-phase15-150k-qlora` — a
text-only checkpoint predating the vision composition. The same file's identity probe has that
model answering *"i am called mythos"* and *"i am built on the kairo 6b base model."*
**2. No vision training ever happened.** The phase-6 log contains three events: download → compose
→ run smoke benchmark → crash on a TLS certificate error fetching `images.cocodataset.org`. The
vision dataset was never built. The "OCR" datasets were text-only by the project's own manifest:
*"OCR rows train interpretation of OCR/image text detection outputs, not direct pixel vision."*
**3. The vision path works anyway.** This was the surprise. A tower trained for Qwen's embedding
space, bolted onto a backbone that has drifted 1.5–5 % away from it, with no re-alignment, scores
**10/10 on OCR and 10/10 on document understanding**. It read a rendered receipt's merchant and
total correctly and extracted all four values from a bar chart. The multimodal claim survives — in
a narrowed, measured form.
## What was measured
| Suite | Piko-9b | `Qwen/Qwen3.5-9B` | Verdict |
|---|---:|---:|---|
| Custom regression suite (70 cases) | **65/70** | **63/70** | No significant difference |
| Smoke evaluation | 8/8 | — | Pass |
| Inference validation (12 checks) | 12/12 | — | Pass |
Per category, both models: OCR 10/10 vs 10/10, documents 10/10 vs 10/10, reasoning 10/10 vs 10/10,
long context 5/5 vs 5/5, tables/charts 9/10 vs 9/10, hallucination/safety 7/10 vs 7/10,
instruction following 9/10 vs 8/10, coding 5/5 vs 4/5.
Every category's 95% Wilson intervals overlap. **Piko-9b is not shown to outperform its base
model.** The +20 % on coding is one example out of five.
Performance, 4-bit NF4 on an RTX 5070 Ti: 101 s cold load, 7.37 GB resident, 29–36 tok/s at batch
1, 81–85 at batch 4, ~5,500 tok/s prefill, 13 ms image preprocessing. Decode rate is **flat across
context length** — the hybrid stack working as designed.
## The operational discovery
**`device_map="auto"` silently corrupts this model.** On a GPU too small to hold it, layers offload
to CPU, the linear-attention recurrent state breaks, and every prompt returns `!!!!!!!!!!` with no
error raised. First observed as an apparently catastrophic model failure; isolated by loading the
same weights in 4-bit fully resident, which produced correct output immediately.
This is now the first item in the troubleshooting guide, the first check in the smoke evaluation
(exit code 2), a dedicated regression test, and a guard in every example script. It is the single
most likely way a user's deployment breaks.
## Licensing
The language backbone descends from **MIT**-licensed `Ornith-1.0-9B`; the vision tower is
**Apache-2.0** from `Qwen3.5-9B`. Apache-2.0 for the combined work is permissible **only with the
MIT notice retained** — the previous release shipped an Apache-2.0 `LICENSE` with no attribution
to either upstream. A `NOTICE` file now carries both.
Training-data licensing for the fine-tuning stages **could not be established**: manifests record
categories and row counts, but several sources are named only by local filename.
## Still live and unfixed
`config.json` on the Hub publishes two absolute paths from the author's machine inside its
`piko_composition` block. The corrected block is specified in
[`repository_audit.md`](repository_audit.md) §8. **That fix has not been applied** — it requires
editing the model repository, which was not done without authorisation.
## What was not done, and why
| Item | Reason |
|---|---|
| IFEval, MMLU-Pro, GSM8K, HumanEval, OCRBench, DocVQA, ChartQA, TextVQA, MMMU | 1–3 h per model each, doubled for the baseline. Scripts provided and runnable |
| bfloat16 unquantized profiling | Needs ~22 GB; test GPU has 15.92 GB |
| 8-bit | Not exercised |
| GGUF / AWQ / GPTQ artefacts | Not produced; nothing claimed. Scripts provided |
| Video input | No fixture built; documented as untested |
| Context beyond ~32K | Hardware limit |
| Byte-comparing the Hub checkpoint | Verified by file name and size only; 21 GB re-download not performed |
| Testing the six unmerged adapters | Out of scope, but worth doing — nobody has checked whether they help |
## An honest read of the outcome
The repository is now accurate, reproducible, and complete. But the central measured fact is that
**Piko-9b performs the same as the model whose vision tower it borrowed.** A user choosing between
them should probably take `Qwen/Qwen3.5-9B`: better documented, better supported, and its vision
tower and language backbone were trained together.
What Piko-9b genuinely has is a different language backbone with a different fine-tuning history,
and a demonstration — worth something on its own — that a vision tower can survive being
transplanted onto a drifted backbone without re-alignment.
That is a defensible thing to publish. It is not what the previous model card said.