Piko-9b / reports /final_summary.md
Dexy2's picture
Rewrite model card around verified evidence; correct misattributed benchmarks and config path leak
0810902 verified
|
Raw
History Blame Contribute Delete
5.98 kB

Final Summary — Piko-9b Release Work

2026-07-29.

What Piko-9b turned out to be

A tensor-level composition of two existing open checkpoints, not a trained model:

  • Language backbone (9,197,093,888 params, 427 tensors) — bitwise identical to wraithfast-phase15-150k-full-ft, a fine-tune of deepreinforce-ai/Ornith-1.0-9B produced by five merged QLoRA stages.
  • Vision tower and merger (456,010,480 params, 348 tensors) — bitwise identical to Qwen/Qwen3.5-9B, copied without modification.

Zero of the 775 tensors are unaccounted for. The six QLoRA adapters trained under the "Piko" name exist on disk with non-zero deltas and are absent from the published weights: the published tensors match the pre-adapter checkpoint exactly.

The three findings that changed the release

1. The published benchmark numbers belonged to a different model. All nine headline scores matched two local result files to the digit, both recording base_model: wraithfast-phase14-100k-full-ft, adapter: wraithfast-phase15-150k-qlora — a text-only checkpoint predating the vision composition. The same file's identity probe has that model answering "i am called mythos" and "i am built on the kairo 6b base model."

2. No vision training ever happened. The phase-6 log contains three events: download → compose → run smoke benchmark → crash on a TLS certificate error fetching images.cocodataset.org. The vision dataset was never built. The "OCR" datasets were text-only by the project's own manifest: "OCR rows train interpretation of OCR/image text detection outputs, not direct pixel vision."

3. The vision path works anyway. This was the surprise. A tower trained for Qwen's embedding space, bolted onto a backbone that has drifted 1.5–5 % away from it, with no re-alignment, scores 10/10 on OCR and 10/10 on document understanding. It read a rendered receipt's merchant and total correctly and extracted all four values from a bar chart. The multimodal claim survives — in a narrowed, measured form.

What was measured

Suite Piko-9b Qwen/Qwen3.5-9B Verdict
Custom regression suite (70 cases) 65/70 63/70 No significant difference
Smoke evaluation 8/8 Pass
Inference validation (12 checks) 12/12 Pass

Per category, both models: OCR 10/10 vs 10/10, documents 10/10 vs 10/10, reasoning 10/10 vs 10/10, long context 5/5 vs 5/5, tables/charts 9/10 vs 9/10, hallucination/safety 7/10 vs 7/10, instruction following 9/10 vs 8/10, coding 5/5 vs 4/5.

Every category's 95% Wilson intervals overlap. Piko-9b is not shown to outperform its base model. The +20 % on coding is one example out of five.

Performance, 4-bit NF4 on an RTX 5070 Ti: 101 s cold load, 7.37 GB resident, 29–36 tok/s at batch 1, 81–85 at batch 4, ~5,500 tok/s prefill, 13 ms image preprocessing. Decode rate is flat across context length — the hybrid stack working as designed.

The operational discovery

device_map="auto" silently corrupts this model. On a GPU too small to hold it, layers offload to CPU, the linear-attention recurrent state breaks, and every prompt returns !!!!!!!!!! with no error raised. First observed as an apparently catastrophic model failure; isolated by loading the same weights in 4-bit fully resident, which produced correct output immediately.

This is now the first item in the troubleshooting guide, the first check in the smoke evaluation (exit code 2), a dedicated regression test, and a guard in every example script. It is the single most likely way a user's deployment breaks.

Licensing

The language backbone descends from MIT-licensed Ornith-1.0-9B; the vision tower is Apache-2.0 from Qwen3.5-9B. Apache-2.0 for the combined work is permissible only with the MIT notice retained — the previous release shipped an Apache-2.0 LICENSE with no attribution to either upstream. A NOTICE file now carries both.

Training-data licensing for the fine-tuning stages could not be established: manifests record categories and row counts, but several sources are named only by local filename.

Still live and unfixed

config.json on the Hub publishes two absolute paths from the author's machine inside its piko_composition block. The corrected block is specified in repository_audit.md §8. That fix has not been applied — it requires editing the model repository, which was not done without authorisation.

What was not done, and why

Item Reason
IFEval, MMLU-Pro, GSM8K, HumanEval, OCRBench, DocVQA, ChartQA, TextVQA, MMMU 1–3 h per model each, doubled for the baseline. Scripts provided and runnable
bfloat16 unquantized profiling Needs ~22 GB; test GPU has 15.92 GB
8-bit Not exercised
GGUF / AWQ / GPTQ artefacts Not produced; nothing claimed. Scripts provided
Video input No fixture built; documented as untested
Context beyond ~32K Hardware limit
Byte-comparing the Hub checkpoint Verified by file name and size only; 21 GB re-download not performed
Testing the six unmerged adapters Out of scope, but worth doing — nobody has checked whether they help

An honest read of the outcome

The repository is now accurate, reproducible, and complete. But the central measured fact is that Piko-9b performs the same as the model whose vision tower it borrowed. A user choosing between them should probably take Qwen/Qwen3.5-9B: better documented, better supported, and its vision tower and language backbone were trained together.

What Piko-9b genuinely has is a different language backbone with a different fine-tuning history, and a demonstration — worth something on its own — that a vision tower can survive being transplanted onto a drifted backbone without re-alignment.

That is a defensible thing to publish. It is not what the previous model card said.