ApacheOne's picture
Publish packed-Q1 native LoRA reproducibility release
bf45b91 verified
|
Raw
History Blame Contribute Delete
4.4 kB
metadata
language:
  - en
license: other
library_name: llama.cpp
pipeline_tag: text-generation
base_model: prism-ml/Bonsai-27B-gguf
tags:
  - llama.cpp
  - lora
  - q1
  - q1_0
  - packed-quantization
  - native-training
  - checkpoint-resume
  - reproducibility
  - cuda

Bonsai-27B packed-Q1 native LoRA training release

This is a reproducible engineering release for native LoRA training directly against the packed Q1_0 Bonsai-27B base model in a pinned Prism/Mintplex llama.cpp tree.

The repository contains adapters, deterministic checkpoints, dataset shards, source changes, clean-build evidence, verification scripts, reports and an optional Linux/CUDA build. The 3.8 GB base GGUF is not redistributed.

Release status

Stage Result
Step 8: native multi-target packed-Q1 LoRA PASS
Step 9: deterministic dataset pipeline PASS
Step 10: native training and exact checkpointing PASS
Step 11: correctness and quality acceptance PASS
Step 12: performance and stability acceptance PASS
Step 13: clean reproducibility build PASS

What is proven

  • Packed Q1_0 base weights remain immutable.
  • Six LoRA tensors are optimized without a persistent expanded base model.
  • Checkpoint save/reload is exact.
  • Interrupted accumulation resumes deterministically.
  • Independent resume branches produced byte-identical checkpoints, adapters and state fingerprints.
  • Load/unload, quick soak, RAM and VRAM checks passed.
  • The patch applies to commit 7529fdaaf99ffdc5ca71ace9c7409a56b27ad92f and selected targets build in a clean worktree.

LoRA targets

Base tensor Family Rank Alpha Base shape
blk.0.ssm_alpha.weight SSM 4 8 48 × 5120
blk.11.attn_k.weight attention 4 8 1024 × 5120
blk.0.ffn_down.weight FFN 4 8 5120 × 17408

Total trainable LoRA values: 135,360.

Quality result

This was a tiny proof run, not a quality fine-tune.

  • Train mean-token NLL, full minus zero adapter: 0.00021834747786
  • Validation mean-token NLL, full minus zero adapter: 9.73401665583e-06
  • Validation improved: no

The adapter is numerically connected and stable, but the proof run is too small to establish a quality gain.

Performance snapshot

Measurement Result
Base end-to-end generation 3.692 tokens/s
Trained-adapter end-to-end generation 3.379 tokens/s
Base fresh-process load 0.817 GiB/s
Trained fresh-process load 0.748 GiB/s
Peak process RSS 3980.6 MiB
Peak observed GPU memory 9972.0 MiB
Load/unload cycles 3
Quick soak 192 requested tokens
Training microsteps/minute 0.843330
Training updates/minute 0.562220

Primary artifacts

Artifact Size SHA256
step10_adapter_final.gguf 529.50 KiB 6ba256d4e1836a177a8c7519d05a95639d47f4ee181fb6fcd438eee40805549f
multi_target_adapter_initial.gguf 529.50 KiB 84d835c964c4a0ae32fed9be2106124ff9319b74317ba9cdb0ea21069150edb5
checkpoint_latest.p10ck 2.07 MiB 1ba5545af5c33aacd7c1d2b395afe0e53e4fac5f1fd07ea56353c5a0c9fa64b9
quick_resume_adapter.gguf 529.50 KiB c6c163276a50c144315543a49ecdccd725608b5e3a653733a4a589369e0233d2
quick_resume_checkpoint.p10ck 2.07 MiB 7c8e9b000eac8fa3ca60883b76f6a6c2ffe6126d97b3c861dfef70dcc4e3080c

Base model identity

  • Repository: prism-ml/Bonsai-27B-gguf
  • Revision: 0469926cd8878dbddb5b883740ab3df060058696
  • Required file: Bonsai-27B-Q1_0.gguf
  • Bytes: 3803452480
  • SHA256: 17ef842e47450caeb8eaa3ebfbbab5d2f2278b62b79be107985fb69a2f819aa0
python -m pip install -r reproduce/requirements.txt
python reproduce/download_models.py
python reproduce/restore_and_build.py
python reproduce/verify_release.py

Load the adapter

./binaries/linux-colab-cuda/llama-completion   -m /path/to/Bonsai-27B-Q1_0.gguf   --lora adapters/step10_adapter_final.gguf   -p "User: Explain deterministic checkpoint resume.\nAssistant:"   -n 64 --temp 0 -ngl 999 -no-cnv

Limitations

  • This is a proof-of-training release, not a broadly trained production LoRA.
  • Validation NLL was slightly worse by 9.73401665583e-06.
  • The Linux build is environment-specific.
  • The quick soak is not an hours-long production endurance test.
  • Upstream source and base-model licenses and terms apply.