--- license: apache-2.0 language: - en pipeline_tag: text-generation library_name: peft base_model: Qwen/Qwen3.5-0.8B base_model_relation: adapter tags: - base_model:adapter:Qwen/Qwen3.5-0.8B - lora - sft - transformers - trl - unsloth szl: doctrine: v11-LOCKED lean: 749/14/163 lambda: Conjecture 1 — advisory, never a theorem artifact_class: ADAPTER originality: FINETUNE_DISCLOSED_BASE sku: CHASKI-R2 quant: bf16-lora qlora: false weights: AVAILABLE evals: none-this-run publication_eligible: false never_overwrite: SZLHOLDINGS/chaski --- # Chaski-R2 adapter (bf16 LoRA) Owner-GPU recut on NVIDIA GeForce RTX 5050 Laptop (8GB). Original SZL cut of disclosed Apache `Qwen/Qwen3.5-0.8B`. **Not QLoRA.** Unsloth 2026-08 does not recommend QLoRA on Qwen3.5 (dense or MoE) because of higher-than-normal quantization differences. This is a **separate SKU**. It does **not** overwrite live `SZLHOLDINGS/chaski` and is **not** `SZLHOLDINGS/chaski-5050` (that kit is r=16 α=16 on doctrine SFT). This SKU is r=16 α=32 on `chaski_r2/train.jsonl` only. ## Honest status ## The cut Round-2 is a first-class citizen in this estate. We do not overwrite R1. We add a sibling. A lineage you can walk. R1 stays up. R2 is the next knot. ### Silhouette → leave → SZL | Leader | Take, then tweak | |---|---| | Anthropic | Versioned constitutions. | | NVIDIA | Recipe rerun. | | Unsloth | Another FastLanguageModel job. | Nobody else ships this combination. That is the point of a one-of-one. ## Intended use Lineage walk. Compare, do not silently replace. ## Limitations - proposal-only Canonical GitHub: [`szl-holdings/szl-forge`](https://github.com/szl-holdings/szl-forge/blob/main/chaski/) | | | |---|---| | **Base** | `Qwen/Qwen3.5-0.8B` | | **Method** | Unsloth bf16 LoRA (`load_in_4bit=False`, `load_in_16bit=True`) | | **LoRA** | r=16, α=32, seed 11, response-only CE | | **Dataset** | `chaski_r2/train.jsonl` (32 rows). Named-N gates held out of gradients. | | **Epochs / steps** | 3 epochs, batch 1, grad accum 2 | | **Train loss** | MEASURED `0.7656` — train metric, **not an eval** | | **Train runtime** | MEASURED on RTX 5050 Laptop 8GB | | **Adapter sha256** | `440340ce29e19344c0625d0adfe820b277cdb0e24099d4e612f88ad6b3cf49c6` | | **Evals** | see `training_receipt.local.json`; do not treat train loss as JSON-draft/refusal | | **publication_eligible** | Hub PUT of adapter bytes is LIVE; eval gates remain labeled in the receipt | | **Jobs** | local-5050 owner metal; HF Jobs not fired from the GitHub kit | | **Ollama / llama-server** | `llama-server` is missing. No tok/s claimed. | Train loss is not a JSON-draft or refusal gate. Not 5/5 or 6/6. Lab load forbidden. House CPU lab stays signed Khipu GGUF. ### Framework versions - PEFT 0.19.1