--- license: apache-2.0 language: - en - ru tags: - code - documentation - summarization - distillation - qwen3 - gguf - llama.cpp pipeline_tag: text-generation base_model: Qwen/Qwen3-4B library_name: gguf --- # code-daemon-summary-v1 Compact **bilingual (English / Russian) code-documentation generator** — a 4B GGUF model that writes: - **one-sentence entity descriptions** for a source file (functions, classes, fields) as a clean markdown bullet list; - **module overviews** — short prose plus ASCII architecture / data-flow diagrams; - **hierarchical codebase summaries** — subsystem, product-level and whole-project digests built from smaller summaries. Runs anywhere llama.cpp runs. Available quants: `Q5_K_M` (~2.8 GB, highest fidelity), `Q4_K_M` (~2.4 GB, recommended), `Q3_K_M` (~2.0 GB, for 6 GB-VRAM GPUs). The output language follows the request — both languages were distilled first-class, not translated. | Task | Output shape | |---|---| | Entity documentation | `- **Name**: one-sentence description.` per entity | | Module overview | `## Overview` prose + `## Architecture` / `## Flow` ASCII diagrams | | Hierarchical summaries | paragraph-length subsystem/product/strategy digests | ## Quick start (llama.cpp) The bundled chat template already pins the model's **non-thinking mode** — use it as a normal ChatML model. If you build raw prompts yourself, end them with the assistant tag followed by an empty think block (`<|im_start|>assistant\n\n\n\n\n`) — that is the format the model was trained with. Greedy decoding (temperature 0) recommended; stop on `<|im_end|>`. ```bash llama-cli -m code-daemon-summary-v1-Q4_K_M.gguf -c 8192 --temp 0 \ -p '<|im_start|>system You write one-sentence descriptions for code entities of a single file. Output ONLY a markdown bullet list, ONE bullet per entity: - ****: .<|im_end|> <|im_start|>user Entities: parseArray, encodeValue. File excerpt: <...><|im_end|> <|im_start|>assistant ' ``` ## Evaluation Held-out prompts via a deterministic content-hash split (provably never trained on), sampled across the full bilingual corpus; reference = the 7B teacher's output on the same prompts. | Slice | n | ROUGE-L | sem-cos* | empty outputs | |---|---|---|---|---| | Entity docs, mixed EN+RU | 300 | 0.618 | **0.904** | 0 | | Entity docs, EN-only sample | 300 | 0.636 | 0.897 | 0 | | Hierarchical summaries | 7 | 0.331 | 0.893 | 0 | \* all-MiniLM-L6-v2 cosine between student and teacher outputs — the paraphrase-aware signal; token metrics understate free-form prose. **Language fidelity:** 56/57 Russian-reference examples answered in Russian, 242/243 English in English. Residual gap to the teacher: slightly more verbose (len-ratio ~1.3–1.6). ## Speed Measured on a laptop RTX 5060 (8 GB), llama.cpp CUDA, Q4_K_M, `n_ctx=8192`, `n_batch=2048`, single stream — the way this model is actually used, since a hierarchical summary is one prompt rolling up its children and has nothing to batch against. | Workload | Input → output | Rate | |---|---|---| | Hierarchical summary | 250 → 250 tok | **103 tok/s** | | Long-form entity doc | 900 → 1500 tok | **95 tok/s** | | Prefill (both) | — | **~3 300 tok/s** | **Memory:** ~4.7 GB resident at `n_ctx=8192` — 2.4 GB weights plus KV and compute buffers. On an 8 GB card that leaves room for one small co-resident model, not two; decode collapses to roughly a third of these rates if the card is oversubscribed, so give it the headroom rather than the extra worker. ## How it was made - **Base:** [`Qwen/Qwen3-4B`](https://huggingface.co/Qwen/Qwen3-4B) (Apache-2.0), 36 layers, ChatML, 151 936-token vocab. - **Teacher:** [`Qwen/Qwen2.5-7B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct). - **Method:** sequence-level knowledge distillation (SeqKD) — LoRA SFT on ~11K teacher `(prompt → response)` traces, 3 epochs, prompt tokens masked, then merged into the base. An on-policy DistiLLM-2 refinement pass was also trained and **rejected on evaluation** (it degraded long-prompt behaviour); v1 is the SeqKD checkpoint. - The model powers the long-output documentation stages of a code-intelligence daemon; it is a purpose-built component, not a general assistant — outside this task distribution its behaviour is undefined. ## Training data Teacher traces generated over a mixed corpus: ~90 open-source repositories (Zig, C/C++, C#, TypeScript/JS, Python, Go, Rust, Kotlin, Swift, Java) for English and a commercial C#/TS/Python codebase for Russian. ~11K kept traces after dedup and corruption filtering. No third-party labeled dataset is used. ## License & attribution **Apache-2.0** — matches the Qwen3-4B base and the Qwen2.5-7B-Instruct teacher (both Alibaba / Qwen team, Apache-2.0). Not legal advice. Base and teacher © the Qwen team; please also honour their model cards.