Kiwen-27B / README.md
beyoru's picture
Update README.md
a903c7f verified
|
Raw
History Blame Contribute Delete
4.86 kB
---
license: apache-2.0
language:
- vi
- en
- zh
base_model:
- Qwen/Qwen3.8-27B
library_name: transformers
tags:
- legal
- kimi-k3
- long-reasoning
- efficient-reasoning
- coding
- math
- stem
---
# Kiwen-27B
A fine-tune of [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) trained on long chain-of-thought traces from Kimi K3.
This is **not a broadly stronger model**. The fine-tune primarily teaches Qwen3.8-27B to terminate its reasoning and produce an answer before exhausting its generation budget.
## Model Summary
Qwen3.8-27B is a strong reasoning model that can sometimes spend its entire token budget deliberating without producing a final answer.
On GSM8K, this behavioral change produces a **+10.31 percentage-point improvement** under flexible extraction.
### Key results
| Benchmark | Qwen3.8-27B | **Kiwen-27B** | Delta | t |
| ---------------------------------------- | ----------: | ---------------: | ---------: | --------: |
| **GSM8K** exact_match (flexible) | 0.7407 | **0.8438** | **+10.31** | **+6.58** |
| **GSM8K** exact_match (strict) | 0.6672 | **0.7278** | **+6.07** | **+3.40** |
| IFEval prompt-level strict | 0.8226 | **0.8429** | +2.03 | +0.90 |
| IFEval prompt-level loose | 0.8447 | **0.8669** | +2.22 | +1.04 |
| IFEval instruction-level strict | 0.8393 | **0.8645** | +2.52 | — |
| IFEval instruction-level loose | 0.8537 | **0.8801** | +2.64 | — |
| VMLU val (744 Vietnamese MMLU questions) | 83.47 | **86.02** | +2.55 | +1.37 |
Evaluated using [lm-evaluation-harness](https://github.com/EleutherAI/lm-evaluation-harness) 0.4.12 with both models served on identical hardware:
* **Hardware:** 1× H200
* **Inference:** SGLang
* **Model dtype:** BF16
* **KV cache:** FP8
* **Generation limit:** `max_gen_toks=4096`
* **Prompts:** identical
* **Decoding settings:** identical
The largest improvement is on GSM8K. Other benchmarks show smaller but consistently positive changes, with no measured regression.
## The Mechanism: Learning When to Stop
GSM8K is evaluated using two extraction methods:
* **Strict match:** requires the canonical `#### <number>` answer format.
* **Flexible extraction:** extracts the final number from the generated response.
The improvement under flexible extraction is substantially larger:
> **+10.31 points flexible vs. +6.07 points strict**
This asymmetry suggests that much of the improvement comes from **successfully producing an answer before the generation budget is exhausted**, rather than from a fundamental improvement in mathematical reasoning.
## Out-of-Domain Regression
To test whether the fine-tuning damaged unrelated capabilities, both models were evaluated on 100 held-out Vietnamese enterprise documents covering:
* Accounting
* CRM
* Internal RAG workloads
The metric was token-level negative log-likelihood:
| Model | NLL / token |
| ------------ | ----------: |
| Qwen3.8-27B | 3.3390 |
| Kiwen-27B | 3.3543 |
Paired difference:
* **Δ:** +0.0067
* **Standard error:** 0.0041
* **t:** 1.65
## Usage
```python
from transformers import AutoModelForImageTextToText, AutoProcessor
model = AutoModelForImageTextToText.from_pretrained(
"beyoru/Kiwen-27B",
dtype="bfloat16",
device_map="auto",
)
proc = AutoProcessor.from_pretrained(
"beyoru/Kiwen-27B"
)
```
### MTP Tensors
MTP tensors are absent from the released model.
`transformers` declares:
```python
_keys_to_ignore_on_load_unexpected = [r"^mtp.*"]
```
Therefore, the multi-token-prediction block is dropped by HF-based merging, including this release.
For speculative decoding, use an external draft model such as [z-lab/Qwen3.8-27B-DFlash2](https://huggingface.co/z-lab/Qwen3.8-27B-DFlash2) rather than NEXTN.
A draft model trained against the original weights may have a lower acceptance rate against the merged Kiwen weights. This does **not** affect correctness because speculative decoding uses rejection verification.
### Benchmark Caveat
VMLU validation is publicly distributed with answer keys and may have some representation in pretraining data.
Its absolute score should therefore be interpreted cautiously.
However, the **base-versus-fine-tuned comparison remains informative**, since both models are exposed to the same potential contamination.
## Citation
```bibtex
@misc{kiwen27bk3,
title = {Kiwen-27B},
author = {beyoru},
year = {2026},
url = {https://huggingface.co/beyoru/Kiwen-27B-K3}
}
```
## License & Attribution
Built on:
* [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) — Apache-2.0
Kiwen-27B is released under the **Apache-2.0** license.