Qwen3-4B-R1-SFT / README.md
DrEntropy's picture
Qwen3-4B-Instruct-2507 SFT on rasbt R1-trace dataset (9549 ex, LoRA + trainable <think> tokens; merged)
f90926e verified
|
Raw
History Blame Contribute Delete
463 Bytes
---
base_model: Qwen/Qwen3-4B-Instruct-2507
library_name: peft
pipeline_tag: text-generation
tags:
- base_model:adapter:Qwen/Qwen3-4B-Instruct-2507
- lora
- sft
- transformers
- trl
---
# Model Card for Model ID
This is a a **experimental** research artifact only.
Trained on rasbt/math_distill/data/deepseek-r1-math-train_4000.json
### Out-of-Scope Use
THIS IS **RESEARCH ARTIFACT** and should not intended for use.
### Framework versions
- PEFT 0.19.1