Qwen3-0.6B-JSON-SFT-GRPO

์œ„ SFT ๋ชจ๋ธ์— GRPO(RLVR) ์ถ”๊ฐ€ ํ•™์Šต (์‹œ๋‚˜๋ฆฌ์˜ค 1 ์ตœ์ข…)

2026-07-19 sllm_practice ์ปค๋ฆฌํ˜๋Ÿผ ๊ฐœ์„  ์‹คํ—˜ ์‚ฐ์ถœ๋ฌผ (fa2e1ce์˜ experiments/results/REPORT.md ์ฐธ์กฐ). ํ‰๊ฐ€ ์ง€ํ‘œ๋Š” json_eval.py (parse_rate / schema_compliance / semantics_pass_rate / field_f1_mean).

LoRA r16/alpha32, num_generations 4, beta 0.04, 600 ํ”„๋กฌํ”„ํŠธ, ๋ณ‘ํ•ฉ๋ณธ. old_eval field_F1 54.4% (SFT 51.4% ๋Œ€๋น„ +3.0%p), eval_reward 1.48โ†’1.89. ํ•™์Šต 62.9๋ถ„(A40).

Downloads last month
26
Safetensors
Model size
0.6B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for NotoriousH2/Qwen3-0.6B-JSON-SFT-GRPO

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1142)
this model