Qwen3-4B-R1-SFT / README.md
DrEntropy's picture
Qwen3-4B-Instruct-2507 SFT on rasbt R1-trace dataset (9549 ex, LoRA + trainable <think> tokens; merged)
f90926e verified
|
Raw
History Blame Contribute Delete
463 Bytes
metadata
base_model: Qwen/Qwen3-4B-Instruct-2507
library_name: peft
pipeline_tag: text-generation
tags:
  - base_model:adapter:Qwen/Qwen3-4B-Instruct-2507
  - lora
  - sft
  - transformers
  - trl

Model Card for Model ID

This is a a experimental research artifact only. Trained on rasbt/math_distill/data/deepseek-r1-math-train_4000.json

Out-of-Scope Use

THIS IS RESEARCH ARTIFACT and should not intended for use.

Framework versions

  • PEFT 0.19.1