Commit
498a838
·
verified ·
1 Parent(s): b37328b

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +35 -0
README.md ADDED
@@ -0,0 +1,35 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen3-0.6B
3
+ language:
4
+ - en
5
+ library_name: transformers
6
+ license: apache-2.0
7
+ pipeline_tag: text-generation
8
+ tags:
9
+ - metacognitive-behavioral-tuning
10
+ - multi-hop-qa
11
+ - reasoning
12
+ - sft
13
+ - grpo
14
+ - MBT-R
15
+ ---
16
+
17
+ # Qwen3-0.6B · MBT-R
18
+
19
+ MBT-R (main table) — **final checkpoint (SFT → GRPO)**. Base: [`Qwen/Qwen3-0.6B`](https://huggingface.co/Qwen/Qwen3-0.6B).
20
+ Paper: *Metacognitive Behavioral Tuning of Large Language Models for Multi-Hop Question Answering*.
21
+
22
+ - **Method**: MBT-R (Refinement): the student's own reasoning traces are rewritten into the 5-phase structure for SFT, then GRPO.
23
+ - **Base model**: `Qwen/Qwen3-0.6B`
24
+ - **Training**: SFT (LR 1e-4, BS 128, HotpotQA) → GRPO
25
+ - **Benchmarks**: HotpotQA (ID), MuSiQue / 2WikiMultiHopQA (OOD)
26
+
27
+ ## Usage
28
+
29
+ ```python
30
+ from transformers import AutoModelForCausalLM, AutoTokenizer
31
+
32
+ repo = "metacognitive-behavioral-tuning/Qwen3-0.6B-MBT-R"
33
+ tok = AutoTokenizer.from_pretrained(repo)
34
+ model = AutoModelForCausalLM.from_pretrained(repo, dtype="bfloat16", device_map="auto")
35
+ ```