c7w commited on
Commit
e2ce9be
·
verified ·
1 Parent(s): ab453ba

Update English model card

Browse files
Files changed (1) hide show
  1. README.md +18 -9
README.md CHANGED
@@ -18,10 +18,8 @@ tags:
18
 
19
  This is the code-domain teacher in the Open-MOPD pipeline. It starts from
20
  `BytedTsinghua-SIA/Open-MOPD-SmolLM3-3B-MixSFT` and is trained only on code
21
- prompts with verifiable rewards using GRPO.
22
-
23
- - Project page: https://bytedtsinghua-sia.github.io/Open-MOPD/
24
- - Code: https://github.com/BytedTsinghua-SIA/Open-MOPD
25
 
26
  Training uses global batch size 128, mini-batch size 32, learning rate `1e-6`,
27
  rollout group size 16, a 30,000-token response limit, no KL penalty, and
@@ -34,7 +32,9 @@ accuracy-based group filtering.
34
  | **RL-Code teacher** | **22.16** | **21.31** | **21.73** |
35
  | MixSFT starting point | 15.99 | 19.20 | 17.60 |
36
 
37
- Results are the mean over 10 samples per problem with temperature 1.0.
 
 
38
 
39
  The code portion of the RL prompt mixture explicitly excludes LiveCodeBench.
40
  The decontamination record is available in
@@ -50,8 +50,17 @@ tokenizer = AutoTokenizer.from_pretrained(model_id)
50
  model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")
51
  ```
52
 
53
- ## Role and scope
 
 
 
 
 
 
54
 
55
- This checkpoint is the routed code teacher used by Open-MOPD. The paper
56
- reports its code results above and does not make claims about its performance
57
- outside the code domain.
 
 
 
 
18
 
19
  This is the code-domain teacher in the Open-MOPD pipeline. It starts from
20
  `BytedTsinghua-SIA/Open-MOPD-SmolLM3-3B-MixSFT` and is trained only on code
21
+ prompts with verifiable rewards using GRPO. This release corresponds to
22
+ training step 180.
 
 
23
 
24
  Training uses global batch size 128, mini-batch size 32, learning rate `1e-6`,
25
  rollout group size 16, a 30,000-token response limit, no KL penalty, and
 
32
  | **RL-Code teacher** | **22.16** | **21.31** | **21.73** |
33
  | MixSFT starting point | 15.99 | 19.20 | 17.60 |
34
 
35
+ Results use avg@10 rather than best@10, with temperature 1.0,
36
+ `max_model_len=32768`, `top_p=0.95`, `top_k=-1`, and
37
+ `stop_token_ids=[128012]`.
38
 
39
  The code portion of the RL prompt mixture explicitly excludes LiveCodeBench.
40
  The decontamination record is available in
 
50
  model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")
51
  ```
52
 
53
+ ## Intended use and limitations
54
+
55
+ This is a domain teacher intended for distillation, not a general-purpose
56
+ assistant. It was optimized only on code and can perform worse than MixSFT on
57
+ other domains.
58
+
59
+ ## Model specifications
60
 
61
+ - Architecture: `SmolLM3ForCausalLM`
62
+ - Parameters: approximately 3B
63
+ - Layers: 36
64
+ - Vocabulary size: 128,256
65
+ - Weights: BF16, approximately 6.2 GB
66
+ - Includes tokenizer and chat template