PEFT
Safetensors
bimabk commited on
Commit
1f55afe
·
verified ·
1 Parent(s): d2cff3c

Upload task output 1

Browse files
README.md ADDED
@@ -0,0 +1,202 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: None
3
+ library_name: peft
4
+ ---
5
+
6
+ # Model Card for Model ID
7
+
8
+ <!-- Provide a quick summary of what the model is/does. -->
9
+
10
+
11
+
12
+ ## Model Details
13
+
14
+ ### Model Description
15
+
16
+ <!-- Provide a longer summary of what this model is. -->
17
+
18
+
19
+
20
+ - **Developed by:** [More Information Needed]
21
+ - **Funded by [optional]:** [More Information Needed]
22
+ - **Shared by [optional]:** [More Information Needed]
23
+ - **Model type:** [More Information Needed]
24
+ - **Language(s) (NLP):** [More Information Needed]
25
+ - **License:** [More Information Needed]
26
+ - **Finetuned from model [optional]:** [More Information Needed]
27
+
28
+ ### Model Sources [optional]
29
+
30
+ <!-- Provide the basic links for the model. -->
31
+
32
+ - **Repository:** [More Information Needed]
33
+ - **Paper [optional]:** [More Information Needed]
34
+ - **Demo [optional]:** [More Information Needed]
35
+
36
+ ## Uses
37
+
38
+ <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
39
+
40
+ ### Direct Use
41
+
42
+ <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
43
+
44
+ [More Information Needed]
45
+
46
+ ### Downstream Use [optional]
47
+
48
+ <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
49
+
50
+ [More Information Needed]
51
+
52
+ ### Out-of-Scope Use
53
+
54
+ <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
55
+
56
+ [More Information Needed]
57
+
58
+ ## Bias, Risks, and Limitations
59
+
60
+ <!-- This section is meant to convey both technical and sociotechnical limitations. -->
61
+
62
+ [More Information Needed]
63
+
64
+ ### Recommendations
65
+
66
+ <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
67
+
68
+ Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
69
+
70
+ ## How to Get Started with the Model
71
+
72
+ Use the code below to get started with the model.
73
+
74
+ [More Information Needed]
75
+
76
+ ## Training Details
77
+
78
+ ### Training Data
79
+
80
+ <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
81
+
82
+ [More Information Needed]
83
+
84
+ ### Training Procedure
85
+
86
+ <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
87
+
88
+ #### Preprocessing [optional]
89
+
90
+ [More Information Needed]
91
+
92
+
93
+ #### Training Hyperparameters
94
+
95
+ - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
96
+
97
+ #### Speeds, Sizes, Times [optional]
98
+
99
+ <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
100
+
101
+ [More Information Needed]
102
+
103
+ ## Evaluation
104
+
105
+ <!-- This section describes the evaluation protocols and provides the results. -->
106
+
107
+ ### Testing Data, Factors & Metrics
108
+
109
+ #### Testing Data
110
+
111
+ <!-- This should link to a Dataset Card if possible. -->
112
+
113
+ [More Information Needed]
114
+
115
+ #### Factors
116
+
117
+ <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
118
+
119
+ [More Information Needed]
120
+
121
+ #### Metrics
122
+
123
+ <!-- These are the evaluation metrics being used, ideally with a description of why. -->
124
+
125
+ [More Information Needed]
126
+
127
+ ### Results
128
+
129
+ [More Information Needed]
130
+
131
+ #### Summary
132
+
133
+
134
+
135
+ ## Model Examination [optional]
136
+
137
+ <!-- Relevant interpretability work for the model goes here -->
138
+
139
+ [More Information Needed]
140
+
141
+ ## Environmental Impact
142
+
143
+ <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
144
+
145
+ Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
146
+
147
+ - **Hardware Type:** [More Information Needed]
148
+ - **Hours used:** [More Information Needed]
149
+ - **Cloud Provider:** [More Information Needed]
150
+ - **Compute Region:** [More Information Needed]
151
+ - **Carbon Emitted:** [More Information Needed]
152
+
153
+ ## Technical Specifications [optional]
154
+
155
+ ### Model Architecture and Objective
156
+
157
+ [More Information Needed]
158
+
159
+ ### Compute Infrastructure
160
+
161
+ [More Information Needed]
162
+
163
+ #### Hardware
164
+
165
+ [More Information Needed]
166
+
167
+ #### Software
168
+
169
+ [More Information Needed]
170
+
171
+ ## Citation [optional]
172
+
173
+ <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
174
+
175
+ **BibTeX:**
176
+
177
+ [More Information Needed]
178
+
179
+ **APA:**
180
+
181
+ [More Information Needed]
182
+
183
+ ## Glossary [optional]
184
+
185
+ <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
186
+
187
+ [More Information Needed]
188
+
189
+ ## More Information [optional]
190
+
191
+ [More Information Needed]
192
+
193
+ ## Model Card Authors [optional]
194
+
195
+ [More Information Needed]
196
+
197
+ ## Model Card Contact
198
+
199
+ [More Information Needed]
200
+ ### Framework versions
201
+
202
+ - PEFT 0.15.1
adapter_config.json ADDED
@@ -0,0 +1,39 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alpha_pattern": {},
3
+ "auto_mapping": null,
4
+ "base_model_name_or_path": null,
5
+ "bias": "none",
6
+ "corda_config": null,
7
+ "eva_config": null,
8
+ "exclude_modules": null,
9
+ "fan_in_fan_out": false,
10
+ "inference_mode": true,
11
+ "init_lora_weights": true,
12
+ "layer_replication": null,
13
+ "layers_pattern": null,
14
+ "layers_to_transform": null,
15
+ "loftq_config": {},
16
+ "lora_alpha": 512,
17
+ "lora_bias": false,
18
+ "lora_dropout": 0.1,
19
+ "megatron_config": null,
20
+ "megatron_core": "megatron.core",
21
+ "modules_to_save": null,
22
+ "peft_type": "LORA",
23
+ "r": 128,
24
+ "rank_pattern": {},
25
+ "revision": null,
26
+ "target_modules": [
27
+ "v_proj",
28
+ "q_proj",
29
+ "o_proj",
30
+ "down_proj",
31
+ "k_proj",
32
+ "up_proj",
33
+ "gate_proj"
34
+ ],
35
+ "task_type": "CAUSAL_LM",
36
+ "trainable_token_indices": null,
37
+ "use_dora": false,
38
+ "use_rslora": false
39
+ }
model.safetensors → adapter_model.safetensors RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:53447a1435746b351d2547246976f23d4bd67be8e89273c04d9c8ab278455af4
3
- size 3087467144
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e726bb5d630d0e328f9dadab3f3514fda4eb8025c341ab9daa92b90f54cff7a8
3
+ size 1279323952
added_tokens.json DELETED
@@ -1,24 +0,0 @@
1
- {
2
- "</tool_call>": 151658,
3
- "<tool_call>": 151657,
4
- "<|box_end|>": 151649,
5
- "<|box_start|>": 151648,
6
- "<|endoftext|>": 151643,
7
- "<|file_sep|>": 151664,
8
- "<|fim_middle|>": 151660,
9
- "<|fim_pad|>": 151662,
10
- "<|fim_prefix|>": 151659,
11
- "<|fim_suffix|>": 151661,
12
- "<|im_end|>": 151645,
13
- "<|im_start|>": 151644,
14
- "<|image_pad|>": 151655,
15
- "<|object_ref_end|>": 151647,
16
- "<|object_ref_start|>": 151646,
17
- "<|quad_end|>": 151651,
18
- "<|quad_start|>": 151650,
19
- "<|repo_name|>": 151663,
20
- "<|video_pad|>": 151656,
21
- "<|vision_end|>": 151653,
22
- "<|vision_pad|>": 151654,
23
- "<|vision_start|>": 151652
24
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
config.json DELETED
@@ -1,59 +0,0 @@
1
- {
2
- "architectures": [
3
- "Qwen2ForCausalLM"
4
- ],
5
- "attention_dropout": 0.0,
6
- "bos_token_id": 151643,
7
- "dtype": "float16",
8
- "eos_token_id": 151645,
9
- "hidden_act": "silu",
10
- "hidden_size": 1536,
11
- "initializer_range": 0.02,
12
- "intermediate_size": 8960,
13
- "layer_types": [
14
- "full_attention",
15
- "full_attention",
16
- "full_attention",
17
- "full_attention",
18
- "full_attention",
19
- "full_attention",
20
- "full_attention",
21
- "full_attention",
22
- "full_attention",
23
- "full_attention",
24
- "full_attention",
25
- "full_attention",
26
- "full_attention",
27
- "full_attention",
28
- "full_attention",
29
- "full_attention",
30
- "full_attention",
31
- "full_attention",
32
- "full_attention",
33
- "full_attention",
34
- "full_attention",
35
- "full_attention",
36
- "full_attention",
37
- "full_attention",
38
- "full_attention",
39
- "full_attention",
40
- "full_attention",
41
- "full_attention"
42
- ],
43
- "max_position_embeddings": 32768,
44
- "max_window_layers": 21,
45
- "model_type": "qwen2",
46
- "num_attention_heads": 12,
47
- "num_hidden_layers": 28,
48
- "num_key_value_heads": 2,
49
- "rms_norm_eps": 1e-06,
50
- "rope_scaling": null,
51
- "rope_theta": 1000000.0,
52
- "sliding_window": null,
53
- "tie_word_embeddings": true,
54
- "torch_dtype": "bfloat16",
55
- "transformers_version": "4.51.3",
56
- "use_cache": true,
57
- "use_sliding_window": false,
58
- "vocab_size": 151936
59
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
generation_config.json DELETED
@@ -1,5 +0,0 @@
1
- {
2
- "temperature": null,
3
- "top_p": null,
4
- "transformers_version": "4.51.3"
5
- }
 
 
 
 
 
 
loss.txt ADDED
@@ -0,0 +1 @@
 
 
1
+ 448,no_eval
merges.txt DELETED
The diff for this file is too large to render. See raw diff
 
special_tokens_map.json CHANGED
@@ -1,28 +1,27 @@
1
  {
2
  "additional_special_tokens": [
3
- "<|im_start|>",
4
- "<|im_end|>",
5
- "<|object_ref_start|>",
6
- "<|object_ref_end|>",
7
- "<|box_start|>",
8
- "<|box_end|>",
9
- "<|quad_start|>",
10
- "<|quad_end|>",
11
- "<|vision_start|>",
12
- "<|vision_end|>",
13
- "<|vision_pad|>",
14
- "<|image_pad|>",
15
- "<|video_pad|>"
16
  ],
 
 
 
 
 
 
 
17
  "eos_token": {
18
- "content": "<|im_end|>",
19
  "lstrip": false,
20
  "normalized": false,
21
  "rstrip": false,
22
  "single_word": false
23
  },
24
- "pad_token": {
25
- "content": "<|endoftext|>",
 
26
  "lstrip": false,
27
  "normalized": false,
28
  "rstrip": false,
 
1
  {
2
  "additional_special_tokens": [
3
+ "<PRE>",
4
+ "<MID>",
5
+ "<SUF>",
6
+ "<EOT>"
 
 
 
 
 
 
 
 
 
7
  ],
8
+ "bos_token": {
9
+ "content": "<s>",
10
+ "lstrip": false,
11
+ "normalized": false,
12
+ "rstrip": false,
13
+ "single_word": false
14
+ },
15
  "eos_token": {
16
+ "content": "</s>",
17
  "lstrip": false,
18
  "normalized": false,
19
  "rstrip": false,
20
  "single_word": false
21
  },
22
+ "pad_token": "</s>",
23
+ "unk_token": {
24
+ "content": "<unk>",
25
  "lstrip": false,
26
  "normalized": false,
27
  "rstrip": false,
tokenizer.json CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:9c5ae00e602b8860cbd784ba82a8aa14e8feecec692e7076590d014d7b7fdafa
3
- size 11421896
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2206073a6598893988ad5f1d96aa193d274fd6ad2a8d8c7ab8f56b29b6d4d0aa
3
+ size 3620731
tokenizer.model ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:45ccb9c8b6b561889acea59191d66986d314e7cbd6a78abc6e49b139ca91c1e6
3
+ size 500058
tokenizer_config.json CHANGED
@@ -1,208 +1,85 @@
1
  {
2
- "add_bos_token": false,
3
- "add_prefix_space": false,
4
  "added_tokens_decoder": {
5
- "151643": {
6
- "content": "<|endoftext|>",
7
  "lstrip": false,
8
  "normalized": false,
9
  "rstrip": false,
10
  "single_word": false,
11
  "special": true
12
  },
13
- "151644": {
14
- "content": "<|im_start|>",
15
  "lstrip": false,
16
  "normalized": false,
17
  "rstrip": false,
18
  "single_word": false,
19
  "special": true
20
  },
21
- "151645": {
22
- "content": "<|im_end|>",
23
  "lstrip": false,
24
  "normalized": false,
25
  "rstrip": false,
26
  "single_word": false,
27
  "special": true
28
  },
29
- "151646": {
30
- "content": "<|object_ref_start|>",
31
  "lstrip": false,
32
  "normalized": false,
33
  "rstrip": false,
34
  "single_word": false,
35
  "special": true
36
  },
37
- "151647": {
38
- "content": "<|object_ref_end|>",
39
  "lstrip": false,
40
  "normalized": false,
41
  "rstrip": false,
42
  "single_word": false,
43
  "special": true
44
  },
45
- "151648": {
46
- "content": "<|box_start|>",
47
  "lstrip": false,
48
  "normalized": false,
49
  "rstrip": false,
50
  "single_word": false,
51
  "special": true
52
  },
53
- "151649": {
54
- "content": "<|box_end|>",
55
  "lstrip": false,
56
  "normalized": false,
57
  "rstrip": false,
58
  "single_word": false,
59
  "special": true
60
- },
61
- "151650": {
62
- "content": "<|quad_start|>",
63
- "lstrip": false,
64
- "normalized": false,
65
- "rstrip": false,
66
- "single_word": false,
67
- "special": true
68
- },
69
- "151651": {
70
- "content": "<|quad_end|>",
71
- "lstrip": false,
72
- "normalized": false,
73
- "rstrip": false,
74
- "single_word": false,
75
- "special": true
76
- },
77
- "151652": {
78
- "content": "<|vision_start|>",
79
- "lstrip": false,
80
- "normalized": false,
81
- "rstrip": false,
82
- "single_word": false,
83
- "special": true
84
- },
85
- "151653": {
86
- "content": "<|vision_end|>",
87
- "lstrip": false,
88
- "normalized": false,
89
- "rstrip": false,
90
- "single_word": false,
91
- "special": true
92
- },
93
- "151654": {
94
- "content": "<|vision_pad|>",
95
- "lstrip": false,
96
- "normalized": false,
97
- "rstrip": false,
98
- "single_word": false,
99
- "special": true
100
- },
101
- "151655": {
102
- "content": "<|image_pad|>",
103
- "lstrip": false,
104
- "normalized": false,
105
- "rstrip": false,
106
- "single_word": false,
107
- "special": true
108
- },
109
- "151656": {
110
- "content": "<|video_pad|>",
111
- "lstrip": false,
112
- "normalized": false,
113
- "rstrip": false,
114
- "single_word": false,
115
- "special": true
116
- },
117
- "151657": {
118
- "content": "<tool_call>",
119
- "lstrip": false,
120
- "normalized": false,
121
- "rstrip": false,
122
- "single_word": false,
123
- "special": false
124
- },
125
- "151658": {
126
- "content": "</tool_call>",
127
- "lstrip": false,
128
- "normalized": false,
129
- "rstrip": false,
130
- "single_word": false,
131
- "special": false
132
- },
133
- "151659": {
134
- "content": "<|fim_prefix|>",
135
- "lstrip": false,
136
- "normalized": false,
137
- "rstrip": false,
138
- "single_word": false,
139
- "special": false
140
- },
141
- "151660": {
142
- "content": "<|fim_middle|>",
143
- "lstrip": false,
144
- "normalized": false,
145
- "rstrip": false,
146
- "single_word": false,
147
- "special": false
148
- },
149
- "151661": {
150
- "content": "<|fim_suffix|>",
151
- "lstrip": false,
152
- "normalized": false,
153
- "rstrip": false,
154
- "single_word": false,
155
- "special": false
156
- },
157
- "151662": {
158
- "content": "<|fim_pad|>",
159
- "lstrip": false,
160
- "normalized": false,
161
- "rstrip": false,
162
- "single_word": false,
163
- "special": false
164
- },
165
- "151663": {
166
- "content": "<|repo_name|>",
167
- "lstrip": false,
168
- "normalized": false,
169
- "rstrip": false,
170
- "single_word": false,
171
- "special": false
172
- },
173
- "151664": {
174
- "content": "<|file_sep|>",
175
- "lstrip": false,
176
- "normalized": false,
177
- "rstrip": false,
178
- "single_word": false,
179
- "special": false
180
  }
181
  },
182
  "additional_special_tokens": [
183
- "<|im_start|>",
184
- "<|im_end|>",
185
- "<|object_ref_start|>",
186
- "<|object_ref_end|>",
187
- "<|box_start|>",
188
- "<|box_end|>",
189
- "<|quad_start|>",
190
- "<|quad_end|>",
191
- "<|vision_start|>",
192
- "<|vision_end|>",
193
- "<|vision_pad|>",
194
- "<|image_pad|>",
195
- "<|video_pad|>"
196
  ],
197
- "bos_token": null,
198
- "chat_template": "{%- if tools %}\n {{- '<|im_start|>system\\n' }}\n {%- if messages[0]['role'] == 'system' %}\n {{- messages[0]['content'] }}\n {%- else %}\n {{- 'You are Qwen, created by Alibaba Cloud. You are a helpful assistant.' }}\n {%- endif %}\n {{- \"\\n\\n# Tools\\n\\nYou may call one or more functions to assist with the user query.\\n\\nYou are provided with function signatures within <tools></tools> XML tags:\\n<tools>\" }}\n {%- for tool in tools %}\n {{- \"\\n\" }}\n {{- tool | tojson }}\n {%- endfor %}\n {{- \"\\n</tools>\\n\\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\\n<tool_call>\\n{\\\"name\\\": <function-name>, \\\"arguments\\\": <args-json-object>}\\n</tool_call><|im_end|>\\n\" }}\n{%- else %}\n {%- if messages[0]['role'] == 'system' %}\n {{- '<|im_start|>system\\n' + messages[0]['content'] + '<|im_end|>\\n' }}\n {%- else %}\n {{- '<|im_start|>system\\nYou are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>\\n' }}\n {%- endif %}\n{%- endif %}\n{%- for message in messages %}\n {%- if (message.role == \"user\") or (message.role == \"system\" and not loop.first) or (message.role == \"assistant\" and not message.tool_calls) %}\n {{- '<|im_start|>' + message.role + '\\n' + message.content + '<|im_end|>' + '\\n' }}\n {%- elif message.role == \"assistant\" %}\n {{- '<|im_start|>' + message.role }}\n {%- if message.content %}\n {{- '\\n' + message.content }}\n {%- endif %}\n {%- for tool_call in message.tool_calls %}\n {%- if tool_call.function is defined %}\n {%- set tool_call = tool_call.function %}\n {%- endif %}\n {{- '\\n<tool_call>\\n{\"name\": \"' }}\n {{- tool_call.name }}\n {{- '\", \"arguments\": ' }}\n {{- tool_call.arguments | tojson }}\n {{- '}\\n</tool_call>' }}\n {%- endfor %}\n {{- '<|im_end|>\\n' }}\n {%- elif message.role == \"tool\" %}\n {%- if (loop.index0 == 0) or (messages[loop.index0 - 1].role != \"tool\") %}\n {{- '<|im_start|>user' }}\n {%- endif %}\n {{- '\\n<tool_response>\\n' }}\n {{- message.content }}\n {{- '\\n</tool_response>' }}\n {%- if loop.last or (messages[loop.index0 + 1].role != \"tool\") %}\n {{- '<|im_end|>\\n' }}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n {{- '<|im_start|>assistant\\n' }}\n{%- endif %}\n",
199
  "clean_up_tokenization_spaces": false,
200
- "eos_token": "<|im_end|>",
201
- "errors": "replace",
202
  "extra_special_tokens": {},
203
- "model_max_length": 131072,
204
- "pad_token": "<|endoftext|>",
205
- "split_special_tokens": false,
206
- "tokenizer_class": "Qwen2Tokenizer",
207
- "unk_token": null
 
 
 
 
 
 
208
  }
 
1
  {
2
+ "add_bos_token": true,
3
+ "add_eos_token": false,
4
  "added_tokens_decoder": {
5
+ "0": {
6
+ "content": "<unk>",
7
  "lstrip": false,
8
  "normalized": false,
9
  "rstrip": false,
10
  "single_word": false,
11
  "special": true
12
  },
13
+ "1": {
14
+ "content": "<s>",
15
  "lstrip": false,
16
  "normalized": false,
17
  "rstrip": false,
18
  "single_word": false,
19
  "special": true
20
  },
21
+ "2": {
22
+ "content": "</s>",
23
  "lstrip": false,
24
  "normalized": false,
25
  "rstrip": false,
26
  "single_word": false,
27
  "special": true
28
  },
29
+ "32007": {
30
+ "content": "<PRE>",
31
  "lstrip": false,
32
  "normalized": false,
33
  "rstrip": false,
34
  "single_word": false,
35
  "special": true
36
  },
37
+ "32008": {
38
+ "content": "<SUF>",
39
  "lstrip": false,
40
  "normalized": false,
41
  "rstrip": false,
42
  "single_word": false,
43
  "special": true
44
  },
45
+ "32009": {
46
+ "content": "<MID>",
47
  "lstrip": false,
48
  "normalized": false,
49
  "rstrip": false,
50
  "single_word": false,
51
  "special": true
52
  },
53
+ "32010": {
54
+ "content": "<EOT>",
55
  "lstrip": false,
56
  "normalized": false,
57
  "rstrip": false,
58
  "single_word": false,
59
  "special": true
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
60
  }
61
  },
62
  "additional_special_tokens": [
63
+ "<PRE>",
64
+ "<MID>",
65
+ "<SUF>",
66
+ "<EOT>"
 
 
 
 
 
 
 
 
 
67
  ],
68
+ "bos_token": "<s>",
69
+ "chat_template": "{% if messages[0]['role'] == 'system' %}{% set loop_messages = messages[1:] %}{% set system_message = messages[0]['content'] %}{% else %}{% set loop_messages = messages %}{% set system_message = false %}{% endif %}{% for message in loop_messages %}{% if (message['role'] == 'user') != (loop.index0 % 2 == 0) %}{{ raise_exception('Conversation roles must alternate user/assistant/user/assistant/...') }}{% endif %}{% if loop.index0 == 0 and system_message != false %}{% set content = '<<SYS>>\\n' + system_message + '\\n<</SYS>>\\n\\n' + message['content'] %}{% else %}{% set content = message['content'] %}{% endif %}{% if message['role'] == 'user' %}{{ bos_token + '[INST] ' + content | trim + ' [/INST]' }}{% elif message['role'] == 'assistant' %}{{ ' ' + content | trim + ' ' + eos_token }}{% endif %}{% endfor %}",
70
  "clean_up_tokenization_spaces": false,
71
+ "eos_token": "</s>",
72
+ "eot_token": "▁<EOT>",
73
  "extra_special_tokens": {},
74
+ "fill_token": "<FILL_ME>",
75
+ "legacy": null,
76
+ "middle_token": "▁<MID>",
77
+ "model_max_length": 1000000000000000019884624838656,
78
+ "pad_token": "</s>",
79
+ "prefix_token": "▁<PRE>",
80
+ "sp_model_kwargs": {},
81
+ "suffix_token": "▁<SUF>",
82
+ "tokenizer_class": "CodeLlamaTokenizer",
83
+ "unk_token": "<unk>",
84
+ "use_default_system_prompt": false
85
  }
trainer_state.json ADDED
@@ -0,0 +1,746 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "best_global_step": null,
3
+ "best_metric": null,
4
+ "best_model_checkpoint": null,
5
+ "epoch": 0.47357293868921774,
6
+ "eval_steps": 500,
7
+ "global_step": 448,
8
+ "is_hyper_param_search": false,
9
+ "is_local_process_zero": true,
10
+ "is_world_process_zero": true,
11
+ "log_history": [
12
+ {
13
+ "epoch": 0.005285412262156448,
14
+ "grad_norm": 8.903882026672363,
15
+ "learning_rate": 9.166551619047618e-06,
16
+ "loss": 1.2181,
17
+ "num_tokens": 152231.0,
18
+ "step": 5
19
+ },
20
+ {
21
+ "epoch": 0.010570824524312896,
22
+ "grad_norm": 1.0083556175231934,
23
+ "learning_rate": 2.062474114285714e-05,
24
+ "loss": 0.6681,
25
+ "num_tokens": 317223.0,
26
+ "step": 10
27
+ },
28
+ {
29
+ "epoch": 0.015856236786469344,
30
+ "grad_norm": 0.746456503868103,
31
+ "learning_rate": 3.2082930666666666e-05,
32
+ "loss": 0.4249,
33
+ "num_tokens": 473767.0,
34
+ "step": 15
35
+ },
36
+ {
37
+ "epoch": 0.021141649048625793,
38
+ "grad_norm": 0.6790260076522827,
39
+ "learning_rate": 4.3541120190476185e-05,
40
+ "loss": 0.3982,
41
+ "num_tokens": 659649.0,
42
+ "step": 20
43
+ },
44
+ {
45
+ "epoch": 0.026427061310782242,
46
+ "grad_norm": 0.7228266596794128,
47
+ "learning_rate": 5.499930971428571e-05,
48
+ "loss": 0.3887,
49
+ "num_tokens": 832364.0,
50
+ "step": 25
51
+ },
52
+ {
53
+ "epoch": 0.03171247357293869,
54
+ "grad_norm": 0.6723021268844604,
55
+ "learning_rate": 6.645749923809523e-05,
56
+ "loss": 0.3467,
57
+ "num_tokens": 991414.0,
58
+ "step": 30
59
+ },
60
+ {
61
+ "epoch": 0.03699788583509514,
62
+ "grad_norm": 0.5942872166633606,
63
+ "learning_rate": 7.791568876190476e-05,
64
+ "loss": 0.3578,
65
+ "num_tokens": 1158859.0,
66
+ "step": 35
67
+ },
68
+ {
69
+ "epoch": 0.042283298097251586,
70
+ "grad_norm": 0.6299146413803101,
71
+ "learning_rate": 8.020446518214365e-05,
72
+ "loss": 0.3567,
73
+ "num_tokens": 1312426.0,
74
+ "step": 40
75
+ },
76
+ {
77
+ "epoch": 0.04756871035940803,
78
+ "grad_norm": 0.43629753589630127,
79
+ "learning_rate": 8.0192841334398e-05,
80
+ "loss": 0.3251,
81
+ "num_tokens": 1488799.0,
82
+ "step": 45
83
+ },
84
+ {
85
+ "epoch": 0.052854122621564484,
86
+ "grad_norm": 0.7118616700172424,
87
+ "learning_rate": 8.017227973373715e-05,
88
+ "loss": 0.3148,
89
+ "num_tokens": 1659228.0,
90
+ "step": 50
91
+ },
92
+ {
93
+ "epoch": 0.05813953488372093,
94
+ "grad_norm": 0.4712906777858734,
95
+ "learning_rate": 8.014278649308742e-05,
96
+ "loss": 0.3528,
97
+ "num_tokens": 1824722.0,
98
+ "step": 55
99
+ },
100
+ {
101
+ "epoch": 0.06342494714587738,
102
+ "grad_norm": 0.5827367305755615,
103
+ "learning_rate": 8.010437038073538e-05,
104
+ "loss": 0.3307,
105
+ "num_tokens": 1999570.0,
106
+ "step": 60
107
+ },
108
+ {
109
+ "epoch": 0.06871035940803383,
110
+ "grad_norm": 0.46786293387413025,
111
+ "learning_rate": 8.005704281772099e-05,
112
+ "loss": 0.3722,
113
+ "num_tokens": 2149278.0,
114
+ "step": 65
115
+ },
116
+ {
117
+ "epoch": 0.07399577167019028,
118
+ "grad_norm": 0.46910208463668823,
119
+ "learning_rate": 8.000081787444232e-05,
120
+ "loss": 0.3281,
121
+ "num_tokens": 2321362.0,
122
+ "step": 70
123
+ },
124
+ {
125
+ "epoch": 0.07928118393234672,
126
+ "grad_norm": 0.46508923172950745,
127
+ "learning_rate": 7.993571226647224e-05,
128
+ "loss": 0.31,
129
+ "num_tokens": 2513422.0,
130
+ "step": 75
131
+ },
132
+ {
133
+ "epoch": 0.08456659619450317,
134
+ "grad_norm": 0.3090206980705261,
135
+ "learning_rate": 7.98617453495891e-05,
136
+ "loss": 0.3267,
137
+ "num_tokens": 2682030.0,
138
+ "step": 80
139
+ },
140
+ {
141
+ "epoch": 0.08985200845665962,
142
+ "grad_norm": 0.4037857949733734,
143
+ "learning_rate": 7.977893911402208e-05,
144
+ "loss": 0.3497,
145
+ "num_tokens": 2838889.0,
146
+ "step": 85
147
+ },
148
+ {
149
+ "epoch": 0.09513742071881606,
150
+ "grad_norm": 0.3317049443721771,
151
+ "learning_rate": 7.968731817791378e-05,
152
+ "loss": 0.4027,
153
+ "num_tokens": 2984527.0,
154
+ "step": 90
155
+ },
156
+ {
157
+ "epoch": 0.10042283298097252,
158
+ "grad_norm": 0.4948605000972748,
159
+ "learning_rate": 7.958690978000108e-05,
160
+ "loss": 0.3301,
161
+ "num_tokens": 3129850.0,
162
+ "step": 95
163
+ },
164
+ {
165
+ "epoch": 0.10570824524312897,
166
+ "grad_norm": 0.5033330917358398,
167
+ "learning_rate": 7.947774377151723e-05,
168
+ "loss": 0.3574,
169
+ "num_tokens": 3289081.0,
170
+ "step": 100
171
+ },
172
+ {
173
+ "epoch": 0.1109936575052854,
174
+ "grad_norm": 0.6163113117218018,
175
+ "learning_rate": 7.935985260731712e-05,
176
+ "loss": 0.3465,
177
+ "num_tokens": 3456305.0,
178
+ "step": 105
179
+ },
180
+ {
181
+ "epoch": 0.11627906976744186,
182
+ "grad_norm": 0.3303355574607849,
183
+ "learning_rate": 7.923327133622843e-05,
184
+ "loss": 0.3013,
185
+ "num_tokens": 3632845.0,
186
+ "step": 110
187
+ },
188
+ {
189
+ "epoch": 0.12156448202959831,
190
+ "grad_norm": 0.3407858908176422,
191
+ "learning_rate": 7.909803759063184e-05,
192
+ "loss": 0.3059,
193
+ "num_tokens": 3829460.0,
194
+ "step": 115
195
+ },
196
+ {
197
+ "epoch": 0.12684989429175475,
198
+ "grad_norm": 0.39908474683761597,
199
+ "learning_rate": 7.895419157527279e-05,
200
+ "loss": 0.3627,
201
+ "num_tokens": 3992992.0,
202
+ "step": 120
203
+ },
204
+ {
205
+ "epoch": 0.1321353065539112,
206
+ "grad_norm": 0.29309067130088806,
207
+ "learning_rate": 7.880177605530884e-05,
208
+ "loss": 0.2892,
209
+ "num_tokens": 4164780.0,
210
+ "step": 125
211
+ },
212
+ {
213
+ "epoch": 0.13742071881606766,
214
+ "grad_norm": 0.37892845273017883,
215
+ "learning_rate": 7.864083634359562e-05,
216
+ "loss": 0.3028,
217
+ "num_tokens": 4322872.0,
218
+ "step": 130
219
+ },
220
+ {
221
+ "epoch": 0.1427061310782241,
222
+ "grad_norm": 0.34533941745758057,
223
+ "learning_rate": 7.847142028721538e-05,
224
+ "loss": 0.3528,
225
+ "num_tokens": 4484774.0,
226
+ "step": 135
227
+ },
228
+ {
229
+ "epoch": 0.14799154334038056,
230
+ "grad_norm": 0.47042974829673767,
231
+ "learning_rate": 7.829357825325212e-05,
232
+ "loss": 0.3279,
233
+ "num_tokens": 4661609.0,
234
+ "step": 140
235
+ },
236
+ {
237
+ "epoch": 0.15327695560253699,
238
+ "grad_norm": 0.3045842945575714,
239
+ "learning_rate": 7.810736311381762e-05,
240
+ "loss": 0.3445,
241
+ "num_tokens": 4827667.0,
242
+ "step": 145
243
+ },
244
+ {
245
+ "epoch": 0.15856236786469344,
246
+ "grad_norm": 0.4161529541015625,
247
+ "learning_rate": 7.791283023033264e-05,
248
+ "loss": 0.3148,
249
+ "num_tokens": 5000163.0,
250
+ "step": 150
251
+ },
252
+ {
253
+ "epoch": 0.1638477801268499,
254
+ "grad_norm": 0.4866187274456024,
255
+ "learning_rate": 7.771003743706797e-05,
256
+ "loss": 0.3225,
257
+ "num_tokens": 5168441.0,
258
+ "step": 155
259
+ },
260
+ {
261
+ "epoch": 0.16913319238900634,
262
+ "grad_norm": 0.35074010491371155,
263
+ "learning_rate": 7.749904502395058e-05,
264
+ "loss": 0.3089,
265
+ "num_tokens": 5338454.0,
266
+ "step": 160
267
+ },
268
+ {
269
+ "epoch": 0.1744186046511628,
270
+ "grad_norm": 0.34164947271347046,
271
+ "learning_rate": 7.727991571863935e-05,
272
+ "loss": 0.3025,
273
+ "num_tokens": 5512939.0,
274
+ "step": 165
275
+ },
276
+ {
277
+ "epoch": 0.17970401691331925,
278
+ "grad_norm": 0.43914279341697693,
279
+ "learning_rate": 7.705271466787641e-05,
280
+ "loss": 0.3334,
281
+ "num_tokens": 5665993.0,
282
+ "step": 170
283
+ },
284
+ {
285
+ "epoch": 0.1849894291754757,
286
+ "grad_norm": 0.405638724565506,
287
+ "learning_rate": 7.681750941811905e-05,
288
+ "loss": 0.3739,
289
+ "num_tokens": 5824925.0,
290
+ "step": 175
291
+ },
292
+ {
293
+ "epoch": 0.19027484143763213,
294
+ "grad_norm": 0.42802131175994873,
295
+ "learning_rate": 7.657436989545827e-05,
296
+ "loss": 0.3284,
297
+ "num_tokens": 5982641.0,
298
+ "step": 180
299
+ },
300
+ {
301
+ "epoch": 0.19556025369978858,
302
+ "grad_norm": 0.41498205065727234,
303
+ "learning_rate": 7.632336838482996e-05,
304
+ "loss": 0.3022,
305
+ "num_tokens": 6149598.0,
306
+ "step": 185
307
+ },
308
+ {
309
+ "epoch": 0.20084566596194503,
310
+ "grad_norm": 0.40431639552116394,
311
+ "learning_rate": 7.60645795085246e-05,
312
+ "loss": 0.3159,
313
+ "num_tokens": 6320515.0,
314
+ "step": 190
315
+ },
316
+ {
317
+ "epoch": 0.20613107822410148,
318
+ "grad_norm": 0.32176268100738525,
319
+ "learning_rate": 7.579808020400232e-05,
320
+ "loss": 0.2938,
321
+ "num_tokens": 6492543.0,
322
+ "step": 195
323
+ },
324
+ {
325
+ "epoch": 0.21141649048625794,
326
+ "grad_norm": 0.4644106924533844,
327
+ "learning_rate": 7.55239497010194e-05,
328
+ "loss": 0.3349,
329
+ "num_tokens": 6672523.0,
330
+ "step": 200
331
+ },
332
+ {
333
+ "epoch": 0.2167019027484144,
334
+ "grad_norm": 0.35432812571525574,
335
+ "learning_rate": 7.52422694980736e-05,
336
+ "loss": 0.3041,
337
+ "num_tokens": 6844107.0,
338
+ "step": 205
339
+ },
340
+ {
341
+ "epoch": 0.2219873150105708,
342
+ "grad_norm": 0.44315239787101746,
343
+ "learning_rate": 7.495312333817455e-05,
344
+ "loss": 0.2612,
345
+ "num_tokens": 7007468.0,
346
+ "step": 210
347
+ },
348
+ {
349
+ "epoch": 0.22727272727272727,
350
+ "grad_norm": 0.36817148327827454,
351
+ "learning_rate": 7.465659718394734e-05,
352
+ "loss": 0.3221,
353
+ "num_tokens": 7156859.0,
354
+ "step": 215
355
+ },
356
+ {
357
+ "epoch": 0.23255813953488372,
358
+ "grad_norm": 0.39306601881980896,
359
+ "learning_rate": 7.43527791920758e-05,
360
+ "loss": 0.3426,
361
+ "num_tokens": 7311920.0,
362
+ "step": 220
363
+ },
364
+ {
365
+ "epoch": 0.23784355179704017,
366
+ "grad_norm": 0.39284104108810425,
367
+ "learning_rate": 7.404175968709388e-05,
368
+ "loss": 0.3232,
369
+ "num_tokens": 7461910.0,
370
+ "step": 225
371
+ },
372
+ {
373
+ "epoch": 0.24312896405919662,
374
+ "grad_norm": 0.3719322085380554,
375
+ "learning_rate": 7.372363113453213e-05,
376
+ "loss": 0.3277,
377
+ "num_tokens": 7629200.0,
378
+ "step": 230
379
+ },
380
+ {
381
+ "epoch": 0.24841437632135308,
382
+ "grad_norm": 0.3720446527004242,
383
+ "learning_rate": 7.339848811342796e-05,
384
+ "loss": 0.3122,
385
+ "num_tokens": 7790842.0,
386
+ "step": 235
387
+ },
388
+ {
389
+ "epoch": 0.2536997885835095,
390
+ "grad_norm": 0.34994134306907654,
391
+ "learning_rate": 7.306642728820755e-05,
392
+ "loss": 0.3404,
393
+ "num_tokens": 7961093.0,
394
+ "step": 240
395
+ },
396
+ {
397
+ "epoch": 0.25898520084566595,
398
+ "grad_norm": 0.379666268825531,
399
+ "learning_rate": 7.272754737994752e-05,
400
+ "loss": 0.3254,
401
+ "num_tokens": 8118785.0,
402
+ "step": 245
403
+ },
404
+ {
405
+ "epoch": 0.2642706131078224,
406
+ "grad_norm": 0.5034386515617371,
407
+ "learning_rate": 7.238194913702544e-05,
408
+ "loss": 0.3185,
409
+ "num_tokens": 8279200.0,
410
+ "step": 250
411
+ },
412
+ {
413
+ "epoch": 0.26955602536997886,
414
+ "grad_norm": 0.38826555013656616,
415
+ "learning_rate": 7.202973530516749e-05,
416
+ "loss": 0.3021,
417
+ "num_tokens": 8487897.0,
418
+ "step": 255
419
+ },
420
+ {
421
+ "epoch": 0.2748414376321353,
422
+ "grad_norm": 0.3884974420070648,
423
+ "learning_rate": 7.167101059690238e-05,
424
+ "loss": 0.3221,
425
+ "num_tokens": 8636638.0,
426
+ "step": 260
427
+ },
428
+ {
429
+ "epoch": 0.28012684989429176,
430
+ "grad_norm": 0.4962138533592224,
431
+ "learning_rate": 7.130588166043048e-05,
432
+ "loss": 0.3485,
433
+ "num_tokens": 8798198.0,
434
+ "step": 265
435
+ },
436
+ {
437
+ "epoch": 0.2854122621564482,
438
+ "grad_norm": 0.4131626486778259,
439
+ "learning_rate": 7.093445704791747e-05,
440
+ "loss": 0.2897,
441
+ "num_tokens": 8990965.0,
442
+ "step": 270
443
+ },
444
+ {
445
+ "epoch": 0.29069767441860467,
446
+ "grad_norm": 0.38657164573669434,
447
+ "learning_rate": 7.055684718322205e-05,
448
+ "loss": 0.3492,
449
+ "num_tokens": 9152215.0,
450
+ "step": 275
451
+ },
452
+ {
453
+ "epoch": 0.2959830866807611,
454
+ "grad_norm": 0.2804111838340759,
455
+ "learning_rate": 7.017316432906707e-05,
456
+ "loss": 0.3243,
457
+ "num_tokens": 9311537.0,
458
+ "step": 280
459
+ },
460
+ {
461
+ "epoch": 0.3012684989429176,
462
+ "grad_norm": 0.3840746283531189,
463
+ "learning_rate": 6.978352255366406e-05,
464
+ "loss": 0.3657,
465
+ "num_tokens": 9446553.0,
466
+ "step": 285
467
+ },
468
+ {
469
+ "epoch": 0.30655391120507397,
470
+ "grad_norm": 0.29615920782089233,
471
+ "learning_rate": 6.938803769680094e-05,
472
+ "loss": 0.3144,
473
+ "num_tokens": 9625894.0,
474
+ "step": 290
475
+ },
476
+ {
477
+ "epoch": 0.3118393234672304,
478
+ "grad_norm": 0.3887929320335388,
479
+ "learning_rate": 6.898682733540313e-05,
480
+ "loss": 0.2967,
481
+ "num_tokens": 9820246.0,
482
+ "step": 295
483
+ },
484
+ {
485
+ "epoch": 0.3171247357293869,
486
+ "grad_norm": 0.4418380856513977,
487
+ "learning_rate": 6.8580010748578e-05,
488
+ "loss": 0.3299,
489
+ "num_tokens": 9995021.0,
490
+ "step": 300
491
+ },
492
+ {
493
+ "epoch": 0.3224101479915433,
494
+ "grad_norm": 0.327908992767334,
495
+ "learning_rate": 6.816770888215352e-05,
496
+ "loss": 0.2879,
497
+ "num_tokens": 10186048.0,
498
+ "step": 305
499
+ },
500
+ {
501
+ "epoch": 0.3276955602536998,
502
+ "grad_norm": 0.31937530636787415,
503
+ "learning_rate": 6.775004431272132e-05,
504
+ "loss": 0.3504,
505
+ "num_tokens": 10343442.0,
506
+ "step": 310
507
+ },
508
+ {
509
+ "epoch": 0.33298097251585623,
510
+ "grad_norm": 0.3883694112300873,
511
+ "learning_rate": 6.732714121119478e-05,
512
+ "loss": 0.3188,
513
+ "num_tokens": 10505725.0,
514
+ "step": 315
515
+ },
516
+ {
517
+ "epoch": 0.3382663847780127,
518
+ "grad_norm": 0.520715594291687,
519
+ "learning_rate": 6.68991253058933e-05,
520
+ "loss": 0.3348,
521
+ "num_tokens": 10685454.0,
522
+ "step": 320
523
+ },
524
+ {
525
+ "epoch": 0.34355179704016914,
526
+ "grad_norm": 0.3867979943752289,
527
+ "learning_rate": 6.646612384516355e-05,
528
+ "loss": 0.2958,
529
+ "num_tokens": 10839868.0,
530
+ "step": 325
531
+ },
532
+ {
533
+ "epoch": 0.3488372093023256,
534
+ "grad_norm": 0.4362066686153412,
535
+ "learning_rate": 6.602826555954866e-05,
536
+ "loss": 0.3052,
537
+ "num_tokens": 10987360.0,
538
+ "step": 330
539
+ },
540
+ {
541
+ "epoch": 0.35412262156448204,
542
+ "grad_norm": 0.3443554639816284,
543
+ "learning_rate": 6.558568062351694e-05,
544
+ "loss": 0.3133,
545
+ "num_tokens": 11147325.0,
546
+ "step": 335
547
+ },
548
+ {
549
+ "epoch": 0.3594080338266385,
550
+ "grad_norm": 0.3375532329082489,
551
+ "learning_rate": 6.513850061676129e-05,
552
+ "loss": 0.3187,
553
+ "num_tokens": 11328945.0,
554
+ "step": 340
555
+ },
556
+ {
557
+ "epoch": 0.36469344608879495,
558
+ "grad_norm": 0.48123109340667725,
559
+ "learning_rate": 6.468685848508066e-05,
560
+ "loss": 0.3234,
561
+ "num_tokens": 11491151.0,
562
+ "step": 345
563
+ },
564
+ {
565
+ "epoch": 0.3699788583509514,
566
+ "grad_norm": 0.350399911403656,
567
+ "learning_rate": 6.423088850085563e-05,
568
+ "loss": 0.2754,
569
+ "num_tokens": 11659065.0,
570
+ "step": 350
571
+ },
572
+ {
573
+ "epoch": 0.3752642706131078,
574
+ "grad_norm": 0.3584173917770386,
575
+ "learning_rate": 6.377072622312942e-05,
576
+ "loss": 0.3336,
577
+ "num_tokens": 11863689.0,
578
+ "step": 355
579
+ },
580
+ {
581
+ "epoch": 0.38054968287526425,
582
+ "grad_norm": 0.3181978464126587,
583
+ "learning_rate": 6.330650845730648e-05,
584
+ "loss": 0.3418,
585
+ "num_tokens": 12006544.0,
586
+ "step": 360
587
+ },
588
+ {
589
+ "epoch": 0.3858350951374207,
590
+ "grad_norm": 0.30926552414894104,
591
+ "learning_rate": 6.283837321448044e-05,
592
+ "loss": 0.3248,
593
+ "num_tokens": 12155215.0,
594
+ "step": 365
595
+ },
596
+ {
597
+ "epoch": 0.39112050739957716,
598
+ "grad_norm": 0.30399248003959656,
599
+ "learning_rate": 6.236645967040363e-05,
600
+ "loss": 0.3199,
601
+ "num_tokens": 12304714.0,
602
+ "step": 370
603
+ },
604
+ {
605
+ "epoch": 0.3964059196617336,
606
+ "grad_norm": 0.3858378827571869,
607
+ "learning_rate": 6.189090812411056e-05,
608
+ "loss": 0.2905,
609
+ "num_tokens": 12467243.0,
610
+ "step": 375
611
+ },
612
+ {
613
+ "epoch": 0.40169133192389006,
614
+ "grad_norm": 0.3069511950016022,
615
+ "learning_rate": 6.141185995620703e-05,
616
+ "loss": 0.3314,
617
+ "num_tokens": 12641866.0,
618
+ "step": 380
619
+ },
620
+ {
621
+ "epoch": 0.4069767441860465,
622
+ "grad_norm": 0.31766510009765625,
623
+ "learning_rate": 6.0929457586838143e-05,
624
+ "loss": 0.3084,
625
+ "num_tokens": 12805214.0,
626
+ "step": 385
627
+ },
628
+ {
629
+ "epoch": 0.41226215644820297,
630
+ "grad_norm": 0.5858482122421265,
631
+ "learning_rate": 6.044384443334691e-05,
632
+ "loss": 0.2878,
633
+ "num_tokens": 12967228.0,
634
+ "step": 390
635
+ },
636
+ {
637
+ "epoch": 0.4175475687103594,
638
+ "grad_norm": 0.34514257311820984,
639
+ "learning_rate": 5.9955164867636644e-05,
640
+ "loss": 0.3471,
641
+ "num_tokens": 13104829.0,
642
+ "step": 395
643
+ },
644
+ {
645
+ "epoch": 0.42283298097251587,
646
+ "grad_norm": 0.4240577816963196,
647
+ "learning_rate": 5.9463564173249374e-05,
648
+ "loss": 0.2973,
649
+ "num_tokens": 13288474.0,
650
+ "step": 400
651
+ },
652
+ {
653
+ "epoch": 0.4281183932346723,
654
+ "grad_norm": 0.3454580008983612,
655
+ "learning_rate": 5.896918850217336e-05,
656
+ "loss": 0.3241,
657
+ "num_tokens": 13451876.0,
658
+ "step": 405
659
+ },
660
+ {
661
+ "epoch": 0.4334038054968288,
662
+ "grad_norm": 0.3258069157600403,
663
+ "learning_rate": 5.8472184831392364e-05,
664
+ "loss": 0.2968,
665
+ "num_tokens": 13633506.0,
666
+ "step": 410
667
+ },
668
+ {
669
+ "epoch": 0.43868921775898523,
670
+ "grad_norm": 0.35886386036872864,
671
+ "learning_rate": 5.797270091918965e-05,
672
+ "loss": 0.3173,
673
+ "num_tokens": 13794654.0,
674
+ "step": 415
675
+ },
676
+ {
677
+ "epoch": 0.4439746300211416,
678
+ "grad_norm": 0.5046164393424988,
679
+ "learning_rate": 5.747088526121975e-05,
680
+ "loss": 0.3439,
681
+ "num_tokens": 13928763.0,
682
+ "step": 420
683
+ },
684
+ {
685
+ "epoch": 0.4492600422832981,
686
+ "grad_norm": 0.34629306197166443,
687
+ "learning_rate": 5.696688704636091e-05,
688
+ "loss": 0.3095,
689
+ "num_tokens": 14091776.0,
690
+ "step": 425
691
+ },
692
+ {
693
+ "epoch": 0.45454545454545453,
694
+ "grad_norm": 0.35323911905288696,
695
+ "learning_rate": 5.6460856112361576e-05,
696
+ "loss": 0.3053,
697
+ "num_tokens": 14283410.0,
698
+ "step": 430
699
+ },
700
+ {
701
+ "epoch": 0.459830866807611,
702
+ "grad_norm": 0.3977009356021881,
703
+ "learning_rate": 5.595294290129389e-05,
704
+ "loss": 0.2885,
705
+ "num_tokens": 14450115.0,
706
+ "step": 435
707
+ },
708
+ {
709
+ "epoch": 0.46511627906976744,
710
+ "grad_norm": 0.34842291474342346,
711
+ "learning_rate": 5.5443298414827514e-05,
712
+ "loss": 0.2904,
713
+ "num_tokens": 14629092.0,
714
+ "step": 440
715
+ },
716
+ {
717
+ "epoch": 0.4704016913319239,
718
+ "grad_norm": 0.3181321322917938,
719
+ "learning_rate": 5.4932074169337124e-05,
720
+ "loss": 0.3289,
721
+ "num_tokens": 14766334.0,
722
+ "step": 445
723
+ }
724
+ ],
725
+ "logging_steps": 5,
726
+ "max_steps": 946,
727
+ "num_input_tokens_seen": 0,
728
+ "num_train_epochs": 1,
729
+ "save_steps": 500,
730
+ "stateful_callbacks": {
731
+ "TrainerControl": {
732
+ "args": {
733
+ "should_epoch_stop": false,
734
+ "should_evaluate": false,
735
+ "should_log": false,
736
+ "should_save": true,
737
+ "should_training_stop": true
738
+ },
739
+ "attributes": {}
740
+ }
741
+ },
742
+ "total_flos": 1.8225814430631854e+18,
743
+ "train_batch_size": 14,
744
+ "trial_name": null,
745
+ "trial_params": null
746
+ }
training_args.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0a21309dd07b6449b2b5b552ad09b8dfd16bc1c50fad4a087fce233aa5408f3b
3
+ size 5816
vocab.json DELETED
The diff for this file is too large to render. See raw diff