ValasaiChander commited on
Commit
8ab646e
·
verified ·
1 Parent(s): 719d9e5

Delete models/stance_tracker/checkpoint-5255

Browse files
models/stance_tracker/checkpoint-5255/README.md DELETED
@@ -1,202 +0,0 @@
1
- ---
2
- base_model: cross-encoder/nli-deberta-v3-small
3
- library_name: peft
4
- ---
5
-
6
- # Model Card for Model ID
7
-
8
- <!-- Provide a quick summary of what the model is/does. -->
9
-
10
-
11
-
12
- ## Model Details
13
-
14
- ### Model Description
15
-
16
- <!-- Provide a longer summary of what this model is. -->
17
-
18
-
19
-
20
- - **Developed by:** [More Information Needed]
21
- - **Funded by [optional]:** [More Information Needed]
22
- - **Shared by [optional]:** [More Information Needed]
23
- - **Model type:** [More Information Needed]
24
- - **Language(s) (NLP):** [More Information Needed]
25
- - **License:** [More Information Needed]
26
- - **Finetuned from model [optional]:** [More Information Needed]
27
-
28
- ### Model Sources [optional]
29
-
30
- <!-- Provide the basic links for the model. -->
31
-
32
- - **Repository:** [More Information Needed]
33
- - **Paper [optional]:** [More Information Needed]
34
- - **Demo [optional]:** [More Information Needed]
35
-
36
- ## Uses
37
-
38
- <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
39
-
40
- ### Direct Use
41
-
42
- <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
43
-
44
- [More Information Needed]
45
-
46
- ### Downstream Use [optional]
47
-
48
- <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
49
-
50
- [More Information Needed]
51
-
52
- ### Out-of-Scope Use
53
-
54
- <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
55
-
56
- [More Information Needed]
57
-
58
- ## Bias, Risks, and Limitations
59
-
60
- <!-- This section is meant to convey both technical and sociotechnical limitations. -->
61
-
62
- [More Information Needed]
63
-
64
- ### Recommendations
65
-
66
- <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
67
-
68
- Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
69
-
70
- ## How to Get Started with the Model
71
-
72
- Use the code below to get started with the model.
73
-
74
- [More Information Needed]
75
-
76
- ## Training Details
77
-
78
- ### Training Data
79
-
80
- <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
81
-
82
- [More Information Needed]
83
-
84
- ### Training Procedure
85
-
86
- <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
87
-
88
- #### Preprocessing [optional]
89
-
90
- [More Information Needed]
91
-
92
-
93
- #### Training Hyperparameters
94
-
95
- - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
96
-
97
- #### Speeds, Sizes, Times [optional]
98
-
99
- <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
100
-
101
- [More Information Needed]
102
-
103
- ## Evaluation
104
-
105
- <!-- This section describes the evaluation protocols and provides the results. -->
106
-
107
- ### Testing Data, Factors & Metrics
108
-
109
- #### Testing Data
110
-
111
- <!-- This should link to a Dataset Card if possible. -->
112
-
113
- [More Information Needed]
114
-
115
- #### Factors
116
-
117
- <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
118
-
119
- [More Information Needed]
120
-
121
- #### Metrics
122
-
123
- <!-- These are the evaluation metrics being used, ideally with a description of why. -->
124
-
125
- [More Information Needed]
126
-
127
- ### Results
128
-
129
- [More Information Needed]
130
-
131
- #### Summary
132
-
133
-
134
-
135
- ## Model Examination [optional]
136
-
137
- <!-- Relevant interpretability work for the model goes here -->
138
-
139
- [More Information Needed]
140
-
141
- ## Environmental Impact
142
-
143
- <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
144
-
145
- Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
146
-
147
- - **Hardware Type:** [More Information Needed]
148
- - **Hours used:** [More Information Needed]
149
- - **Cloud Provider:** [More Information Needed]
150
- - **Compute Region:** [More Information Needed]
151
- - **Carbon Emitted:** [More Information Needed]
152
-
153
- ## Technical Specifications [optional]
154
-
155
- ### Model Architecture and Objective
156
-
157
- [More Information Needed]
158
-
159
- ### Compute Infrastructure
160
-
161
- [More Information Needed]
162
-
163
- #### Hardware
164
-
165
- [More Information Needed]
166
-
167
- #### Software
168
-
169
- [More Information Needed]
170
-
171
- ## Citation [optional]
172
-
173
- <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
174
-
175
- **BibTeX:**
176
-
177
- [More Information Needed]
178
-
179
- **APA:**
180
-
181
- [More Information Needed]
182
-
183
- ## Glossary [optional]
184
-
185
- <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
186
-
187
- [More Information Needed]
188
-
189
- ## More Information [optional]
190
-
191
- [More Information Needed]
192
-
193
- ## Model Card Authors [optional]
194
-
195
- [More Information Needed]
196
-
197
- ## Model Card Contact
198
-
199
- [More Information Needed]
200
- ### Framework versions
201
-
202
- - PEFT 0.12.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
models/stance_tracker/checkpoint-5255/adapter_config.json DELETED
@@ -1,33 +0,0 @@
1
- {
2
- "alpha_pattern": {},
3
- "auto_mapping": null,
4
- "base_model_name_or_path": "cross-encoder/nli-deberta-v3-small",
5
- "bias": "none",
6
- "fan_in_fan_out": false,
7
- "inference_mode": true,
8
- "init_lora_weights": true,
9
- "layer_replication": null,
10
- "layers_pattern": null,
11
- "layers_to_transform": null,
12
- "loftq_config": {},
13
- "lora_alpha": 32,
14
- "lora_dropout": 0.1,
15
- "megatron_config": null,
16
- "megatron_core": "megatron.core",
17
- "modules_to_save": [
18
- "classifier",
19
- "score"
20
- ],
21
- "peft_type": "LORA",
22
- "r": 16,
23
- "rank_pattern": {},
24
- "revision": null,
25
- "target_modules": [
26
- "query_proj",
27
- "key_proj",
28
- "value_proj"
29
- ],
30
- "task_type": "SEQ_CLS",
31
- "use_dora": false,
32
- "use_rslora": false
33
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
models/stance_tracker/checkpoint-5255/adapter_model.safetensors DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:910c12283603a9508e3ef6eff567ca28443ea71c49466b78efbb572eee1af17d
3
- size 1784220
 
 
 
 
models/stance_tracker/checkpoint-5255/added_tokens.json DELETED
@@ -1,3 +0,0 @@
1
- {
2
- "[MASK]": 128000
3
- }
 
 
 
 
models/stance_tracker/checkpoint-5255/optimizer.pt DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:2493364078eb423afbdbbe0a97b86d0c4a13d3e2421a882b75d4a5a47398f6a5
3
- size 3590091
 
 
 
 
models/stance_tracker/checkpoint-5255/rng_state.pth DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:2d16468ea7d34ff50d4ff8edabcee5e9986716b2a4228da4d040cec6e2b202bc
3
- size 14645
 
 
 
 
models/stance_tracker/checkpoint-5255/scheduler.pt DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:aa68dd0ca704e46527cf15047d30da476cb73df6764cb5f5dfc2226ea090b38f
3
- size 1465
 
 
 
 
models/stance_tracker/checkpoint-5255/special_tokens_map.json DELETED
@@ -1,51 +0,0 @@
1
- {
2
- "bos_token": {
3
- "content": "[CLS]",
4
- "lstrip": false,
5
- "normalized": false,
6
- "rstrip": false,
7
- "single_word": false
8
- },
9
- "cls_token": {
10
- "content": "[CLS]",
11
- "lstrip": false,
12
- "normalized": false,
13
- "rstrip": false,
14
- "single_word": false
15
- },
16
- "eos_token": {
17
- "content": "[SEP]",
18
- "lstrip": false,
19
- "normalized": false,
20
- "rstrip": false,
21
- "single_word": false
22
- },
23
- "mask_token": {
24
- "content": "[MASK]",
25
- "lstrip": false,
26
- "normalized": false,
27
- "rstrip": false,
28
- "single_word": false
29
- },
30
- "pad_token": {
31
- "content": "[PAD]",
32
- "lstrip": false,
33
- "normalized": false,
34
- "rstrip": false,
35
- "single_word": false
36
- },
37
- "sep_token": {
38
- "content": "[SEP]",
39
- "lstrip": false,
40
- "normalized": false,
41
- "rstrip": false,
42
- "single_word": false
43
- },
44
- "unk_token": {
45
- "content": "[UNK]",
46
- "lstrip": false,
47
- "normalized": true,
48
- "rstrip": false,
49
- "single_word": false
50
- }
51
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
models/stance_tracker/checkpoint-5255/spm.model DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:c679fbf93643d19aab7ee10c0b99e460bdbc02fedf34b92b05af343b4af586fd
3
- size 2464616
 
 
 
 
models/stance_tracker/checkpoint-5255/tokenizer.json DELETED
The diff for this file is too large to render. See raw diff
 
models/stance_tracker/checkpoint-5255/tokenizer_config.json DELETED
@@ -1,59 +0,0 @@
1
- {
2
- "added_tokens_decoder": {
3
- "0": {
4
- "content": "[PAD]",
5
- "lstrip": false,
6
- "normalized": false,
7
- "rstrip": false,
8
- "single_word": false,
9
- "special": true
10
- },
11
- "1": {
12
- "content": "[CLS]",
13
- "lstrip": false,
14
- "normalized": false,
15
- "rstrip": false,
16
- "single_word": false,
17
- "special": true
18
- },
19
- "2": {
20
- "content": "[SEP]",
21
- "lstrip": false,
22
- "normalized": false,
23
- "rstrip": false,
24
- "single_word": false,
25
- "special": true
26
- },
27
- "3": {
28
- "content": "[UNK]",
29
- "lstrip": false,
30
- "normalized": true,
31
- "rstrip": false,
32
- "single_word": false,
33
- "special": true
34
- },
35
- "128000": {
36
- "content": "[MASK]",
37
- "lstrip": false,
38
- "normalized": false,
39
- "rstrip": false,
40
- "single_word": false,
41
- "special": true
42
- }
43
- },
44
- "bos_token": "[CLS]",
45
- "clean_up_tokenization_spaces": false,
46
- "cls_token": "[CLS]",
47
- "do_lower_case": false,
48
- "eos_token": "[SEP]",
49
- "extra_special_tokens": {},
50
- "mask_token": "[MASK]",
51
- "model_max_length": 512,
52
- "pad_token": "[PAD]",
53
- "sep_token": "[SEP]",
54
- "sp_model_kwargs": {},
55
- "split_by_punct": false,
56
- "tokenizer_class": "DebertaV2Tokenizer",
57
- "unk_token": "[UNK]",
58
- "vocab_type": "spm"
59
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
models/stance_tracker/checkpoint-5255/trainer_state.json DELETED
@@ -1,456 +0,0 @@
1
- {
2
- "best_metric": 0.8411684097031022,
3
- "best_model_checkpoint": "models/stance_tracker/checkpoint-5255",
4
- "epoch": 5.0,
5
- "eval_steps": 500,
6
- "global_step": 5255,
7
- "is_hyper_param_search": false,
8
- "is_local_process_zero": true,
9
- "is_world_process_zero": true,
10
- "log_history": [
11
- {
12
- "epoch": 0.09514747859181731,
13
- "grad_norm": 1180798.875,
14
- "learning_rate": 1.9011406844106464e-05,
15
- "loss": 4.8322,
16
- "step": 100
17
- },
18
- {
19
- "epoch": 0.19029495718363462,
20
- "grad_norm": 73790.9375,
21
- "learning_rate": 3.802281368821293e-05,
22
- "loss": 1.802,
23
- "step": 200
24
- },
25
- {
26
- "epoch": 0.285442435775452,
27
- "grad_norm": 106478.0546875,
28
- "learning_rate": 5.703422053231939e-05,
29
- "loss": 0.9777,
30
- "step": 300
31
- },
32
- {
33
- "epoch": 0.38058991436726924,
34
- "grad_norm": 246882.859375,
35
- "learning_rate": 7.604562737642586e-05,
36
- "loss": 0.8782,
37
- "step": 400
38
- },
39
- {
40
- "epoch": 0.47573739295908657,
41
- "grad_norm": 193025.796875,
42
- "learning_rate": 9.505703422053233e-05,
43
- "loss": 0.7562,
44
- "step": 500
45
- },
46
- {
47
- "epoch": 0.570884871550904,
48
- "grad_norm": 136630.859375,
49
- "learning_rate": 9.843518714315923e-05,
50
- "loss": 0.603,
51
- "step": 600
52
- },
53
- {
54
- "epoch": 0.6660323501427212,
55
- "grad_norm": 175707.984375,
56
- "learning_rate": 9.63205751744555e-05,
57
- "loss": 0.5252,
58
- "step": 700
59
- },
60
- {
61
- "epoch": 0.7611798287345385,
62
- "grad_norm": 199550.734375,
63
- "learning_rate": 9.420596320575174e-05,
64
- "loss": 0.4753,
65
- "step": 800
66
- },
67
- {
68
- "epoch": 0.8563273073263559,
69
- "grad_norm": 147047.765625,
70
- "learning_rate": 9.2091351237048e-05,
71
- "loss": 0.4839,
72
- "step": 900
73
- },
74
- {
75
- "epoch": 0.9514747859181731,
76
- "grad_norm": 170235.53125,
77
- "learning_rate": 8.997673926834426e-05,
78
- "loss": 0.4608,
79
- "step": 1000
80
- },
81
- {
82
- "epoch": 1.0,
83
- "eval_accuracy": 0.8310730430644777,
84
- "eval_f1_macro": 0.8119798637764705,
85
- "eval_loss": 0.397123783826828,
86
- "eval_runtime": 29.8024,
87
- "eval_samples_per_second": 141.029,
88
- "eval_steps_per_second": 4.429,
89
- "step": 1051
90
- },
91
- {
92
- "epoch": 1.0466222645099905,
93
- "grad_norm": 154165.984375,
94
- "learning_rate": 8.786212729964052e-05,
95
- "loss": 0.4329,
96
- "step": 1100
97
- },
98
- {
99
- "epoch": 1.141769743101808,
100
- "grad_norm": 167387.453125,
101
- "learning_rate": 8.574751533093678e-05,
102
- "loss": 0.4237,
103
- "step": 1200
104
- },
105
- {
106
- "epoch": 1.236917221693625,
107
- "grad_norm": 78232.890625,
108
- "learning_rate": 8.363290336223304e-05,
109
- "loss": 0.4274,
110
- "step": 1300
111
- },
112
- {
113
- "epoch": 1.3320647002854424,
114
- "grad_norm": 157848.515625,
115
- "learning_rate": 8.151829139352929e-05,
116
- "loss": 0.4101,
117
- "step": 1400
118
- },
119
- {
120
- "epoch": 1.4272121788772598,
121
- "grad_norm": 83062.8359375,
122
- "learning_rate": 7.940367942482555e-05,
123
- "loss": 0.4046,
124
- "step": 1500
125
- },
126
- {
127
- "epoch": 1.5223596574690772,
128
- "grad_norm": 217329.375,
129
- "learning_rate": 7.728906745612181e-05,
130
- "loss": 0.405,
131
- "step": 1600
132
- },
133
- {
134
- "epoch": 1.6175071360608944,
135
- "grad_norm": 219439.40625,
136
- "learning_rate": 7.517445548741806e-05,
137
- "loss": 0.3994,
138
- "step": 1700
139
- },
140
- {
141
- "epoch": 1.7126546146527117,
142
- "grad_norm": 94135.8046875,
143
- "learning_rate": 7.305984351871432e-05,
144
- "loss": 0.3957,
145
- "step": 1800
146
- },
147
- {
148
- "epoch": 1.8078020932445291,
149
- "grad_norm": 267690.875,
150
- "learning_rate": 7.094523155001058e-05,
151
- "loss": 0.4029,
152
- "step": 1900
153
- },
154
- {
155
- "epoch": 1.9029495718363463,
156
- "grad_norm": 289756.53125,
157
- "learning_rate": 6.883061958130684e-05,
158
- "loss": 0.3973,
159
- "step": 2000
160
- },
161
- {
162
- "epoch": 1.9980970504281637,
163
- "grad_norm": 230541.8125,
164
- "learning_rate": 6.67160076126031e-05,
165
- "loss": 0.3772,
166
- "step": 2100
167
- },
168
- {
169
- "epoch": 2.0,
170
- "eval_accuracy": 0.8505829169640733,
171
- "eval_f1_macro": 0.8312500853233197,
172
- "eval_loss": 0.3541408181190491,
173
- "eval_runtime": 30.0593,
174
- "eval_samples_per_second": 139.824,
175
- "eval_steps_per_second": 4.391,
176
- "step": 2102
177
- },
178
- {
179
- "epoch": 2.093244529019981,
180
- "grad_norm": 163747.640625,
181
- "learning_rate": 6.460139564389934e-05,
182
- "loss": 0.3638,
183
- "step": 2200
184
- },
185
- {
186
- "epoch": 2.188392007611798,
187
- "grad_norm": 269249.65625,
188
- "learning_rate": 6.24867836751956e-05,
189
- "loss": 0.367,
190
- "step": 2300
191
- },
192
- {
193
- "epoch": 2.283539486203616,
194
- "grad_norm": 131646.234375,
195
- "learning_rate": 6.0372171706491863e-05,
196
- "loss": 0.3669,
197
- "step": 2400
198
- },
199
- {
200
- "epoch": 2.378686964795433,
201
- "grad_norm": 163041.328125,
202
- "learning_rate": 5.825755973778812e-05,
203
- "loss": 0.3556,
204
- "step": 2500
205
- },
206
- {
207
- "epoch": 2.47383444338725,
208
- "grad_norm": 183448.84375,
209
- "learning_rate": 5.614294776908438e-05,
210
- "loss": 0.3565,
211
- "step": 2600
212
- },
213
- {
214
- "epoch": 2.5689819219790677,
215
- "grad_norm": 223062.75,
216
- "learning_rate": 5.402833580038064e-05,
217
- "loss": 0.3728,
218
- "step": 2700
219
- },
220
- {
221
- "epoch": 2.664129400570885,
222
- "grad_norm": 125715.6640625,
223
- "learning_rate": 5.1913723831676884e-05,
224
- "loss": 0.3654,
225
- "step": 2800
226
- },
227
- {
228
- "epoch": 2.759276879162702,
229
- "grad_norm": 179914.296875,
230
- "learning_rate": 4.9799111862973144e-05,
231
- "loss": 0.3571,
232
- "step": 2900
233
- },
234
- {
235
- "epoch": 2.8544243577545196,
236
- "grad_norm": 184159.671875,
237
- "learning_rate": 4.7684499894269404e-05,
238
- "loss": 0.3499,
239
- "step": 3000
240
- },
241
- {
242
- "epoch": 2.949571836346337,
243
- "grad_norm": 123080.375,
244
- "learning_rate": 4.5569887925565665e-05,
245
- "loss": 0.3556,
246
- "step": 3100
247
- },
248
- {
249
- "epoch": 3.0,
250
- "eval_accuracy": 0.8624791815369974,
251
- "eval_f1_macro": 0.8367348377738546,
252
- "eval_loss": 0.32585257291793823,
253
- "eval_runtime": 29.8922,
254
- "eval_samples_per_second": 140.605,
255
- "eval_steps_per_second": 4.416,
256
- "step": 3153
257
- },
258
- {
259
- "epoch": 3.044719314938154,
260
- "grad_norm": 232850.921875,
261
- "learning_rate": 4.345527595686192e-05,
262
- "loss": 0.3615,
263
- "step": 3200
264
- },
265
- {
266
- "epoch": 3.1398667935299716,
267
- "grad_norm": 123978.3515625,
268
- "learning_rate": 4.134066398815817e-05,
269
- "loss": 0.3419,
270
- "step": 3300
271
- },
272
- {
273
- "epoch": 3.2350142721217887,
274
- "grad_norm": 143989.796875,
275
- "learning_rate": 3.922605201945443e-05,
276
- "loss": 0.3339,
277
- "step": 3400
278
- },
279
- {
280
- "epoch": 3.3301617507136063,
281
- "grad_norm": 131409.40625,
282
- "learning_rate": 3.711144005075069e-05,
283
- "loss": 0.3272,
284
- "step": 3500
285
- },
286
- {
287
- "epoch": 3.4253092293054235,
288
- "grad_norm": 295399.4375,
289
- "learning_rate": 3.4996828082046945e-05,
290
- "loss": 0.3529,
291
- "step": 3600
292
- },
293
- {
294
- "epoch": 3.5204567078972406,
295
- "grad_norm": 279721.375,
296
- "learning_rate": 3.28822161133432e-05,
297
- "loss": 0.3484,
298
- "step": 3700
299
- },
300
- {
301
- "epoch": 3.6156041864890582,
302
- "grad_norm": 234246.640625,
303
- "learning_rate": 3.076760414463946e-05,
304
- "loss": 0.3274,
305
- "step": 3800
306
- },
307
- {
308
- "epoch": 3.7107516650808754,
309
- "grad_norm": 215512.46875,
310
- "learning_rate": 2.8652992175935716e-05,
311
- "loss": 0.3289,
312
- "step": 3900
313
- },
314
- {
315
- "epoch": 3.8058991436726926,
316
- "grad_norm": 193919.359375,
317
- "learning_rate": 2.6538380207231972e-05,
318
- "loss": 0.3427,
319
- "step": 4000
320
- },
321
- {
322
- "epoch": 3.90104662226451,
323
- "grad_norm": 266348.53125,
324
- "learning_rate": 2.442376823852823e-05,
325
- "loss": 0.3488,
326
- "step": 4100
327
- },
328
- {
329
- "epoch": 3.9961941008563273,
330
- "grad_norm": 157798.1875,
331
- "learning_rate": 2.230915626982449e-05,
332
- "loss": 0.327,
333
- "step": 4200
334
- },
335
- {
336
- "epoch": 4.0,
337
- "eval_accuracy": 0.8639067332857483,
338
- "eval_f1_macro": 0.8404976856282999,
339
- "eval_loss": 0.3265078663825989,
340
- "eval_runtime": 29.6057,
341
- "eval_samples_per_second": 141.966,
342
- "eval_steps_per_second": 4.459,
343
- "step": 4204
344
- },
345
- {
346
- "epoch": 4.0913415794481445,
347
- "grad_norm": 93288.796875,
348
- "learning_rate": 2.0194544301120746e-05,
349
- "loss": 0.3462,
350
- "step": 4300
351
- },
352
- {
353
- "epoch": 4.186489058039962,
354
- "grad_norm": 155794.875,
355
- "learning_rate": 1.8079932332417003e-05,
356
- "loss": 0.3407,
357
- "step": 4400
358
- },
359
- {
360
- "epoch": 4.28163653663178,
361
- "grad_norm": 180881.90625,
362
- "learning_rate": 1.596532036371326e-05,
363
- "loss": 0.3082,
364
- "step": 4500
365
- },
366
- {
367
- "epoch": 4.376784015223596,
368
- "grad_norm": 150323.046875,
369
- "learning_rate": 1.3850708395009515e-05,
370
- "loss": 0.315,
371
- "step": 4600
372
- },
373
- {
374
- "epoch": 4.471931493815414,
375
- "grad_norm": 123327.1796875,
376
- "learning_rate": 1.1736096426305774e-05,
377
- "loss": 0.3391,
378
- "step": 4700
379
- },
380
- {
381
- "epoch": 4.567078972407232,
382
- "grad_norm": 160687.984375,
383
- "learning_rate": 9.62148445760203e-06,
384
- "loss": 0.3163,
385
- "step": 4800
386
- },
387
- {
388
- "epoch": 4.662226450999048,
389
- "grad_norm": 59321.7109375,
390
- "learning_rate": 7.506872488898288e-06,
391
- "loss": 0.3198,
392
- "step": 4900
393
- },
394
- {
395
- "epoch": 4.757373929590866,
396
- "grad_norm": 117285.140625,
397
- "learning_rate": 5.392260520194545e-06,
398
- "loss": 0.3315,
399
- "step": 5000
400
- },
401
- {
402
- "epoch": 4.8525214081826835,
403
- "grad_norm": 153029.640625,
404
- "learning_rate": 3.2776485514908016e-06,
405
- "loss": 0.3144,
406
- "step": 5100
407
- },
408
- {
409
- "epoch": 4.9476688867745,
410
- "grad_norm": 82521.359375,
411
- "learning_rate": 1.1630365827870586e-06,
412
- "loss": 0.3221,
413
- "step": 5200
414
- },
415
- {
416
- "epoch": 5.0,
417
- "eval_accuracy": 0.8655722103259577,
418
- "eval_f1_macro": 0.8411684097031022,
419
- "eval_loss": 0.319923996925354,
420
- "eval_runtime": 29.7497,
421
- "eval_samples_per_second": 141.279,
422
- "eval_steps_per_second": 4.437,
423
- "step": 5255
424
- }
425
- ],
426
- "logging_steps": 100,
427
- "max_steps": 5255,
428
- "num_input_tokens_seen": 0,
429
- "num_train_epochs": 5,
430
- "save_steps": 500,
431
- "stateful_callbacks": {
432
- "EarlyStoppingCallback": {
433
- "args": {
434
- "early_stopping_patience": 2,
435
- "early_stopping_threshold": 0.0
436
- },
437
- "attributes": {
438
- "early_stopping_patience_counter": 0
439
- }
440
- },
441
- "TrainerControl": {
442
- "args": {
443
- "should_epoch_stop": false,
444
- "should_evaluate": false,
445
- "should_log": false,
446
- "should_save": true,
447
- "should_training_stop": true
448
- },
449
- "attributes": {}
450
- }
451
- },
452
- "total_flos": 1.125063421341696e+16,
453
- "train_batch_size": 32,
454
- "trial_name": null,
455
- "trial_params": null
456
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
models/stance_tracker/checkpoint-5255/training_args.bin DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:b9b570979fe3321d4de2454f36475c8d35e6fd68b5100ed47a7dc03d542ad2b0
3
- size 5585