K2triinK commited on
Commit
df361c1
·
verified ·
1 Parent(s): afad0a6

Add files using upload-large-folder tool

Browse files
This view is limited to 50 files because it contains too many changes.   See raw diff
Files changed (50) hide show
  1. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test1/README.md +58 -0
  2. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/README.md +58 -0
  3. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/README.md +209 -0
  4. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/adapter_config.json +48 -0
  5. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/chat_template.jinja +85 -0
  6. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/tokenizer_config.json +29 -0
  7. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/trainer_state.json +297 -0
  8. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/README.md +209 -0
  9. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/adapter_config.json +48 -0
  10. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/chat_template.jinja +85 -0
  11. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/tokenizer_config.json +29 -0
  12. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/trainer_state.json +388 -0
  13. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/README.md +209 -0
  14. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/adapter_config.json +48 -0
  15. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/chat_template.jinja +85 -0
  16. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/tokenizer_config.json +29 -0
  17. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/trainer_state.json +469 -0
  18. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/README.md +209 -0
  19. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/adapter_config.json +48 -0
  20. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/chat_template.jinja +85 -0
  21. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/tokenizer_config.json +29 -0
  22. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/trainer_state.json +560 -0
  23. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/README.md +209 -0
  24. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/adapter_config.json +48 -0
  25. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/chat_template.jinja +85 -0
  26. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/tokenizer_config.json +29 -0
  27. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/trainer_state.json +651 -0
  28. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/README.md +209 -0
  29. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/adapter_config.json +48 -0
  30. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/chat_template.jinja +85 -0
  31. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/tokenizer_config.json +29 -0
  32. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/trainer_state.json +742 -0
  33. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/README.md +209 -0
  34. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/adapter_config.json +48 -0
  35. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/chat_template.jinja +85 -0
  36. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/tokenizer_config.json +29 -0
  37. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/trainer_state.json +833 -0
  38. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/README.md +209 -0
  39. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/adapter_config.json +48 -0
  40. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/chat_template.jinja +85 -0
  41. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/tokenizer_config.json +29 -0
  42. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/trainer_state.json +115 -0
  43. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3890/README.md +209 -0
  44. productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3890/adapter_config.json +48 -0
  45. random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/README.md +209 -0
  46. random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/adapter_config.json +48 -0
  47. random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/chat_template.jinja +85 -0
  48. random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/tokenizer_config.json +29 -0
  49. random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/trainer_state.json +307 -0
  50. random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1632/README.md +209 -0
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test1/README.md ADDED
@@ -0,0 +1,58 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen3-14B-Base
3
+ library_name: transformers
4
+ model_name: Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test1
5
+ tags:
6
+ - generated_from_trainer
7
+ - trl
8
+ - sft
9
+ licence: license
10
+ ---
11
+
12
+ # Model Card for Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test1
13
+
14
+ This model is a fine-tuned version of [Qwen/Qwen3-14B-Base](https://huggingface.co/Qwen/Qwen3-14B-Base).
15
+ It has been trained using [TRL](https://github.com/huggingface/trl).
16
+
17
+ ## Quick start
18
+
19
+ ```python
20
+ from transformers import pipeline
21
+
22
+ question = "If you had a time machine, but could only go to the past or the future once and never return, which would you choose and why?"
23
+ generator = pipeline("text-generation", model="None", device="cuda")
24
+ output = generator([{"role": "user", "content": question}], max_new_tokens=128, return_full_text=False)[0]
25
+ print(output["generated_text"])
26
+ ```
27
+
28
+ ## Training procedure
29
+
30
+ [<img src="https://raw.githubusercontent.com/wandb/assets/main/wandb-github-badge-28.svg" alt="Visualize in Weights & Biases" width="150" height="24"/>](https://wandb.ai/katriin-kukk/Cross_lingual_morphological_generalization/runs/tjg90wvc)
31
+
32
+
33
+
34
+ This model was trained with SFT.
35
+
36
+ ### Framework versions
37
+
38
+ - TRL: 0.29.0
39
+ - Transformers: 5.5.4
40
+ - Pytorch: 2.10.0
41
+ - Datasets: 4.6.1
42
+ - Tokenizers: 0.22.2
43
+
44
+ ## Citations
45
+
46
+
47
+
48
+ Cite TRL as:
49
+
50
+ ```bibtex
51
+ @software{vonwerra2020trl,
52
+ title = {{TRL: Transformers Reinforcement Learning}},
53
+ author = {von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and Beeching, Edward and Thrush, Tristan and Lambert, Nathan and Huang, Shengyi and Rasul, Kashif and Gallouédec, Quentin},
54
+ license = {Apache-2.0},
55
+ url = {https://github.com/huggingface/trl},
56
+ year = {2020}
57
+ }
58
+ ```
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/README.md ADDED
@@ -0,0 +1,58 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen3-14B-Base
3
+ library_name: transformers
4
+ model_name: Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2
5
+ tags:
6
+ - generated_from_trainer
7
+ - trl
8
+ - sft
9
+ licence: license
10
+ ---
11
+
12
+ # Model Card for Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2
13
+
14
+ This model is a fine-tuned version of [Qwen/Qwen3-14B-Base](https://huggingface.co/Qwen/Qwen3-14B-Base).
15
+ It has been trained using [TRL](https://github.com/huggingface/trl).
16
+
17
+ ## Quick start
18
+
19
+ ```python
20
+ from transformers import pipeline
21
+
22
+ question = "If you had a time machine, but could only go to the past or the future once and never return, which would you choose and why?"
23
+ generator = pipeline("text-generation", model="None", device="cuda")
24
+ output = generator([{"role": "user", "content": question}], max_new_tokens=128, return_full_text=False)[0]
25
+ print(output["generated_text"])
26
+ ```
27
+
28
+ ## Training procedure
29
+
30
+ [<img src="https://raw.githubusercontent.com/wandb/assets/main/wandb-github-badge-28.svg" alt="Visualize in Weights & Biases" width="150" height="24"/>](https://wandb.ai/katriin-kukk/Cross_lingual_morphological_generalization/runs/3ulga1iu)
31
+
32
+
33
+
34
+ This model was trained with SFT.
35
+
36
+ ### Framework versions
37
+
38
+ - TRL: 0.29.0
39
+ - Transformers: 5.5.4
40
+ - Pytorch: 2.10.0
41
+ - Datasets: 4.6.1
42
+ - Tokenizers: 0.22.2
43
+
44
+ ## Citations
45
+
46
+
47
+
48
+ Cite TRL as:
49
+
50
+ ```bibtex
51
+ @software{vonwerra2020trl,
52
+ title = {{TRL: Transformers Reinforcement Learning}},
53
+ author = {von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and Beeching, Edward and Thrush, Tristan and Lambert, Nathan and Huang, Shengyi and Rasul, Kashif and Gallouédec, Quentin},
54
+ license = {Apache-2.0},
55
+ url = {https://github.com/huggingface/trl},
56
+ year = {2020}
57
+ }
58
+ ```
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/README.md ADDED
@@ -0,0 +1,209 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen3-14B-Base
3
+ library_name: peft
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - base_model:adapter:Qwen/Qwen3-14B-Base
7
+ - lora
8
+ - sft
9
+ - transformers
10
+ - trl
11
+ ---
12
+
13
+ # Model Card for Model ID
14
+
15
+ <!-- Provide a quick summary of what the model is/does. -->
16
+
17
+
18
+
19
+ ## Model Details
20
+
21
+ ### Model Description
22
+
23
+ <!-- Provide a longer summary of what this model is. -->
24
+
25
+
26
+
27
+ - **Developed by:** [More Information Needed]
28
+ - **Funded by [optional]:** [More Information Needed]
29
+ - **Shared by [optional]:** [More Information Needed]
30
+ - **Model type:** [More Information Needed]
31
+ - **Language(s) (NLP):** [More Information Needed]
32
+ - **License:** [More Information Needed]
33
+ - **Finetuned from model [optional]:** [More Information Needed]
34
+
35
+ ### Model Sources [optional]
36
+
37
+ <!-- Provide the basic links for the model. -->
38
+
39
+ - **Repository:** [More Information Needed]
40
+ - **Paper [optional]:** [More Information Needed]
41
+ - **Demo [optional]:** [More Information Needed]
42
+
43
+ ## Uses
44
+
45
+ <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
46
+
47
+ ### Direct Use
48
+
49
+ <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
50
+
51
+ [More Information Needed]
52
+
53
+ ### Downstream Use [optional]
54
+
55
+ <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
56
+
57
+ [More Information Needed]
58
+
59
+ ### Out-of-Scope Use
60
+
61
+ <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
62
+
63
+ [More Information Needed]
64
+
65
+ ## Bias, Risks, and Limitations
66
+
67
+ <!-- This section is meant to convey both technical and sociotechnical limitations. -->
68
+
69
+ [More Information Needed]
70
+
71
+ ### Recommendations
72
+
73
+ <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
74
+
75
+ Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
76
+
77
+ ## How to Get Started with the Model
78
+
79
+ Use the code below to get started with the model.
80
+
81
+ [More Information Needed]
82
+
83
+ ## Training Details
84
+
85
+ ### Training Data
86
+
87
+ <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
88
+
89
+ [More Information Needed]
90
+
91
+ ### Training Procedure
92
+
93
+ <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
94
+
95
+ #### Preprocessing [optional]
96
+
97
+ [More Information Needed]
98
+
99
+
100
+ #### Training Hyperparameters
101
+
102
+ - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
103
+
104
+ #### Speeds, Sizes, Times [optional]
105
+
106
+ <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
107
+
108
+ [More Information Needed]
109
+
110
+ ## Evaluation
111
+
112
+ <!-- This section describes the evaluation protocols and provides the results. -->
113
+
114
+ ### Testing Data, Factors & Metrics
115
+
116
+ #### Testing Data
117
+
118
+ <!-- This should link to a Dataset Card if possible. -->
119
+
120
+ [More Information Needed]
121
+
122
+ #### Factors
123
+
124
+ <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
125
+
126
+ [More Information Needed]
127
+
128
+ #### Metrics
129
+
130
+ <!-- These are the evaluation metrics being used, ideally with a description of why. -->
131
+
132
+ [More Information Needed]
133
+
134
+ ### Results
135
+
136
+ [More Information Needed]
137
+
138
+ #### Summary
139
+
140
+
141
+
142
+ ## Model Examination [optional]
143
+
144
+ <!-- Relevant interpretability work for the model goes here -->
145
+
146
+ [More Information Needed]
147
+
148
+ ## Environmental Impact
149
+
150
+ <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
151
+
152
+ Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
153
+
154
+ - **Hardware Type:** [More Information Needed]
155
+ - **Hours used:** [More Information Needed]
156
+ - **Cloud Provider:** [More Information Needed]
157
+ - **Compute Region:** [More Information Needed]
158
+ - **Carbon Emitted:** [More Information Needed]
159
+
160
+ ## Technical Specifications [optional]
161
+
162
+ ### Model Architecture and Objective
163
+
164
+ [More Information Needed]
165
+
166
+ ### Compute Infrastructure
167
+
168
+ [More Information Needed]
169
+
170
+ #### Hardware
171
+
172
+ [More Information Needed]
173
+
174
+ #### Software
175
+
176
+ [More Information Needed]
177
+
178
+ ## Citation [optional]
179
+
180
+ <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
181
+
182
+ **BibTeX:**
183
+
184
+ [More Information Needed]
185
+
186
+ **APA:**
187
+
188
+ [More Information Needed]
189
+
190
+ ## Glossary [optional]
191
+
192
+ <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
193
+
194
+ [More Information Needed]
195
+
196
+ ## More Information [optional]
197
+
198
+ [More Information Needed]
199
+
200
+ ## Model Card Authors [optional]
201
+
202
+ [More Information Needed]
203
+
204
+ ## Model Card Contact
205
+
206
+ [More Information Needed]
207
+ ### Framework versions
208
+
209
+ - PEFT 0.19.1
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/adapter_config.json ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "Qwen/Qwen3-14B-Base",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "layer_replication": null,
16
+ "layers_pattern": null,
17
+ "layers_to_transform": null,
18
+ "loftq_config": {},
19
+ "lora_alpha": 256,
20
+ "lora_bias": false,
21
+ "lora_dropout": 0.00237968804112545,
22
+ "lora_ga_config": null,
23
+ "megatron_config": null,
24
+ "megatron_core": "megatron.core",
25
+ "modules_to_save": null,
26
+ "peft_type": "LORA",
27
+ "peft_version": "0.19.1",
28
+ "qalora_group_size": 16,
29
+ "r": 128,
30
+ "rank_pattern": {},
31
+ "revision": null,
32
+ "target_modules": [
33
+ "gate_proj",
34
+ "q_proj",
35
+ "up_proj",
36
+ "o_proj",
37
+ "v_proj",
38
+ "k_proj",
39
+ "down_proj"
40
+ ],
41
+ "target_parameters": null,
42
+ "task_type": "CAUSAL_LM",
43
+ "trainable_token_indices": null,
44
+ "use_bdlora": null,
45
+ "use_dora": false,
46
+ "use_qalora": false,
47
+ "use_rslora": false
48
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/chat_template.jinja ADDED
@@ -0,0 +1,85 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- if tools %}
2
+ {{- '<|im_start|>system\n' }}
3
+ {%- if messages[0].role == 'system' %}
4
+ {{- messages[0].content + '\n\n' }}
5
+ {%- endif %}
6
+ {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
7
+ {%- for tool in tools %}
8
+ {{- "\n" }}
9
+ {{- tool | tojson }}
10
+ {%- endfor %}
11
+ {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
12
+ {%- else %}
13
+ {%- if messages[0].role == 'system' %}
14
+ {{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
15
+ {%- endif %}
16
+ {%- endif %}
17
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
18
+ {%- for message in messages[::-1] %}
19
+ {%- set index = (messages|length - 1) - loop.index0 %}
20
+ {%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
21
+ {%- set ns.multi_step_tool = false %}
22
+ {%- set ns.last_query_index = index %}
23
+ {%- endif %}
24
+ {%- endfor %}
25
+ {%- for message in messages %}
26
+ {%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
27
+ {{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
28
+ {%- elif message.role == "assistant" %}
29
+ {%- set content = message.content %}
30
+ {%- set reasoning_content = '' %}
31
+ {%- if message.reasoning_content is defined and message.reasoning_content is not none %}
32
+ {%- set reasoning_content = message.reasoning_content %}
33
+ {%- else %}
34
+ {%- if '</think>' in message.content %}
35
+ {%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
36
+ {%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
37
+ {%- endif %}
38
+ {%- endif %}
39
+ {%- if loop.index0 > ns.last_query_index %}
40
+ {%- if loop.last or (not loop.last and reasoning_content) %}
41
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
42
+ {%- else %}
43
+ {{- '<|im_start|>' + message.role + '\n' + content }}
44
+ {%- endif %}
45
+ {%- else %}
46
+ {{- '<|im_start|>' + message.role + '\n' + content }}
47
+ {%- endif %}
48
+ {%- if message.tool_calls %}
49
+ {%- for tool_call in message.tool_calls %}
50
+ {%- if (loop.first and content) or (not loop.first) %}
51
+ {{- '\n' }}
52
+ {%- endif %}
53
+ {%- if tool_call.function %}
54
+ {%- set tool_call = tool_call.function %}
55
+ {%- endif %}
56
+ {{- '<tool_call>\n{"name": "' }}
57
+ {{- tool_call.name }}
58
+ {{- '", "arguments": ' }}
59
+ {%- if tool_call.arguments is string %}
60
+ {{- tool_call.arguments }}
61
+ {%- else %}
62
+ {{- tool_call.arguments | tojson }}
63
+ {%- endif %}
64
+ {{- '}\n</tool_call>' }}
65
+ {%- endfor %}
66
+ {%- endif %}
67
+ {{- '<|im_end|>\n' }}
68
+ {%- elif message.role == "tool" %}
69
+ {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
70
+ {{- '<|im_start|>user' }}
71
+ {%- endif %}
72
+ {{- '\n<tool_response>\n' }}
73
+ {{- message.content }}
74
+ {{- '\n</tool_response>' }}
75
+ {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
76
+ {{- '<|im_end|>\n' }}
77
+ {%- endif %}
78
+ {%- endif %}
79
+ {%- endfor %}
80
+ {%- if add_generation_prompt %}
81
+ {{- '<|im_start|>assistant\n' }}
82
+ {%- if enable_thinking is defined and enable_thinking is false %}
83
+ {{- '<think>\n\n</think>\n\n' }}
84
+ {%- endif %}
85
+ {%- endif %}
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/tokenizer_config.json ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "backend": "tokenizers",
4
+ "bos_token": null,
5
+ "clean_up_tokenization_spaces": false,
6
+ "eos_token": "<|endoftext|>",
7
+ "errors": "replace",
8
+ "extra_special_tokens": [
9
+ "<|im_start|>",
10
+ "<|im_end|>",
11
+ "<|object_ref_start|>",
12
+ "<|object_ref_end|>",
13
+ "<|box_start|>",
14
+ "<|box_end|>",
15
+ "<|quad_start|>",
16
+ "<|quad_end|>",
17
+ "<|vision_start|>",
18
+ "<|vision_end|>",
19
+ "<|vision_pad|>",
20
+ "<|image_pad|>",
21
+ "<|video_pad|>"
22
+ ],
23
+ "is_local": false,
24
+ "model_max_length": 131072,
25
+ "pad_token": "<|endoftext|>",
26
+ "split_special_tokens": false,
27
+ "tokenizer_class": "Qwen2Tokenizer",
28
+ "unk_token": null
29
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/trainer_state.json ADDED
@@ -0,0 +1,297 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "best_global_step": null,
3
+ "best_metric": null,
4
+ "best_model_checkpoint": null,
5
+ "epoch": 3.0,
6
+ "eval_steps": 500,
7
+ "global_step": 1167,
8
+ "is_hyper_param_search": false,
9
+ "is_local_process_zero": true,
10
+ "is_world_process_zero": true,
11
+ "log_history": [
12
+ {
13
+ "entropy": 1.6110110306739807,
14
+ "epoch": 0.1287001287001287,
15
+ "grad_norm": 0.8611119389533997,
16
+ "learning_rate": 3.4130257962133866e-05,
17
+ "loss": 1.53956298828125,
18
+ "mean_token_accuracy": 0.6652342769503593,
19
+ "num_tokens": 73407.0,
20
+ "step": 50
21
+ },
22
+ {
23
+ "entropy": 0.8662820833921433,
24
+ "epoch": 0.2574002574002574,
25
+ "grad_norm": 0.7405035495758057,
26
+ "learning_rate": 6.895705180104598e-05,
27
+ "loss": 0.7929539489746094,
28
+ "mean_token_accuracy": 0.7786347842216492,
29
+ "num_tokens": 143994.0,
30
+ "step": 100
31
+ },
32
+ {
33
+ "entropy": 0.7912819278240204,
34
+ "epoch": 0.3861003861003861,
35
+ "grad_norm": 0.6105485558509827,
36
+ "learning_rate": 0.00010378384563995809,
37
+ "loss": 0.7236511993408203,
38
+ "mean_token_accuracy": 0.7925467795133591,
39
+ "num_tokens": 216171.0,
40
+ "step": 150
41
+ },
42
+ {
43
+ "entropy": 0.7777709531784057,
44
+ "epoch": 0.5148005148005148,
45
+ "grad_norm": 0.5781793594360352,
46
+ "learning_rate": 0.0001386106394788702,
47
+ "loss": 0.7024919891357422,
48
+ "mean_token_accuracy": 0.7969876372814179,
49
+ "num_tokens": 284702.0,
50
+ "step": 200
51
+ },
52
+ {
53
+ "entropy": 0.7632609683275223,
54
+ "epoch": 0.6435006435006435,
55
+ "grad_norm": 0.6553444862365723,
56
+ "learning_rate": 0.00017343743331778232,
57
+ "loss": 0.7000718688964844,
58
+ "mean_token_accuracy": 0.7994742071628571,
59
+ "num_tokens": 356393.0,
60
+ "step": 250
61
+ },
62
+ {
63
+ "entropy": 0.754165632724762,
64
+ "epoch": 0.7722007722007722,
65
+ "grad_norm": 0.634222149848938,
66
+ "learning_rate": 0.00020826422715669444,
67
+ "loss": 0.6909049987792969,
68
+ "mean_token_accuracy": 0.8016809666156769,
69
+ "num_tokens": 426916.0,
70
+ "step": 300
71
+ },
72
+ {
73
+ "entropy": 0.7530711203813553,
74
+ "epoch": 0.9009009009009009,
75
+ "grad_norm": 0.7020156383514404,
76
+ "learning_rate": 0.00024309102099560653,
77
+ "loss": 0.69554931640625,
78
+ "mean_token_accuracy": 0.8002807641029358,
79
+ "num_tokens": 499603.0,
80
+ "step": 350
81
+ },
82
+ {
83
+ "epoch": 1.0,
84
+ "eval_entropy": 0.5728044164242204,
85
+ "eval_loss": 0.6785851716995239,
86
+ "eval_mean_token_accuracy": 0.798433246993527,
87
+ "eval_num_tokens": 555493.0,
88
+ "eval_runtime": 162.4129,
89
+ "eval_samples_per_second": 9.513,
90
+ "eval_steps_per_second": 1.194,
91
+ "step": 389
92
+ },
93
+ {
94
+ "entropy": 0.7489313946829902,
95
+ "epoch": 1.0283140283140284,
96
+ "grad_norm": 0.7532815933227539,
97
+ "learning_rate": 0.0002709470016827303,
98
+ "loss": 0.6856581878662109,
99
+ "mean_token_accuracy": 0.8023027079273956,
100
+ "num_tokens": 570813.0,
101
+ "step": 400
102
+ },
103
+ {
104
+ "entropy": 0.7238409864902496,
105
+ "epoch": 1.157014157014157,
106
+ "grad_norm": 1.2016215324401855,
107
+ "learning_rate": 0.0002707561443541359,
108
+ "loss": 0.6699818420410156,
109
+ "mean_token_accuracy": 0.8070302194356919,
110
+ "num_tokens": 642956.0,
111
+ "step": 450
112
+ },
113
+ {
114
+ "entropy": 0.7388938587903976,
115
+ "epoch": 1.2857142857142856,
116
+ "grad_norm": 0.7279272079467773,
117
+ "learning_rate": 0.0002702930068622498,
118
+ "loss": 0.6728517150878907,
119
+ "mean_token_accuracy": 0.8049580943584442,
120
+ "num_tokens": 714498.0,
121
+ "step": 500
122
+ },
123
+ {
124
+ "entropy": 0.7284485149383545,
125
+ "epoch": 1.4144144144144144,
126
+ "grad_norm": 0.8563987016677856,
127
+ "learning_rate": 0.0002695585213716931,
128
+ "loss": 0.6657986450195312,
129
+ "mean_token_accuracy": 0.8085588800907135,
130
+ "num_tokens": 785595.0,
131
+ "step": 550
132
+ },
133
+ {
134
+ "entropy": 0.6960959500074386,
135
+ "epoch": 1.5431145431145432,
136
+ "grad_norm": 0.5145474672317505,
137
+ "learning_rate": 0.0002685541661937683,
138
+ "loss": 0.6358638763427734,
139
+ "mean_token_accuracy": 0.8131621342897415,
140
+ "num_tokens": 857551.0,
141
+ "step": 600
142
+ },
143
+ {
144
+ "entropy": 0.7220414417982102,
145
+ "epoch": 1.6718146718146718,
146
+ "grad_norm": 0.9381059408187866,
147
+ "learning_rate": 0.00026728196281103746,
148
+ "loss": 0.6531407928466797,
149
+ "mean_token_accuracy": 0.811619822382927,
150
+ "num_tokens": 926864.0,
151
+ "step": 650
152
+ },
153
+ {
154
+ "entropy": 0.6955459499359131,
155
+ "epoch": 1.8005148005148004,
156
+ "grad_norm": 0.6115108728408813,
157
+ "learning_rate": 0.0002657444718086503,
158
+ "loss": 0.6269588088989257,
159
+ "mean_token_accuracy": 0.8155373805761337,
160
+ "num_tokens": 998824.0,
161
+ "step": 700
162
+ },
163
+ {
164
+ "entropy": 0.6929812705516816,
165
+ "epoch": 1.9292149292149292,
166
+ "grad_norm": 0.749489426612854,
167
+ "learning_rate": 0.0002639447877206115,
168
+ "loss": 0.629054069519043,
169
+ "mean_token_accuracy": 0.8183021235466004,
170
+ "num_tokens": 1069332.0,
171
+ "step": 750
172
+ },
173
+ {
174
+ "epoch": 2.0,
175
+ "eval_entropy": 0.5833694102223387,
176
+ "eval_loss": 0.6328718662261963,
177
+ "eval_mean_token_accuracy": 0.8159044071571114,
178
+ "eval_num_tokens": 1110986.0,
179
+ "eval_runtime": 161.885,
180
+ "eval_samples_per_second": 9.544,
181
+ "eval_steps_per_second": 1.198,
182
+ "step": 778
183
+ },
184
+ {
185
+ "entropy": 0.647241060480927,
186
+ "epoch": 2.056628056628057,
187
+ "grad_norm": 0.7070767879486084,
188
+ "learning_rate": 0.00026188653280135975,
189
+ "loss": 0.5823922348022461,
190
+ "mean_token_accuracy": 0.8265365301960647,
191
+ "num_tokens": 1141195.0,
192
+ "step": 800
193
+ },
194
+ {
195
+ "entropy": 0.5995378407835961,
196
+ "epoch": 2.1853281853281854,
197
+ "grad_norm": 0.8090486526489258,
198
+ "learning_rate": 0.0002595738497351955,
199
+ "loss": 0.5325597763061524,
200
+ "mean_token_accuracy": 0.8369336777925491,
201
+ "num_tokens": 1210708.0,
202
+ "step": 850
203
+ },
204
+ {
205
+ "entropy": 0.6043145382404327,
206
+ "epoch": 2.314028314028314,
207
+ "grad_norm": 0.8279913067817688,
208
+ "learning_rate": 0.00025701139329823054,
209
+ "loss": 0.5414446258544922,
210
+ "mean_token_accuracy": 0.8361396533250809,
211
+ "num_tokens": 1283441.0,
212
+ "step": 900
213
+ },
214
+ {
215
+ "entropy": 0.5953224584460258,
216
+ "epoch": 2.4427284427284426,
217
+ "grad_norm": 0.6075023412704468,
218
+ "learning_rate": 0.00025420432098964183,
219
+ "loss": 0.536654167175293,
220
+ "mean_token_accuracy": 0.8340093129873276,
221
+ "num_tokens": 1356418.0,
222
+ "step": 950
223
+ },
224
+ {
225
+ "entropy": 0.5998479858040809,
226
+ "epoch": 2.571428571428571,
227
+ "grad_norm": 1.0311471223831177,
228
+ "learning_rate": 0.0002511582826510862,
229
+ "loss": 0.5372924423217773,
230
+ "mean_token_accuracy": 0.8366045409440994,
231
+ "num_tokens": 1427797.0,
232
+ "step": 1000
233
+ },
234
+ {
235
+ "entropy": 0.5980158120393753,
236
+ "epoch": 2.7001287001287,
237
+ "grad_norm": 0.5971426367759705,
238
+ "learning_rate": 0.0002478794090951689,
239
+ "loss": 0.5392885208129883,
240
+ "mean_token_accuracy": 0.8347727072238922,
241
+ "num_tokens": 1498082.0,
242
+ "step": 1050
243
+ },
244
+ {
245
+ "entropy": 0.5990680930018425,
246
+ "epoch": 2.828828828828829,
247
+ "grad_norm": 0.5662627220153809,
248
+ "learning_rate": 0.0002443742997658538,
249
+ "loss": 0.5360498428344727,
250
+ "mean_token_accuracy": 0.8371847170591354,
251
+ "num_tokens": 1568798.0,
252
+ "step": 1100
253
+ },
254
+ {
255
+ "entropy": 0.5957039377093315,
256
+ "epoch": 2.9575289575289574,
257
+ "grad_norm": 0.5043798685073853,
258
+ "learning_rate": 0.00024065000945565205,
259
+ "loss": 0.5342231369018555,
260
+ "mean_token_accuracy": 0.8380735236406326,
261
+ "num_tokens": 1643449.0,
262
+ "step": 1150
263
+ },
264
+ {
265
+ "epoch": 3.0,
266
+ "eval_entropy": 0.5346978819861854,
267
+ "eval_loss": 0.6255015134811401,
268
+ "eval_mean_token_accuracy": 0.8202014476368108,
269
+ "eval_num_tokens": 1666479.0,
270
+ "eval_runtime": 161.6098,
271
+ "eval_samples_per_second": 9.56,
272
+ "eval_steps_per_second": 1.2,
273
+ "step": 1167
274
+ }
275
+ ],
276
+ "logging_steps": 50,
277
+ "max_steps": 3890,
278
+ "num_input_tokens_seen": 0,
279
+ "num_train_epochs": 10,
280
+ "save_steps": 500,
281
+ "stateful_callbacks": {
282
+ "TrainerControl": {
283
+ "args": {
284
+ "should_epoch_stop": false,
285
+ "should_evaluate": false,
286
+ "should_log": false,
287
+ "should_save": true,
288
+ "should_training_stop": false
289
+ },
290
+ "attributes": {}
291
+ }
292
+ },
293
+ "total_flos": 2.7917443648582656e+17,
294
+ "train_batch_size": 8,
295
+ "trial_name": null,
296
+ "trial_params": null
297
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/README.md ADDED
@@ -0,0 +1,209 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen3-14B-Base
3
+ library_name: peft
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - base_model:adapter:Qwen/Qwen3-14B-Base
7
+ - lora
8
+ - sft
9
+ - transformers
10
+ - trl
11
+ ---
12
+
13
+ # Model Card for Model ID
14
+
15
+ <!-- Provide a quick summary of what the model is/does. -->
16
+
17
+
18
+
19
+ ## Model Details
20
+
21
+ ### Model Description
22
+
23
+ <!-- Provide a longer summary of what this model is. -->
24
+
25
+
26
+
27
+ - **Developed by:** [More Information Needed]
28
+ - **Funded by [optional]:** [More Information Needed]
29
+ - **Shared by [optional]:** [More Information Needed]
30
+ - **Model type:** [More Information Needed]
31
+ - **Language(s) (NLP):** [More Information Needed]
32
+ - **License:** [More Information Needed]
33
+ - **Finetuned from model [optional]:** [More Information Needed]
34
+
35
+ ### Model Sources [optional]
36
+
37
+ <!-- Provide the basic links for the model. -->
38
+
39
+ - **Repository:** [More Information Needed]
40
+ - **Paper [optional]:** [More Information Needed]
41
+ - **Demo [optional]:** [More Information Needed]
42
+
43
+ ## Uses
44
+
45
+ <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
46
+
47
+ ### Direct Use
48
+
49
+ <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
50
+
51
+ [More Information Needed]
52
+
53
+ ### Downstream Use [optional]
54
+
55
+ <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
56
+
57
+ [More Information Needed]
58
+
59
+ ### Out-of-Scope Use
60
+
61
+ <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
62
+
63
+ [More Information Needed]
64
+
65
+ ## Bias, Risks, and Limitations
66
+
67
+ <!-- This section is meant to convey both technical and sociotechnical limitations. -->
68
+
69
+ [More Information Needed]
70
+
71
+ ### Recommendations
72
+
73
+ <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
74
+
75
+ Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
76
+
77
+ ## How to Get Started with the Model
78
+
79
+ Use the code below to get started with the model.
80
+
81
+ [More Information Needed]
82
+
83
+ ## Training Details
84
+
85
+ ### Training Data
86
+
87
+ <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
88
+
89
+ [More Information Needed]
90
+
91
+ ### Training Procedure
92
+
93
+ <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
94
+
95
+ #### Preprocessing [optional]
96
+
97
+ [More Information Needed]
98
+
99
+
100
+ #### Training Hyperparameters
101
+
102
+ - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
103
+
104
+ #### Speeds, Sizes, Times [optional]
105
+
106
+ <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
107
+
108
+ [More Information Needed]
109
+
110
+ ## Evaluation
111
+
112
+ <!-- This section describes the evaluation protocols and provides the results. -->
113
+
114
+ ### Testing Data, Factors & Metrics
115
+
116
+ #### Testing Data
117
+
118
+ <!-- This should link to a Dataset Card if possible. -->
119
+
120
+ [More Information Needed]
121
+
122
+ #### Factors
123
+
124
+ <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
125
+
126
+ [More Information Needed]
127
+
128
+ #### Metrics
129
+
130
+ <!-- These are the evaluation metrics being used, ideally with a description of why. -->
131
+
132
+ [More Information Needed]
133
+
134
+ ### Results
135
+
136
+ [More Information Needed]
137
+
138
+ #### Summary
139
+
140
+
141
+
142
+ ## Model Examination [optional]
143
+
144
+ <!-- Relevant interpretability work for the model goes here -->
145
+
146
+ [More Information Needed]
147
+
148
+ ## Environmental Impact
149
+
150
+ <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
151
+
152
+ Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
153
+
154
+ - **Hardware Type:** [More Information Needed]
155
+ - **Hours used:** [More Information Needed]
156
+ - **Cloud Provider:** [More Information Needed]
157
+ - **Compute Region:** [More Information Needed]
158
+ - **Carbon Emitted:** [More Information Needed]
159
+
160
+ ## Technical Specifications [optional]
161
+
162
+ ### Model Architecture and Objective
163
+
164
+ [More Information Needed]
165
+
166
+ ### Compute Infrastructure
167
+
168
+ [More Information Needed]
169
+
170
+ #### Hardware
171
+
172
+ [More Information Needed]
173
+
174
+ #### Software
175
+
176
+ [More Information Needed]
177
+
178
+ ## Citation [optional]
179
+
180
+ <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
181
+
182
+ **BibTeX:**
183
+
184
+ [More Information Needed]
185
+
186
+ **APA:**
187
+
188
+ [More Information Needed]
189
+
190
+ ## Glossary [optional]
191
+
192
+ <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
193
+
194
+ [More Information Needed]
195
+
196
+ ## More Information [optional]
197
+
198
+ [More Information Needed]
199
+
200
+ ## Model Card Authors [optional]
201
+
202
+ [More Information Needed]
203
+
204
+ ## Model Card Contact
205
+
206
+ [More Information Needed]
207
+ ### Framework versions
208
+
209
+ - PEFT 0.19.1
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/adapter_config.json ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "Qwen/Qwen3-14B-Base",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "layer_replication": null,
16
+ "layers_pattern": null,
17
+ "layers_to_transform": null,
18
+ "loftq_config": {},
19
+ "lora_alpha": 256,
20
+ "lora_bias": false,
21
+ "lora_dropout": 0.00237968804112545,
22
+ "lora_ga_config": null,
23
+ "megatron_config": null,
24
+ "megatron_core": "megatron.core",
25
+ "modules_to_save": null,
26
+ "peft_type": "LORA",
27
+ "peft_version": "0.19.1",
28
+ "qalora_group_size": 16,
29
+ "r": 128,
30
+ "rank_pattern": {},
31
+ "revision": null,
32
+ "target_modules": [
33
+ "gate_proj",
34
+ "q_proj",
35
+ "up_proj",
36
+ "o_proj",
37
+ "v_proj",
38
+ "k_proj",
39
+ "down_proj"
40
+ ],
41
+ "target_parameters": null,
42
+ "task_type": "CAUSAL_LM",
43
+ "trainable_token_indices": null,
44
+ "use_bdlora": null,
45
+ "use_dora": false,
46
+ "use_qalora": false,
47
+ "use_rslora": false
48
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/chat_template.jinja ADDED
@@ -0,0 +1,85 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- if tools %}
2
+ {{- '<|im_start|>system\n' }}
3
+ {%- if messages[0].role == 'system' %}
4
+ {{- messages[0].content + '\n\n' }}
5
+ {%- endif %}
6
+ {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
7
+ {%- for tool in tools %}
8
+ {{- "\n" }}
9
+ {{- tool | tojson }}
10
+ {%- endfor %}
11
+ {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
12
+ {%- else %}
13
+ {%- if messages[0].role == 'system' %}
14
+ {{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
15
+ {%- endif %}
16
+ {%- endif %}
17
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
18
+ {%- for message in messages[::-1] %}
19
+ {%- set index = (messages|length - 1) - loop.index0 %}
20
+ {%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
21
+ {%- set ns.multi_step_tool = false %}
22
+ {%- set ns.last_query_index = index %}
23
+ {%- endif %}
24
+ {%- endfor %}
25
+ {%- for message in messages %}
26
+ {%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
27
+ {{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
28
+ {%- elif message.role == "assistant" %}
29
+ {%- set content = message.content %}
30
+ {%- set reasoning_content = '' %}
31
+ {%- if message.reasoning_content is defined and message.reasoning_content is not none %}
32
+ {%- set reasoning_content = message.reasoning_content %}
33
+ {%- else %}
34
+ {%- if '</think>' in message.content %}
35
+ {%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
36
+ {%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
37
+ {%- endif %}
38
+ {%- endif %}
39
+ {%- if loop.index0 > ns.last_query_index %}
40
+ {%- if loop.last or (not loop.last and reasoning_content) %}
41
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
42
+ {%- else %}
43
+ {{- '<|im_start|>' + message.role + '\n' + content }}
44
+ {%- endif %}
45
+ {%- else %}
46
+ {{- '<|im_start|>' + message.role + '\n' + content }}
47
+ {%- endif %}
48
+ {%- if message.tool_calls %}
49
+ {%- for tool_call in message.tool_calls %}
50
+ {%- if (loop.first and content) or (not loop.first) %}
51
+ {{- '\n' }}
52
+ {%- endif %}
53
+ {%- if tool_call.function %}
54
+ {%- set tool_call = tool_call.function %}
55
+ {%- endif %}
56
+ {{- '<tool_call>\n{"name": "' }}
57
+ {{- tool_call.name }}
58
+ {{- '", "arguments": ' }}
59
+ {%- if tool_call.arguments is string %}
60
+ {{- tool_call.arguments }}
61
+ {%- else %}
62
+ {{- tool_call.arguments | tojson }}
63
+ {%- endif %}
64
+ {{- '}\n</tool_call>' }}
65
+ {%- endfor %}
66
+ {%- endif %}
67
+ {{- '<|im_end|>\n' }}
68
+ {%- elif message.role == "tool" %}
69
+ {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
70
+ {{- '<|im_start|>user' }}
71
+ {%- endif %}
72
+ {{- '\n<tool_response>\n' }}
73
+ {{- message.content }}
74
+ {{- '\n</tool_response>' }}
75
+ {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
76
+ {{- '<|im_end|>\n' }}
77
+ {%- endif %}
78
+ {%- endif %}
79
+ {%- endfor %}
80
+ {%- if add_generation_prompt %}
81
+ {{- '<|im_start|>assistant\n' }}
82
+ {%- if enable_thinking is defined and enable_thinking is false %}
83
+ {{- '<think>\n\n</think>\n\n' }}
84
+ {%- endif %}
85
+ {%- endif %}
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/tokenizer_config.json ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "backend": "tokenizers",
4
+ "bos_token": null,
5
+ "clean_up_tokenization_spaces": false,
6
+ "eos_token": "<|endoftext|>",
7
+ "errors": "replace",
8
+ "extra_special_tokens": [
9
+ "<|im_start|>",
10
+ "<|im_end|>",
11
+ "<|object_ref_start|>",
12
+ "<|object_ref_end|>",
13
+ "<|box_start|>",
14
+ "<|box_end|>",
15
+ "<|quad_start|>",
16
+ "<|quad_end|>",
17
+ "<|vision_start|>",
18
+ "<|vision_end|>",
19
+ "<|vision_pad|>",
20
+ "<|image_pad|>",
21
+ "<|video_pad|>"
22
+ ],
23
+ "is_local": false,
24
+ "model_max_length": 131072,
25
+ "pad_token": "<|endoftext|>",
26
+ "split_special_tokens": false,
27
+ "tokenizer_class": "Qwen2Tokenizer",
28
+ "unk_token": null
29
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/trainer_state.json ADDED
@@ -0,0 +1,388 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "best_global_step": null,
3
+ "best_metric": null,
4
+ "best_model_checkpoint": null,
5
+ "epoch": 4.0,
6
+ "eval_steps": 500,
7
+ "global_step": 1556,
8
+ "is_hyper_param_search": false,
9
+ "is_local_process_zero": true,
10
+ "is_world_process_zero": true,
11
+ "log_history": [
12
+ {
13
+ "entropy": 1.6110110306739807,
14
+ "epoch": 0.1287001287001287,
15
+ "grad_norm": 0.8611119389533997,
16
+ "learning_rate": 3.4130257962133866e-05,
17
+ "loss": 1.53956298828125,
18
+ "mean_token_accuracy": 0.6652342769503593,
19
+ "num_tokens": 73407.0,
20
+ "step": 50
21
+ },
22
+ {
23
+ "entropy": 0.8662820833921433,
24
+ "epoch": 0.2574002574002574,
25
+ "grad_norm": 0.7405035495758057,
26
+ "learning_rate": 6.895705180104598e-05,
27
+ "loss": 0.7929539489746094,
28
+ "mean_token_accuracy": 0.7786347842216492,
29
+ "num_tokens": 143994.0,
30
+ "step": 100
31
+ },
32
+ {
33
+ "entropy": 0.7912819278240204,
34
+ "epoch": 0.3861003861003861,
35
+ "grad_norm": 0.6105485558509827,
36
+ "learning_rate": 0.00010378384563995809,
37
+ "loss": 0.7236511993408203,
38
+ "mean_token_accuracy": 0.7925467795133591,
39
+ "num_tokens": 216171.0,
40
+ "step": 150
41
+ },
42
+ {
43
+ "entropy": 0.7777709531784057,
44
+ "epoch": 0.5148005148005148,
45
+ "grad_norm": 0.5781793594360352,
46
+ "learning_rate": 0.0001386106394788702,
47
+ "loss": 0.7024919891357422,
48
+ "mean_token_accuracy": 0.7969876372814179,
49
+ "num_tokens": 284702.0,
50
+ "step": 200
51
+ },
52
+ {
53
+ "entropy": 0.7632609683275223,
54
+ "epoch": 0.6435006435006435,
55
+ "grad_norm": 0.6553444862365723,
56
+ "learning_rate": 0.00017343743331778232,
57
+ "loss": 0.7000718688964844,
58
+ "mean_token_accuracy": 0.7994742071628571,
59
+ "num_tokens": 356393.0,
60
+ "step": 250
61
+ },
62
+ {
63
+ "entropy": 0.754165632724762,
64
+ "epoch": 0.7722007722007722,
65
+ "grad_norm": 0.634222149848938,
66
+ "learning_rate": 0.00020826422715669444,
67
+ "loss": 0.6909049987792969,
68
+ "mean_token_accuracy": 0.8016809666156769,
69
+ "num_tokens": 426916.0,
70
+ "step": 300
71
+ },
72
+ {
73
+ "entropy": 0.7530711203813553,
74
+ "epoch": 0.9009009009009009,
75
+ "grad_norm": 0.7020156383514404,
76
+ "learning_rate": 0.00024309102099560653,
77
+ "loss": 0.69554931640625,
78
+ "mean_token_accuracy": 0.8002807641029358,
79
+ "num_tokens": 499603.0,
80
+ "step": 350
81
+ },
82
+ {
83
+ "epoch": 1.0,
84
+ "eval_entropy": 0.5728044164242204,
85
+ "eval_loss": 0.6785851716995239,
86
+ "eval_mean_token_accuracy": 0.798433246993527,
87
+ "eval_num_tokens": 555493.0,
88
+ "eval_runtime": 162.4129,
89
+ "eval_samples_per_second": 9.513,
90
+ "eval_steps_per_second": 1.194,
91
+ "step": 389
92
+ },
93
+ {
94
+ "entropy": 0.7489313946829902,
95
+ "epoch": 1.0283140283140284,
96
+ "grad_norm": 0.7532815933227539,
97
+ "learning_rate": 0.0002709470016827303,
98
+ "loss": 0.6856581878662109,
99
+ "mean_token_accuracy": 0.8023027079273956,
100
+ "num_tokens": 570813.0,
101
+ "step": 400
102
+ },
103
+ {
104
+ "entropy": 0.7238409864902496,
105
+ "epoch": 1.157014157014157,
106
+ "grad_norm": 1.2016215324401855,
107
+ "learning_rate": 0.0002707561443541359,
108
+ "loss": 0.6699818420410156,
109
+ "mean_token_accuracy": 0.8070302194356919,
110
+ "num_tokens": 642956.0,
111
+ "step": 450
112
+ },
113
+ {
114
+ "entropy": 0.7388938587903976,
115
+ "epoch": 1.2857142857142856,
116
+ "grad_norm": 0.7279272079467773,
117
+ "learning_rate": 0.0002702930068622498,
118
+ "loss": 0.6728517150878907,
119
+ "mean_token_accuracy": 0.8049580943584442,
120
+ "num_tokens": 714498.0,
121
+ "step": 500
122
+ },
123
+ {
124
+ "entropy": 0.7284485149383545,
125
+ "epoch": 1.4144144144144144,
126
+ "grad_norm": 0.8563987016677856,
127
+ "learning_rate": 0.0002695585213716931,
128
+ "loss": 0.6657986450195312,
129
+ "mean_token_accuracy": 0.8085588800907135,
130
+ "num_tokens": 785595.0,
131
+ "step": 550
132
+ },
133
+ {
134
+ "entropy": 0.6960959500074386,
135
+ "epoch": 1.5431145431145432,
136
+ "grad_norm": 0.5145474672317505,
137
+ "learning_rate": 0.0002685541661937683,
138
+ "loss": 0.6358638763427734,
139
+ "mean_token_accuracy": 0.8131621342897415,
140
+ "num_tokens": 857551.0,
141
+ "step": 600
142
+ },
143
+ {
144
+ "entropy": 0.7220414417982102,
145
+ "epoch": 1.6718146718146718,
146
+ "grad_norm": 0.9381059408187866,
147
+ "learning_rate": 0.00026728196281103746,
148
+ "loss": 0.6531407928466797,
149
+ "mean_token_accuracy": 0.811619822382927,
150
+ "num_tokens": 926864.0,
151
+ "step": 650
152
+ },
153
+ {
154
+ "entropy": 0.6955459499359131,
155
+ "epoch": 1.8005148005148004,
156
+ "grad_norm": 0.6115108728408813,
157
+ "learning_rate": 0.0002657444718086503,
158
+ "loss": 0.6269588088989257,
159
+ "mean_token_accuracy": 0.8155373805761337,
160
+ "num_tokens": 998824.0,
161
+ "step": 700
162
+ },
163
+ {
164
+ "entropy": 0.6929812705516816,
165
+ "epoch": 1.9292149292149292,
166
+ "grad_norm": 0.749489426612854,
167
+ "learning_rate": 0.0002639447877206115,
168
+ "loss": 0.629054069519043,
169
+ "mean_token_accuracy": 0.8183021235466004,
170
+ "num_tokens": 1069332.0,
171
+ "step": 750
172
+ },
173
+ {
174
+ "epoch": 2.0,
175
+ "eval_entropy": 0.5833694102223387,
176
+ "eval_loss": 0.6328718662261963,
177
+ "eval_mean_token_accuracy": 0.8159044071571114,
178
+ "eval_num_tokens": 1110986.0,
179
+ "eval_runtime": 161.885,
180
+ "eval_samples_per_second": 9.544,
181
+ "eval_steps_per_second": 1.198,
182
+ "step": 778
183
+ },
184
+ {
185
+ "entropy": 0.647241060480927,
186
+ "epoch": 2.056628056628057,
187
+ "grad_norm": 0.7070767879486084,
188
+ "learning_rate": 0.00026188653280135975,
189
+ "loss": 0.5823922348022461,
190
+ "mean_token_accuracy": 0.8265365301960647,
191
+ "num_tokens": 1141195.0,
192
+ "step": 800
193
+ },
194
+ {
195
+ "entropy": 0.5995378407835961,
196
+ "epoch": 2.1853281853281854,
197
+ "grad_norm": 0.8090486526489258,
198
+ "learning_rate": 0.0002595738497351955,
199
+ "loss": 0.5325597763061524,
200
+ "mean_token_accuracy": 0.8369336777925491,
201
+ "num_tokens": 1210708.0,
202
+ "step": 850
203
+ },
204
+ {
205
+ "entropy": 0.6043145382404327,
206
+ "epoch": 2.314028314028314,
207
+ "grad_norm": 0.8279913067817688,
208
+ "learning_rate": 0.00025701139329823054,
209
+ "loss": 0.5414446258544922,
210
+ "mean_token_accuracy": 0.8361396533250809,
211
+ "num_tokens": 1283441.0,
212
+ "step": 900
213
+ },
214
+ {
215
+ "entropy": 0.5953224584460258,
216
+ "epoch": 2.4427284427284426,
217
+ "grad_norm": 0.6075023412704468,
218
+ "learning_rate": 0.00025420432098964183,
219
+ "loss": 0.536654167175293,
220
+ "mean_token_accuracy": 0.8340093129873276,
221
+ "num_tokens": 1356418.0,
222
+ "step": 950
223
+ },
224
+ {
225
+ "entropy": 0.5998479858040809,
226
+ "epoch": 2.571428571428571,
227
+ "grad_norm": 1.0311471223831177,
228
+ "learning_rate": 0.0002511582826510862,
229
+ "loss": 0.5372924423217773,
230
+ "mean_token_accuracy": 0.8366045409440994,
231
+ "num_tokens": 1427797.0,
232
+ "step": 1000
233
+ },
234
+ {
235
+ "entropy": 0.5980158120393753,
236
+ "epoch": 2.7001287001287,
237
+ "grad_norm": 0.5971426367759705,
238
+ "learning_rate": 0.0002478794090951689,
239
+ "loss": 0.5392885208129883,
240
+ "mean_token_accuracy": 0.8347727072238922,
241
+ "num_tokens": 1498082.0,
242
+ "step": 1050
243
+ },
244
+ {
245
+ "entropy": 0.5990680930018425,
246
+ "epoch": 2.828828828828829,
247
+ "grad_norm": 0.5662627220153809,
248
+ "learning_rate": 0.0002443742997658538,
249
+ "loss": 0.5360498428344727,
250
+ "mean_token_accuracy": 0.8371847170591354,
251
+ "num_tokens": 1568798.0,
252
+ "step": 1100
253
+ },
254
+ {
255
+ "entropy": 0.5957039377093315,
256
+ "epoch": 2.9575289575289574,
257
+ "grad_norm": 0.5043798685073853,
258
+ "learning_rate": 0.00024065000945565205,
259
+ "loss": 0.5342231369018555,
260
+ "mean_token_accuracy": 0.8380735236406326,
261
+ "num_tokens": 1643449.0,
262
+ "step": 1150
263
+ },
264
+ {
265
+ "epoch": 3.0,
266
+ "eval_entropy": 0.5346978819861854,
267
+ "eval_loss": 0.6255015134811401,
268
+ "eval_mean_token_accuracy": 0.8202014476368108,
269
+ "eval_num_tokens": 1666479.0,
270
+ "eval_runtime": 161.6098,
271
+ "eval_samples_per_second": 9.56,
272
+ "eval_steps_per_second": 1.2,
273
+ "step": 1167
274
+ },
275
+ {
276
+ "entropy": 0.5212209137401196,
277
+ "epoch": 3.0849420849420848,
278
+ "grad_norm": 0.7095440626144409,
279
+ "learning_rate": 0.00023671403410632178,
280
+ "loss": 0.45311901092529294,
281
+ "mean_token_accuracy": 0.856362871449403,
282
+ "num_tokens": 1713536.0,
283
+ "step": 1200
284
+ },
285
+ {
286
+ "entropy": 0.4701593083143234,
287
+ "epoch": 3.213642213642214,
288
+ "grad_norm": 0.6408083438873291,
289
+ "learning_rate": 0.0002325742957216607,
290
+ "loss": 0.39916397094726563,
291
+ "mean_token_accuracy": 0.8698061722517013,
292
+ "num_tokens": 1785609.0,
293
+ "step": 1250
294
+ },
295
+ {
296
+ "entropy": 0.4766591975092888,
297
+ "epoch": 3.3423423423423424,
298
+ "grad_norm": 0.6415093541145325,
299
+ "learning_rate": 0.0002282391264227552,
300
+ "loss": 0.4116698455810547,
301
+ "mean_token_accuracy": 0.8679435575008392,
302
+ "num_tokens": 1858651.0,
303
+ "step": 1300
304
+ },
305
+ {
306
+ "entropy": 0.4937947469949722,
307
+ "epoch": 3.471042471042471,
308
+ "grad_norm": 0.6549825072288513,
309
+ "learning_rate": 0.00022371725167778054,
310
+ "loss": 0.4296376037597656,
311
+ "mean_token_accuracy": 0.8609357953071595,
312
+ "num_tokens": 1928692.0,
313
+ "step": 1350
314
+ },
315
+ {
316
+ "entropy": 0.4920153194665909,
317
+ "epoch": 3.5997425997425996,
318
+ "grad_norm": 0.6452126502990723,
319
+ "learning_rate": 0.00021901777274010406,
320
+ "loss": 0.4307489013671875,
321
+ "mean_token_accuracy": 0.8606827831268311,
322
+ "num_tokens": 1998668.0,
323
+ "step": 1400
324
+ },
325
+ {
326
+ "entropy": 0.490042342543602,
327
+ "epoch": 3.7284427284427286,
328
+ "grad_norm": 0.5727734565734863,
329
+ "learning_rate": 0.0002141501483300395,
330
+ "loss": 0.4295254135131836,
331
+ "mean_token_accuracy": 0.8616545403003693,
332
+ "num_tokens": 2072809.0,
333
+ "step": 1450
334
+ },
335
+ {
336
+ "entropy": 0.49857193052768706,
337
+ "epoch": 3.857142857142857,
338
+ "grad_norm": 0.7732954025268555,
339
+ "learning_rate": 0.00020912417559712133,
340
+ "loss": 0.4289303207397461,
341
+ "mean_token_accuracy": 0.8616443765163422,
342
+ "num_tokens": 2142475.0,
343
+ "step": 1500
344
+ },
345
+ {
346
+ "entropy": 0.4746784272789955,
347
+ "epoch": 3.985842985842986,
348
+ "grad_norm": 0.5791187882423401,
349
+ "learning_rate": 0.00020394997040121726,
350
+ "loss": 0.4180263900756836,
351
+ "mean_token_accuracy": 0.866080379486084,
352
+ "num_tokens": 2214044.0,
353
+ "step": 1550
354
+ },
355
+ {
356
+ "epoch": 4.0,
357
+ "eval_entropy": 0.4793926059585257,
358
+ "eval_loss": 0.627627968788147,
359
+ "eval_mean_token_accuracy": 0.8233870095813397,
360
+ "eval_num_tokens": 2221972.0,
361
+ "eval_runtime": 162.0225,
362
+ "eval_samples_per_second": 9.536,
363
+ "eval_steps_per_second": 1.197,
364
+ "step": 1556
365
+ }
366
+ ],
367
+ "logging_steps": 50,
368
+ "max_steps": 3890,
369
+ "num_input_tokens_seen": 0,
370
+ "num_train_epochs": 10,
371
+ "save_steps": 500,
372
+ "stateful_callbacks": {
373
+ "TrainerControl": {
374
+ "args": {
375
+ "should_epoch_stop": false,
376
+ "should_evaluate": false,
377
+ "should_log": false,
378
+ "should_save": true,
379
+ "should_training_stop": false
380
+ },
381
+ "attributes": {}
382
+ }
383
+ },
384
+ "total_flos": 3.723284547240653e+17,
385
+ "train_batch_size": 8,
386
+ "trial_name": null,
387
+ "trial_params": null
388
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/README.md ADDED
@@ -0,0 +1,209 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen3-14B-Base
3
+ library_name: peft
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - base_model:adapter:Qwen/Qwen3-14B-Base
7
+ - lora
8
+ - sft
9
+ - transformers
10
+ - trl
11
+ ---
12
+
13
+ # Model Card for Model ID
14
+
15
+ <!-- Provide a quick summary of what the model is/does. -->
16
+
17
+
18
+
19
+ ## Model Details
20
+
21
+ ### Model Description
22
+
23
+ <!-- Provide a longer summary of what this model is. -->
24
+
25
+
26
+
27
+ - **Developed by:** [More Information Needed]
28
+ - **Funded by [optional]:** [More Information Needed]
29
+ - **Shared by [optional]:** [More Information Needed]
30
+ - **Model type:** [More Information Needed]
31
+ - **Language(s) (NLP):** [More Information Needed]
32
+ - **License:** [More Information Needed]
33
+ - **Finetuned from model [optional]:** [More Information Needed]
34
+
35
+ ### Model Sources [optional]
36
+
37
+ <!-- Provide the basic links for the model. -->
38
+
39
+ - **Repository:** [More Information Needed]
40
+ - **Paper [optional]:** [More Information Needed]
41
+ - **Demo [optional]:** [More Information Needed]
42
+
43
+ ## Uses
44
+
45
+ <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
46
+
47
+ ### Direct Use
48
+
49
+ <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
50
+
51
+ [More Information Needed]
52
+
53
+ ### Downstream Use [optional]
54
+
55
+ <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
56
+
57
+ [More Information Needed]
58
+
59
+ ### Out-of-Scope Use
60
+
61
+ <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
62
+
63
+ [More Information Needed]
64
+
65
+ ## Bias, Risks, and Limitations
66
+
67
+ <!-- This section is meant to convey both technical and sociotechnical limitations. -->
68
+
69
+ [More Information Needed]
70
+
71
+ ### Recommendations
72
+
73
+ <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
74
+
75
+ Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
76
+
77
+ ## How to Get Started with the Model
78
+
79
+ Use the code below to get started with the model.
80
+
81
+ [More Information Needed]
82
+
83
+ ## Training Details
84
+
85
+ ### Training Data
86
+
87
+ <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
88
+
89
+ [More Information Needed]
90
+
91
+ ### Training Procedure
92
+
93
+ <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
94
+
95
+ #### Preprocessing [optional]
96
+
97
+ [More Information Needed]
98
+
99
+
100
+ #### Training Hyperparameters
101
+
102
+ - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
103
+
104
+ #### Speeds, Sizes, Times [optional]
105
+
106
+ <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
107
+
108
+ [More Information Needed]
109
+
110
+ ## Evaluation
111
+
112
+ <!-- This section describes the evaluation protocols and provides the results. -->
113
+
114
+ ### Testing Data, Factors & Metrics
115
+
116
+ #### Testing Data
117
+
118
+ <!-- This should link to a Dataset Card if possible. -->
119
+
120
+ [More Information Needed]
121
+
122
+ #### Factors
123
+
124
+ <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
125
+
126
+ [More Information Needed]
127
+
128
+ #### Metrics
129
+
130
+ <!-- These are the evaluation metrics being used, ideally with a description of why. -->
131
+
132
+ [More Information Needed]
133
+
134
+ ### Results
135
+
136
+ [More Information Needed]
137
+
138
+ #### Summary
139
+
140
+
141
+
142
+ ## Model Examination [optional]
143
+
144
+ <!-- Relevant interpretability work for the model goes here -->
145
+
146
+ [More Information Needed]
147
+
148
+ ## Environmental Impact
149
+
150
+ <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
151
+
152
+ Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
153
+
154
+ - **Hardware Type:** [More Information Needed]
155
+ - **Hours used:** [More Information Needed]
156
+ - **Cloud Provider:** [More Information Needed]
157
+ - **Compute Region:** [More Information Needed]
158
+ - **Carbon Emitted:** [More Information Needed]
159
+
160
+ ## Technical Specifications [optional]
161
+
162
+ ### Model Architecture and Objective
163
+
164
+ [More Information Needed]
165
+
166
+ ### Compute Infrastructure
167
+
168
+ [More Information Needed]
169
+
170
+ #### Hardware
171
+
172
+ [More Information Needed]
173
+
174
+ #### Software
175
+
176
+ [More Information Needed]
177
+
178
+ ## Citation [optional]
179
+
180
+ <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
181
+
182
+ **BibTeX:**
183
+
184
+ [More Information Needed]
185
+
186
+ **APA:**
187
+
188
+ [More Information Needed]
189
+
190
+ ## Glossary [optional]
191
+
192
+ <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
193
+
194
+ [More Information Needed]
195
+
196
+ ## More Information [optional]
197
+
198
+ [More Information Needed]
199
+
200
+ ## Model Card Authors [optional]
201
+
202
+ [More Information Needed]
203
+
204
+ ## Model Card Contact
205
+
206
+ [More Information Needed]
207
+ ### Framework versions
208
+
209
+ - PEFT 0.19.1
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/adapter_config.json ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "Qwen/Qwen3-14B-Base",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "layer_replication": null,
16
+ "layers_pattern": null,
17
+ "layers_to_transform": null,
18
+ "loftq_config": {},
19
+ "lora_alpha": 256,
20
+ "lora_bias": false,
21
+ "lora_dropout": 0.00237968804112545,
22
+ "lora_ga_config": null,
23
+ "megatron_config": null,
24
+ "megatron_core": "megatron.core",
25
+ "modules_to_save": null,
26
+ "peft_type": "LORA",
27
+ "peft_version": "0.19.1",
28
+ "qalora_group_size": 16,
29
+ "r": 128,
30
+ "rank_pattern": {},
31
+ "revision": null,
32
+ "target_modules": [
33
+ "gate_proj",
34
+ "q_proj",
35
+ "up_proj",
36
+ "o_proj",
37
+ "v_proj",
38
+ "k_proj",
39
+ "down_proj"
40
+ ],
41
+ "target_parameters": null,
42
+ "task_type": "CAUSAL_LM",
43
+ "trainable_token_indices": null,
44
+ "use_bdlora": null,
45
+ "use_dora": false,
46
+ "use_qalora": false,
47
+ "use_rslora": false
48
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/chat_template.jinja ADDED
@@ -0,0 +1,85 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- if tools %}
2
+ {{- '<|im_start|>system\n' }}
3
+ {%- if messages[0].role == 'system' %}
4
+ {{- messages[0].content + '\n\n' }}
5
+ {%- endif %}
6
+ {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
7
+ {%- for tool in tools %}
8
+ {{- "\n" }}
9
+ {{- tool | tojson }}
10
+ {%- endfor %}
11
+ {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
12
+ {%- else %}
13
+ {%- if messages[0].role == 'system' %}
14
+ {{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
15
+ {%- endif %}
16
+ {%- endif %}
17
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
18
+ {%- for message in messages[::-1] %}
19
+ {%- set index = (messages|length - 1) - loop.index0 %}
20
+ {%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
21
+ {%- set ns.multi_step_tool = false %}
22
+ {%- set ns.last_query_index = index %}
23
+ {%- endif %}
24
+ {%- endfor %}
25
+ {%- for message in messages %}
26
+ {%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
27
+ {{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
28
+ {%- elif message.role == "assistant" %}
29
+ {%- set content = message.content %}
30
+ {%- set reasoning_content = '' %}
31
+ {%- if message.reasoning_content is defined and message.reasoning_content is not none %}
32
+ {%- set reasoning_content = message.reasoning_content %}
33
+ {%- else %}
34
+ {%- if '</think>' in message.content %}
35
+ {%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
36
+ {%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
37
+ {%- endif %}
38
+ {%- endif %}
39
+ {%- if loop.index0 > ns.last_query_index %}
40
+ {%- if loop.last or (not loop.last and reasoning_content) %}
41
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
42
+ {%- else %}
43
+ {{- '<|im_start|>' + message.role + '\n' + content }}
44
+ {%- endif %}
45
+ {%- else %}
46
+ {{- '<|im_start|>' + message.role + '\n' + content }}
47
+ {%- endif %}
48
+ {%- if message.tool_calls %}
49
+ {%- for tool_call in message.tool_calls %}
50
+ {%- if (loop.first and content) or (not loop.first) %}
51
+ {{- '\n' }}
52
+ {%- endif %}
53
+ {%- if tool_call.function %}
54
+ {%- set tool_call = tool_call.function %}
55
+ {%- endif %}
56
+ {{- '<tool_call>\n{"name": "' }}
57
+ {{- tool_call.name }}
58
+ {{- '", "arguments": ' }}
59
+ {%- if tool_call.arguments is string %}
60
+ {{- tool_call.arguments }}
61
+ {%- else %}
62
+ {{- tool_call.arguments | tojson }}
63
+ {%- endif %}
64
+ {{- '}\n</tool_call>' }}
65
+ {%- endfor %}
66
+ {%- endif %}
67
+ {{- '<|im_end|>\n' }}
68
+ {%- elif message.role == "tool" %}
69
+ {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
70
+ {{- '<|im_start|>user' }}
71
+ {%- endif %}
72
+ {{- '\n<tool_response>\n' }}
73
+ {{- message.content }}
74
+ {{- '\n</tool_response>' }}
75
+ {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
76
+ {{- '<|im_end|>\n' }}
77
+ {%- endif %}
78
+ {%- endif %}
79
+ {%- endfor %}
80
+ {%- if add_generation_prompt %}
81
+ {{- '<|im_start|>assistant\n' }}
82
+ {%- if enable_thinking is defined and enable_thinking is false %}
83
+ {{- '<think>\n\n</think>\n\n' }}
84
+ {%- endif %}
85
+ {%- endif %}
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/tokenizer_config.json ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "backend": "tokenizers",
4
+ "bos_token": null,
5
+ "clean_up_tokenization_spaces": false,
6
+ "eos_token": "<|endoftext|>",
7
+ "errors": "replace",
8
+ "extra_special_tokens": [
9
+ "<|im_start|>",
10
+ "<|im_end|>",
11
+ "<|object_ref_start|>",
12
+ "<|object_ref_end|>",
13
+ "<|box_start|>",
14
+ "<|box_end|>",
15
+ "<|quad_start|>",
16
+ "<|quad_end|>",
17
+ "<|vision_start|>",
18
+ "<|vision_end|>",
19
+ "<|vision_pad|>",
20
+ "<|image_pad|>",
21
+ "<|video_pad|>"
22
+ ],
23
+ "is_local": false,
24
+ "model_max_length": 131072,
25
+ "pad_token": "<|endoftext|>",
26
+ "split_special_tokens": false,
27
+ "tokenizer_class": "Qwen2Tokenizer",
28
+ "unk_token": null
29
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/trainer_state.json ADDED
@@ -0,0 +1,469 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "best_global_step": null,
3
+ "best_metric": null,
4
+ "best_model_checkpoint": null,
5
+ "epoch": 5.0,
6
+ "eval_steps": 500,
7
+ "global_step": 1945,
8
+ "is_hyper_param_search": false,
9
+ "is_local_process_zero": true,
10
+ "is_world_process_zero": true,
11
+ "log_history": [
12
+ {
13
+ "entropy": 1.6110110306739807,
14
+ "epoch": 0.1287001287001287,
15
+ "grad_norm": 0.8611119389533997,
16
+ "learning_rate": 3.4130257962133866e-05,
17
+ "loss": 1.53956298828125,
18
+ "mean_token_accuracy": 0.6652342769503593,
19
+ "num_tokens": 73407.0,
20
+ "step": 50
21
+ },
22
+ {
23
+ "entropy": 0.8662820833921433,
24
+ "epoch": 0.2574002574002574,
25
+ "grad_norm": 0.7405035495758057,
26
+ "learning_rate": 6.895705180104598e-05,
27
+ "loss": 0.7929539489746094,
28
+ "mean_token_accuracy": 0.7786347842216492,
29
+ "num_tokens": 143994.0,
30
+ "step": 100
31
+ },
32
+ {
33
+ "entropy": 0.7912819278240204,
34
+ "epoch": 0.3861003861003861,
35
+ "grad_norm": 0.6105485558509827,
36
+ "learning_rate": 0.00010378384563995809,
37
+ "loss": 0.7236511993408203,
38
+ "mean_token_accuracy": 0.7925467795133591,
39
+ "num_tokens": 216171.0,
40
+ "step": 150
41
+ },
42
+ {
43
+ "entropy": 0.7777709531784057,
44
+ "epoch": 0.5148005148005148,
45
+ "grad_norm": 0.5781793594360352,
46
+ "learning_rate": 0.0001386106394788702,
47
+ "loss": 0.7024919891357422,
48
+ "mean_token_accuracy": 0.7969876372814179,
49
+ "num_tokens": 284702.0,
50
+ "step": 200
51
+ },
52
+ {
53
+ "entropy": 0.7632609683275223,
54
+ "epoch": 0.6435006435006435,
55
+ "grad_norm": 0.6553444862365723,
56
+ "learning_rate": 0.00017343743331778232,
57
+ "loss": 0.7000718688964844,
58
+ "mean_token_accuracy": 0.7994742071628571,
59
+ "num_tokens": 356393.0,
60
+ "step": 250
61
+ },
62
+ {
63
+ "entropy": 0.754165632724762,
64
+ "epoch": 0.7722007722007722,
65
+ "grad_norm": 0.634222149848938,
66
+ "learning_rate": 0.00020826422715669444,
67
+ "loss": 0.6909049987792969,
68
+ "mean_token_accuracy": 0.8016809666156769,
69
+ "num_tokens": 426916.0,
70
+ "step": 300
71
+ },
72
+ {
73
+ "entropy": 0.7530711203813553,
74
+ "epoch": 0.9009009009009009,
75
+ "grad_norm": 0.7020156383514404,
76
+ "learning_rate": 0.00024309102099560653,
77
+ "loss": 0.69554931640625,
78
+ "mean_token_accuracy": 0.8002807641029358,
79
+ "num_tokens": 499603.0,
80
+ "step": 350
81
+ },
82
+ {
83
+ "epoch": 1.0,
84
+ "eval_entropy": 0.5728044164242204,
85
+ "eval_loss": 0.6785851716995239,
86
+ "eval_mean_token_accuracy": 0.798433246993527,
87
+ "eval_num_tokens": 555493.0,
88
+ "eval_runtime": 162.4129,
89
+ "eval_samples_per_second": 9.513,
90
+ "eval_steps_per_second": 1.194,
91
+ "step": 389
92
+ },
93
+ {
94
+ "entropy": 0.7489313946829902,
95
+ "epoch": 1.0283140283140284,
96
+ "grad_norm": 0.7532815933227539,
97
+ "learning_rate": 0.0002709470016827303,
98
+ "loss": 0.6856581878662109,
99
+ "mean_token_accuracy": 0.8023027079273956,
100
+ "num_tokens": 570813.0,
101
+ "step": 400
102
+ },
103
+ {
104
+ "entropy": 0.7238409864902496,
105
+ "epoch": 1.157014157014157,
106
+ "grad_norm": 1.2016215324401855,
107
+ "learning_rate": 0.0002707561443541359,
108
+ "loss": 0.6699818420410156,
109
+ "mean_token_accuracy": 0.8070302194356919,
110
+ "num_tokens": 642956.0,
111
+ "step": 450
112
+ },
113
+ {
114
+ "entropy": 0.7388938587903976,
115
+ "epoch": 1.2857142857142856,
116
+ "grad_norm": 0.7279272079467773,
117
+ "learning_rate": 0.0002702930068622498,
118
+ "loss": 0.6728517150878907,
119
+ "mean_token_accuracy": 0.8049580943584442,
120
+ "num_tokens": 714498.0,
121
+ "step": 500
122
+ },
123
+ {
124
+ "entropy": 0.7284485149383545,
125
+ "epoch": 1.4144144144144144,
126
+ "grad_norm": 0.8563987016677856,
127
+ "learning_rate": 0.0002695585213716931,
128
+ "loss": 0.6657986450195312,
129
+ "mean_token_accuracy": 0.8085588800907135,
130
+ "num_tokens": 785595.0,
131
+ "step": 550
132
+ },
133
+ {
134
+ "entropy": 0.6960959500074386,
135
+ "epoch": 1.5431145431145432,
136
+ "grad_norm": 0.5145474672317505,
137
+ "learning_rate": 0.0002685541661937683,
138
+ "loss": 0.6358638763427734,
139
+ "mean_token_accuracy": 0.8131621342897415,
140
+ "num_tokens": 857551.0,
141
+ "step": 600
142
+ },
143
+ {
144
+ "entropy": 0.7220414417982102,
145
+ "epoch": 1.6718146718146718,
146
+ "grad_norm": 0.9381059408187866,
147
+ "learning_rate": 0.00026728196281103746,
148
+ "loss": 0.6531407928466797,
149
+ "mean_token_accuracy": 0.811619822382927,
150
+ "num_tokens": 926864.0,
151
+ "step": 650
152
+ },
153
+ {
154
+ "entropy": 0.6955459499359131,
155
+ "epoch": 1.8005148005148004,
156
+ "grad_norm": 0.6115108728408813,
157
+ "learning_rate": 0.0002657444718086503,
158
+ "loss": 0.6269588088989257,
159
+ "mean_token_accuracy": 0.8155373805761337,
160
+ "num_tokens": 998824.0,
161
+ "step": 700
162
+ },
163
+ {
164
+ "entropy": 0.6929812705516816,
165
+ "epoch": 1.9292149292149292,
166
+ "grad_norm": 0.749489426612854,
167
+ "learning_rate": 0.0002639447877206115,
168
+ "loss": 0.629054069519043,
169
+ "mean_token_accuracy": 0.8183021235466004,
170
+ "num_tokens": 1069332.0,
171
+ "step": 750
172
+ },
173
+ {
174
+ "epoch": 2.0,
175
+ "eval_entropy": 0.5833694102223387,
176
+ "eval_loss": 0.6328718662261963,
177
+ "eval_mean_token_accuracy": 0.8159044071571114,
178
+ "eval_num_tokens": 1110986.0,
179
+ "eval_runtime": 161.885,
180
+ "eval_samples_per_second": 9.544,
181
+ "eval_steps_per_second": 1.198,
182
+ "step": 778
183
+ },
184
+ {
185
+ "entropy": 0.647241060480927,
186
+ "epoch": 2.056628056628057,
187
+ "grad_norm": 0.7070767879486084,
188
+ "learning_rate": 0.00026188653280135975,
189
+ "loss": 0.5823922348022461,
190
+ "mean_token_accuracy": 0.8265365301960647,
191
+ "num_tokens": 1141195.0,
192
+ "step": 800
193
+ },
194
+ {
195
+ "entropy": 0.5995378407835961,
196
+ "epoch": 2.1853281853281854,
197
+ "grad_norm": 0.8090486526489258,
198
+ "learning_rate": 0.0002595738497351955,
199
+ "loss": 0.5325597763061524,
200
+ "mean_token_accuracy": 0.8369336777925491,
201
+ "num_tokens": 1210708.0,
202
+ "step": 850
203
+ },
204
+ {
205
+ "entropy": 0.6043145382404327,
206
+ "epoch": 2.314028314028314,
207
+ "grad_norm": 0.8279913067817688,
208
+ "learning_rate": 0.00025701139329823054,
209
+ "loss": 0.5414446258544922,
210
+ "mean_token_accuracy": 0.8361396533250809,
211
+ "num_tokens": 1283441.0,
212
+ "step": 900
213
+ },
214
+ {
215
+ "entropy": 0.5953224584460258,
216
+ "epoch": 2.4427284427284426,
217
+ "grad_norm": 0.6075023412704468,
218
+ "learning_rate": 0.00025420432098964183,
219
+ "loss": 0.536654167175293,
220
+ "mean_token_accuracy": 0.8340093129873276,
221
+ "num_tokens": 1356418.0,
222
+ "step": 950
223
+ },
224
+ {
225
+ "entropy": 0.5998479858040809,
226
+ "epoch": 2.571428571428571,
227
+ "grad_norm": 1.0311471223831177,
228
+ "learning_rate": 0.0002511582826510862,
229
+ "loss": 0.5372924423217773,
230
+ "mean_token_accuracy": 0.8366045409440994,
231
+ "num_tokens": 1427797.0,
232
+ "step": 1000
233
+ },
234
+ {
235
+ "entropy": 0.5980158120393753,
236
+ "epoch": 2.7001287001287,
237
+ "grad_norm": 0.5971426367759705,
238
+ "learning_rate": 0.0002478794090951689,
239
+ "loss": 0.5392885208129883,
240
+ "mean_token_accuracy": 0.8347727072238922,
241
+ "num_tokens": 1498082.0,
242
+ "step": 1050
243
+ },
244
+ {
245
+ "entropy": 0.5990680930018425,
246
+ "epoch": 2.828828828828829,
247
+ "grad_norm": 0.5662627220153809,
248
+ "learning_rate": 0.0002443742997658538,
249
+ "loss": 0.5360498428344727,
250
+ "mean_token_accuracy": 0.8371847170591354,
251
+ "num_tokens": 1568798.0,
252
+ "step": 1100
253
+ },
254
+ {
255
+ "entropy": 0.5957039377093315,
256
+ "epoch": 2.9575289575289574,
257
+ "grad_norm": 0.5043798685073853,
258
+ "learning_rate": 0.00024065000945565205,
259
+ "loss": 0.5342231369018555,
260
+ "mean_token_accuracy": 0.8380735236406326,
261
+ "num_tokens": 1643449.0,
262
+ "step": 1150
263
+ },
264
+ {
265
+ "epoch": 3.0,
266
+ "eval_entropy": 0.5346978819861854,
267
+ "eval_loss": 0.6255015134811401,
268
+ "eval_mean_token_accuracy": 0.8202014476368108,
269
+ "eval_num_tokens": 1666479.0,
270
+ "eval_runtime": 161.6098,
271
+ "eval_samples_per_second": 9.56,
272
+ "eval_steps_per_second": 1.2,
273
+ "step": 1167
274
+ },
275
+ {
276
+ "entropy": 0.5212209137401196,
277
+ "epoch": 3.0849420849420848,
278
+ "grad_norm": 0.7095440626144409,
279
+ "learning_rate": 0.00023671403410632178,
280
+ "loss": 0.45311901092529294,
281
+ "mean_token_accuracy": 0.856362871449403,
282
+ "num_tokens": 1713536.0,
283
+ "step": 1200
284
+ },
285
+ {
286
+ "entropy": 0.4701593083143234,
287
+ "epoch": 3.213642213642214,
288
+ "grad_norm": 0.6408083438873291,
289
+ "learning_rate": 0.0002325742957216607,
290
+ "loss": 0.39916397094726563,
291
+ "mean_token_accuracy": 0.8698061722517013,
292
+ "num_tokens": 1785609.0,
293
+ "step": 1250
294
+ },
295
+ {
296
+ "entropy": 0.4766591975092888,
297
+ "epoch": 3.3423423423423424,
298
+ "grad_norm": 0.6415093541145325,
299
+ "learning_rate": 0.0002282391264227552,
300
+ "loss": 0.4116698455810547,
301
+ "mean_token_accuracy": 0.8679435575008392,
302
+ "num_tokens": 1858651.0,
303
+ "step": 1300
304
+ },
305
+ {
306
+ "entropy": 0.4937947469949722,
307
+ "epoch": 3.471042471042471,
308
+ "grad_norm": 0.6549825072288513,
309
+ "learning_rate": 0.00022371725167778054,
310
+ "loss": 0.4296376037597656,
311
+ "mean_token_accuracy": 0.8609357953071595,
312
+ "num_tokens": 1928692.0,
313
+ "step": 1350
314
+ },
315
+ {
316
+ "entropy": 0.4920153194665909,
317
+ "epoch": 3.5997425997425996,
318
+ "grad_norm": 0.6452126502990723,
319
+ "learning_rate": 0.00021901777274010406,
320
+ "loss": 0.4307489013671875,
321
+ "mean_token_accuracy": 0.8606827831268311,
322
+ "num_tokens": 1998668.0,
323
+ "step": 1400
324
+ },
325
+ {
326
+ "entropy": 0.490042342543602,
327
+ "epoch": 3.7284427284427286,
328
+ "grad_norm": 0.5727734565734863,
329
+ "learning_rate": 0.0002141501483300395,
330
+ "loss": 0.4295254135131836,
331
+ "mean_token_accuracy": 0.8616545403003693,
332
+ "num_tokens": 2072809.0,
333
+ "step": 1450
334
+ },
335
+ {
336
+ "entropy": 0.49857193052768706,
337
+ "epoch": 3.857142857142857,
338
+ "grad_norm": 0.7732954025268555,
339
+ "learning_rate": 0.00020912417559712133,
340
+ "loss": 0.4289303207397461,
341
+ "mean_token_accuracy": 0.8616443765163422,
342
+ "num_tokens": 2142475.0,
343
+ "step": 1500
344
+ },
345
+ {
346
+ "entropy": 0.4746784272789955,
347
+ "epoch": 3.985842985842986,
348
+ "grad_norm": 0.5791187882423401,
349
+ "learning_rate": 0.00020394997040121726,
350
+ "loss": 0.4180263900756836,
351
+ "mean_token_accuracy": 0.866080379486084,
352
+ "num_tokens": 2214044.0,
353
+ "step": 1550
354
+ },
355
+ {
356
+ "epoch": 4.0,
357
+ "eval_entropy": 0.4793926059585257,
358
+ "eval_loss": 0.627627968788147,
359
+ "eval_mean_token_accuracy": 0.8233870095813397,
360
+ "eval_num_tokens": 2221972.0,
361
+ "eval_runtime": 162.0225,
362
+ "eval_samples_per_second": 9.536,
363
+ "eval_steps_per_second": 1.197,
364
+ "step": 1556
365
+ },
366
+ {
367
+ "entropy": 0.38195489000792454,
368
+ "epoch": 4.113256113256114,
369
+ "grad_norm": 0.6002617478370667,
370
+ "learning_rate": 0.0001986379469521669,
371
+ "loss": 0.30819049835205076,
372
+ "mean_token_accuracy": 0.8977848634575353,
373
+ "num_tokens": 2282164.0,
374
+ "step": 1600
375
+ },
376
+ {
377
+ "entropy": 0.3655787402391434,
378
+ "epoch": 4.241956241956242,
379
+ "grad_norm": 0.7100041508674622,
380
+ "learning_rate": 0.00019319879684892634,
381
+ "loss": 0.29959835052490236,
382
+ "mean_token_accuracy": 0.8991208010911942,
383
+ "num_tokens": 2353213.0,
384
+ "step": 1650
385
+ },
386
+ {
387
+ "entropy": 0.3821141055226326,
388
+ "epoch": 4.370656370656371,
389
+ "grad_norm": 0.5848307013511658,
390
+ "learning_rate": 0.00018764346756040715,
391
+ "loss": 0.313802490234375,
392
+ "mean_token_accuracy": 0.895167955160141,
393
+ "num_tokens": 2425068.0,
394
+ "step": 1700
395
+ },
396
+ {
397
+ "entropy": 0.37083797007799146,
398
+ "epoch": 4.499356499356499,
399
+ "grad_norm": 0.6447024941444397,
400
+ "learning_rate": 0.00018198314039132143,
401
+ "loss": 0.30583988189697264,
402
+ "mean_token_accuracy": 0.8961733293533325,
403
+ "num_tokens": 2498321.0,
404
+ "step": 1750
405
+ },
406
+ {
407
+ "entropy": 0.3791545969247818,
408
+ "epoch": 4.628056628056628,
409
+ "grad_norm": 0.6575382351875305,
410
+ "learning_rate": 0.00017622920797738184,
411
+ "loss": 0.3088031005859375,
412
+ "mean_token_accuracy": 0.8960050916671753,
413
+ "num_tokens": 2570321.0,
414
+ "step": 1800
415
+ },
416
+ {
417
+ "entropy": 0.3946831756830215,
418
+ "epoch": 4.756756756756757,
419
+ "grad_norm": 0.5351552963256836,
420
+ "learning_rate": 0.00017039325135515207,
421
+ "loss": 0.3229162979125977,
422
+ "mean_token_accuracy": 0.8920552498102188,
423
+ "num_tokens": 2642851.0,
424
+ "step": 1850
425
+ },
426
+ {
427
+ "entropy": 0.37298239797353744,
428
+ "epoch": 4.885456885456885,
429
+ "grad_norm": 0.7624587416648865,
430
+ "learning_rate": 0.00016448701665269964,
431
+ "loss": 0.3067934799194336,
432
+ "mean_token_accuracy": 0.8951873427629471,
433
+ "num_tokens": 2715629.0,
434
+ "step": 1900
435
+ },
436
+ {
437
+ "epoch": 5.0,
438
+ "eval_entropy": 0.3738150204887095,
439
+ "eval_loss": 0.7131896615028381,
440
+ "eval_mean_token_accuracy": 0.8207752468045225,
441
+ "eval_num_tokens": 2777465.0,
442
+ "eval_runtime": 162.1244,
443
+ "eval_samples_per_second": 9.53,
444
+ "eval_steps_per_second": 1.197,
445
+ "step": 1945
446
+ }
447
+ ],
448
+ "logging_steps": 50,
449
+ "max_steps": 3890,
450
+ "num_input_tokens_seen": 0,
451
+ "num_train_epochs": 10,
452
+ "save_steps": 500,
453
+ "stateful_callbacks": {
454
+ "TrainerControl": {
455
+ "args": {
456
+ "should_epoch_stop": false,
457
+ "should_evaluate": false,
458
+ "should_log": false,
459
+ "should_save": true,
460
+ "should_training_stop": false
461
+ },
462
+ "attributes": {}
463
+ }
464
+ },
465
+ "total_flos": 4.653233039031091e+17,
466
+ "train_batch_size": 8,
467
+ "trial_name": null,
468
+ "trial_params": null
469
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/README.md ADDED
@@ -0,0 +1,209 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen3-14B-Base
3
+ library_name: peft
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - base_model:adapter:Qwen/Qwen3-14B-Base
7
+ - lora
8
+ - sft
9
+ - transformers
10
+ - trl
11
+ ---
12
+
13
+ # Model Card for Model ID
14
+
15
+ <!-- Provide a quick summary of what the model is/does. -->
16
+
17
+
18
+
19
+ ## Model Details
20
+
21
+ ### Model Description
22
+
23
+ <!-- Provide a longer summary of what this model is. -->
24
+
25
+
26
+
27
+ - **Developed by:** [More Information Needed]
28
+ - **Funded by [optional]:** [More Information Needed]
29
+ - **Shared by [optional]:** [More Information Needed]
30
+ - **Model type:** [More Information Needed]
31
+ - **Language(s) (NLP):** [More Information Needed]
32
+ - **License:** [More Information Needed]
33
+ - **Finetuned from model [optional]:** [More Information Needed]
34
+
35
+ ### Model Sources [optional]
36
+
37
+ <!-- Provide the basic links for the model. -->
38
+
39
+ - **Repository:** [More Information Needed]
40
+ - **Paper [optional]:** [More Information Needed]
41
+ - **Demo [optional]:** [More Information Needed]
42
+
43
+ ## Uses
44
+
45
+ <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
46
+
47
+ ### Direct Use
48
+
49
+ <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
50
+
51
+ [More Information Needed]
52
+
53
+ ### Downstream Use [optional]
54
+
55
+ <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
56
+
57
+ [More Information Needed]
58
+
59
+ ### Out-of-Scope Use
60
+
61
+ <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
62
+
63
+ [More Information Needed]
64
+
65
+ ## Bias, Risks, and Limitations
66
+
67
+ <!-- This section is meant to convey both technical and sociotechnical limitations. -->
68
+
69
+ [More Information Needed]
70
+
71
+ ### Recommendations
72
+
73
+ <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
74
+
75
+ Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
76
+
77
+ ## How to Get Started with the Model
78
+
79
+ Use the code below to get started with the model.
80
+
81
+ [More Information Needed]
82
+
83
+ ## Training Details
84
+
85
+ ### Training Data
86
+
87
+ <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
88
+
89
+ [More Information Needed]
90
+
91
+ ### Training Procedure
92
+
93
+ <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
94
+
95
+ #### Preprocessing [optional]
96
+
97
+ [More Information Needed]
98
+
99
+
100
+ #### Training Hyperparameters
101
+
102
+ - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
103
+
104
+ #### Speeds, Sizes, Times [optional]
105
+
106
+ <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
107
+
108
+ [More Information Needed]
109
+
110
+ ## Evaluation
111
+
112
+ <!-- This section describes the evaluation protocols and provides the results. -->
113
+
114
+ ### Testing Data, Factors & Metrics
115
+
116
+ #### Testing Data
117
+
118
+ <!-- This should link to a Dataset Card if possible. -->
119
+
120
+ [More Information Needed]
121
+
122
+ #### Factors
123
+
124
+ <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
125
+
126
+ [More Information Needed]
127
+
128
+ #### Metrics
129
+
130
+ <!-- These are the evaluation metrics being used, ideally with a description of why. -->
131
+
132
+ [More Information Needed]
133
+
134
+ ### Results
135
+
136
+ [More Information Needed]
137
+
138
+ #### Summary
139
+
140
+
141
+
142
+ ## Model Examination [optional]
143
+
144
+ <!-- Relevant interpretability work for the model goes here -->
145
+
146
+ [More Information Needed]
147
+
148
+ ## Environmental Impact
149
+
150
+ <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
151
+
152
+ Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
153
+
154
+ - **Hardware Type:** [More Information Needed]
155
+ - **Hours used:** [More Information Needed]
156
+ - **Cloud Provider:** [More Information Needed]
157
+ - **Compute Region:** [More Information Needed]
158
+ - **Carbon Emitted:** [More Information Needed]
159
+
160
+ ## Technical Specifications [optional]
161
+
162
+ ### Model Architecture and Objective
163
+
164
+ [More Information Needed]
165
+
166
+ ### Compute Infrastructure
167
+
168
+ [More Information Needed]
169
+
170
+ #### Hardware
171
+
172
+ [More Information Needed]
173
+
174
+ #### Software
175
+
176
+ [More Information Needed]
177
+
178
+ ## Citation [optional]
179
+
180
+ <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
181
+
182
+ **BibTeX:**
183
+
184
+ [More Information Needed]
185
+
186
+ **APA:**
187
+
188
+ [More Information Needed]
189
+
190
+ ## Glossary [optional]
191
+
192
+ <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
193
+
194
+ [More Information Needed]
195
+
196
+ ## More Information [optional]
197
+
198
+ [More Information Needed]
199
+
200
+ ## Model Card Authors [optional]
201
+
202
+ [More Information Needed]
203
+
204
+ ## Model Card Contact
205
+
206
+ [More Information Needed]
207
+ ### Framework versions
208
+
209
+ - PEFT 0.19.1
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/adapter_config.json ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "Qwen/Qwen3-14B-Base",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "layer_replication": null,
16
+ "layers_pattern": null,
17
+ "layers_to_transform": null,
18
+ "loftq_config": {},
19
+ "lora_alpha": 256,
20
+ "lora_bias": false,
21
+ "lora_dropout": 0.00237968804112545,
22
+ "lora_ga_config": null,
23
+ "megatron_config": null,
24
+ "megatron_core": "megatron.core",
25
+ "modules_to_save": null,
26
+ "peft_type": "LORA",
27
+ "peft_version": "0.19.1",
28
+ "qalora_group_size": 16,
29
+ "r": 128,
30
+ "rank_pattern": {},
31
+ "revision": null,
32
+ "target_modules": [
33
+ "gate_proj",
34
+ "q_proj",
35
+ "up_proj",
36
+ "o_proj",
37
+ "v_proj",
38
+ "k_proj",
39
+ "down_proj"
40
+ ],
41
+ "target_parameters": null,
42
+ "task_type": "CAUSAL_LM",
43
+ "trainable_token_indices": null,
44
+ "use_bdlora": null,
45
+ "use_dora": false,
46
+ "use_qalora": false,
47
+ "use_rslora": false
48
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/chat_template.jinja ADDED
@@ -0,0 +1,85 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- if tools %}
2
+ {{- '<|im_start|>system\n' }}
3
+ {%- if messages[0].role == 'system' %}
4
+ {{- messages[0].content + '\n\n' }}
5
+ {%- endif %}
6
+ {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
7
+ {%- for tool in tools %}
8
+ {{- "\n" }}
9
+ {{- tool | tojson }}
10
+ {%- endfor %}
11
+ {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
12
+ {%- else %}
13
+ {%- if messages[0].role == 'system' %}
14
+ {{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
15
+ {%- endif %}
16
+ {%- endif %}
17
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
18
+ {%- for message in messages[::-1] %}
19
+ {%- set index = (messages|length - 1) - loop.index0 %}
20
+ {%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
21
+ {%- set ns.multi_step_tool = false %}
22
+ {%- set ns.last_query_index = index %}
23
+ {%- endif %}
24
+ {%- endfor %}
25
+ {%- for message in messages %}
26
+ {%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
27
+ {{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
28
+ {%- elif message.role == "assistant" %}
29
+ {%- set content = message.content %}
30
+ {%- set reasoning_content = '' %}
31
+ {%- if message.reasoning_content is defined and message.reasoning_content is not none %}
32
+ {%- set reasoning_content = message.reasoning_content %}
33
+ {%- else %}
34
+ {%- if '</think>' in message.content %}
35
+ {%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
36
+ {%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
37
+ {%- endif %}
38
+ {%- endif %}
39
+ {%- if loop.index0 > ns.last_query_index %}
40
+ {%- if loop.last or (not loop.last and reasoning_content) %}
41
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
42
+ {%- else %}
43
+ {{- '<|im_start|>' + message.role + '\n' + content }}
44
+ {%- endif %}
45
+ {%- else %}
46
+ {{- '<|im_start|>' + message.role + '\n' + content }}
47
+ {%- endif %}
48
+ {%- if message.tool_calls %}
49
+ {%- for tool_call in message.tool_calls %}
50
+ {%- if (loop.first and content) or (not loop.first) %}
51
+ {{- '\n' }}
52
+ {%- endif %}
53
+ {%- if tool_call.function %}
54
+ {%- set tool_call = tool_call.function %}
55
+ {%- endif %}
56
+ {{- '<tool_call>\n{"name": "' }}
57
+ {{- tool_call.name }}
58
+ {{- '", "arguments": ' }}
59
+ {%- if tool_call.arguments is string %}
60
+ {{- tool_call.arguments }}
61
+ {%- else %}
62
+ {{- tool_call.arguments | tojson }}
63
+ {%- endif %}
64
+ {{- '}\n</tool_call>' }}
65
+ {%- endfor %}
66
+ {%- endif %}
67
+ {{- '<|im_end|>\n' }}
68
+ {%- elif message.role == "tool" %}
69
+ {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
70
+ {{- '<|im_start|>user' }}
71
+ {%- endif %}
72
+ {{- '\n<tool_response>\n' }}
73
+ {{- message.content }}
74
+ {{- '\n</tool_response>' }}
75
+ {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
76
+ {{- '<|im_end|>\n' }}
77
+ {%- endif %}
78
+ {%- endif %}
79
+ {%- endfor %}
80
+ {%- if add_generation_prompt %}
81
+ {{- '<|im_start|>assistant\n' }}
82
+ {%- if enable_thinking is defined and enable_thinking is false %}
83
+ {{- '<think>\n\n</think>\n\n' }}
84
+ {%- endif %}
85
+ {%- endif %}
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/tokenizer_config.json ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "backend": "tokenizers",
4
+ "bos_token": null,
5
+ "clean_up_tokenization_spaces": false,
6
+ "eos_token": "<|endoftext|>",
7
+ "errors": "replace",
8
+ "extra_special_tokens": [
9
+ "<|im_start|>",
10
+ "<|im_end|>",
11
+ "<|object_ref_start|>",
12
+ "<|object_ref_end|>",
13
+ "<|box_start|>",
14
+ "<|box_end|>",
15
+ "<|quad_start|>",
16
+ "<|quad_end|>",
17
+ "<|vision_start|>",
18
+ "<|vision_end|>",
19
+ "<|vision_pad|>",
20
+ "<|image_pad|>",
21
+ "<|video_pad|>"
22
+ ],
23
+ "is_local": false,
24
+ "model_max_length": 131072,
25
+ "pad_token": "<|endoftext|>",
26
+ "split_special_tokens": false,
27
+ "tokenizer_class": "Qwen2Tokenizer",
28
+ "unk_token": null
29
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/trainer_state.json ADDED
@@ -0,0 +1,560 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "best_global_step": null,
3
+ "best_metric": null,
4
+ "best_model_checkpoint": null,
5
+ "epoch": 6.0,
6
+ "eval_steps": 500,
7
+ "global_step": 2334,
8
+ "is_hyper_param_search": false,
9
+ "is_local_process_zero": true,
10
+ "is_world_process_zero": true,
11
+ "log_history": [
12
+ {
13
+ "entropy": 1.6110110306739807,
14
+ "epoch": 0.1287001287001287,
15
+ "grad_norm": 0.8611119389533997,
16
+ "learning_rate": 3.4130257962133866e-05,
17
+ "loss": 1.53956298828125,
18
+ "mean_token_accuracy": 0.6652342769503593,
19
+ "num_tokens": 73407.0,
20
+ "step": 50
21
+ },
22
+ {
23
+ "entropy": 0.8662820833921433,
24
+ "epoch": 0.2574002574002574,
25
+ "grad_norm": 0.7405035495758057,
26
+ "learning_rate": 6.895705180104598e-05,
27
+ "loss": 0.7929539489746094,
28
+ "mean_token_accuracy": 0.7786347842216492,
29
+ "num_tokens": 143994.0,
30
+ "step": 100
31
+ },
32
+ {
33
+ "entropy": 0.7912819278240204,
34
+ "epoch": 0.3861003861003861,
35
+ "grad_norm": 0.6105485558509827,
36
+ "learning_rate": 0.00010378384563995809,
37
+ "loss": 0.7236511993408203,
38
+ "mean_token_accuracy": 0.7925467795133591,
39
+ "num_tokens": 216171.0,
40
+ "step": 150
41
+ },
42
+ {
43
+ "entropy": 0.7777709531784057,
44
+ "epoch": 0.5148005148005148,
45
+ "grad_norm": 0.5781793594360352,
46
+ "learning_rate": 0.0001386106394788702,
47
+ "loss": 0.7024919891357422,
48
+ "mean_token_accuracy": 0.7969876372814179,
49
+ "num_tokens": 284702.0,
50
+ "step": 200
51
+ },
52
+ {
53
+ "entropy": 0.7632609683275223,
54
+ "epoch": 0.6435006435006435,
55
+ "grad_norm": 0.6553444862365723,
56
+ "learning_rate": 0.00017343743331778232,
57
+ "loss": 0.7000718688964844,
58
+ "mean_token_accuracy": 0.7994742071628571,
59
+ "num_tokens": 356393.0,
60
+ "step": 250
61
+ },
62
+ {
63
+ "entropy": 0.754165632724762,
64
+ "epoch": 0.7722007722007722,
65
+ "grad_norm": 0.634222149848938,
66
+ "learning_rate": 0.00020826422715669444,
67
+ "loss": 0.6909049987792969,
68
+ "mean_token_accuracy": 0.8016809666156769,
69
+ "num_tokens": 426916.0,
70
+ "step": 300
71
+ },
72
+ {
73
+ "entropy": 0.7530711203813553,
74
+ "epoch": 0.9009009009009009,
75
+ "grad_norm": 0.7020156383514404,
76
+ "learning_rate": 0.00024309102099560653,
77
+ "loss": 0.69554931640625,
78
+ "mean_token_accuracy": 0.8002807641029358,
79
+ "num_tokens": 499603.0,
80
+ "step": 350
81
+ },
82
+ {
83
+ "epoch": 1.0,
84
+ "eval_entropy": 0.5728044164242204,
85
+ "eval_loss": 0.6785851716995239,
86
+ "eval_mean_token_accuracy": 0.798433246993527,
87
+ "eval_num_tokens": 555493.0,
88
+ "eval_runtime": 162.4129,
89
+ "eval_samples_per_second": 9.513,
90
+ "eval_steps_per_second": 1.194,
91
+ "step": 389
92
+ },
93
+ {
94
+ "entropy": 0.7489313946829902,
95
+ "epoch": 1.0283140283140284,
96
+ "grad_norm": 0.7532815933227539,
97
+ "learning_rate": 0.0002709470016827303,
98
+ "loss": 0.6856581878662109,
99
+ "mean_token_accuracy": 0.8023027079273956,
100
+ "num_tokens": 570813.0,
101
+ "step": 400
102
+ },
103
+ {
104
+ "entropy": 0.7238409864902496,
105
+ "epoch": 1.157014157014157,
106
+ "grad_norm": 1.2016215324401855,
107
+ "learning_rate": 0.0002707561443541359,
108
+ "loss": 0.6699818420410156,
109
+ "mean_token_accuracy": 0.8070302194356919,
110
+ "num_tokens": 642956.0,
111
+ "step": 450
112
+ },
113
+ {
114
+ "entropy": 0.7388938587903976,
115
+ "epoch": 1.2857142857142856,
116
+ "grad_norm": 0.7279272079467773,
117
+ "learning_rate": 0.0002702930068622498,
118
+ "loss": 0.6728517150878907,
119
+ "mean_token_accuracy": 0.8049580943584442,
120
+ "num_tokens": 714498.0,
121
+ "step": 500
122
+ },
123
+ {
124
+ "entropy": 0.7284485149383545,
125
+ "epoch": 1.4144144144144144,
126
+ "grad_norm": 0.8563987016677856,
127
+ "learning_rate": 0.0002695585213716931,
128
+ "loss": 0.6657986450195312,
129
+ "mean_token_accuracy": 0.8085588800907135,
130
+ "num_tokens": 785595.0,
131
+ "step": 550
132
+ },
133
+ {
134
+ "entropy": 0.6960959500074386,
135
+ "epoch": 1.5431145431145432,
136
+ "grad_norm": 0.5145474672317505,
137
+ "learning_rate": 0.0002685541661937683,
138
+ "loss": 0.6358638763427734,
139
+ "mean_token_accuracy": 0.8131621342897415,
140
+ "num_tokens": 857551.0,
141
+ "step": 600
142
+ },
143
+ {
144
+ "entropy": 0.7220414417982102,
145
+ "epoch": 1.6718146718146718,
146
+ "grad_norm": 0.9381059408187866,
147
+ "learning_rate": 0.00026728196281103746,
148
+ "loss": 0.6531407928466797,
149
+ "mean_token_accuracy": 0.811619822382927,
150
+ "num_tokens": 926864.0,
151
+ "step": 650
152
+ },
153
+ {
154
+ "entropy": 0.6955459499359131,
155
+ "epoch": 1.8005148005148004,
156
+ "grad_norm": 0.6115108728408813,
157
+ "learning_rate": 0.0002657444718086503,
158
+ "loss": 0.6269588088989257,
159
+ "mean_token_accuracy": 0.8155373805761337,
160
+ "num_tokens": 998824.0,
161
+ "step": 700
162
+ },
163
+ {
164
+ "entropy": 0.6929812705516816,
165
+ "epoch": 1.9292149292149292,
166
+ "grad_norm": 0.749489426612854,
167
+ "learning_rate": 0.0002639447877206115,
168
+ "loss": 0.629054069519043,
169
+ "mean_token_accuracy": 0.8183021235466004,
170
+ "num_tokens": 1069332.0,
171
+ "step": 750
172
+ },
173
+ {
174
+ "epoch": 2.0,
175
+ "eval_entropy": 0.5833694102223387,
176
+ "eval_loss": 0.6328718662261963,
177
+ "eval_mean_token_accuracy": 0.8159044071571114,
178
+ "eval_num_tokens": 1110986.0,
179
+ "eval_runtime": 161.885,
180
+ "eval_samples_per_second": 9.544,
181
+ "eval_steps_per_second": 1.198,
182
+ "step": 778
183
+ },
184
+ {
185
+ "entropy": 0.647241060480927,
186
+ "epoch": 2.056628056628057,
187
+ "grad_norm": 0.7070767879486084,
188
+ "learning_rate": 0.00026188653280135975,
189
+ "loss": 0.5823922348022461,
190
+ "mean_token_accuracy": 0.8265365301960647,
191
+ "num_tokens": 1141195.0,
192
+ "step": 800
193
+ },
194
+ {
195
+ "entropy": 0.5995378407835961,
196
+ "epoch": 2.1853281853281854,
197
+ "grad_norm": 0.8090486526489258,
198
+ "learning_rate": 0.0002595738497351955,
199
+ "loss": 0.5325597763061524,
200
+ "mean_token_accuracy": 0.8369336777925491,
201
+ "num_tokens": 1210708.0,
202
+ "step": 850
203
+ },
204
+ {
205
+ "entropy": 0.6043145382404327,
206
+ "epoch": 2.314028314028314,
207
+ "grad_norm": 0.8279913067817688,
208
+ "learning_rate": 0.00025701139329823054,
209
+ "loss": 0.5414446258544922,
210
+ "mean_token_accuracy": 0.8361396533250809,
211
+ "num_tokens": 1283441.0,
212
+ "step": 900
213
+ },
214
+ {
215
+ "entropy": 0.5953224584460258,
216
+ "epoch": 2.4427284427284426,
217
+ "grad_norm": 0.6075023412704468,
218
+ "learning_rate": 0.00025420432098964183,
219
+ "loss": 0.536654167175293,
220
+ "mean_token_accuracy": 0.8340093129873276,
221
+ "num_tokens": 1356418.0,
222
+ "step": 950
223
+ },
224
+ {
225
+ "entropy": 0.5998479858040809,
226
+ "epoch": 2.571428571428571,
227
+ "grad_norm": 1.0311471223831177,
228
+ "learning_rate": 0.0002511582826510862,
229
+ "loss": 0.5372924423217773,
230
+ "mean_token_accuracy": 0.8366045409440994,
231
+ "num_tokens": 1427797.0,
232
+ "step": 1000
233
+ },
234
+ {
235
+ "entropy": 0.5980158120393753,
236
+ "epoch": 2.7001287001287,
237
+ "grad_norm": 0.5971426367759705,
238
+ "learning_rate": 0.0002478794090951689,
239
+ "loss": 0.5392885208129883,
240
+ "mean_token_accuracy": 0.8347727072238922,
241
+ "num_tokens": 1498082.0,
242
+ "step": 1050
243
+ },
244
+ {
245
+ "entropy": 0.5990680930018425,
246
+ "epoch": 2.828828828828829,
247
+ "grad_norm": 0.5662627220153809,
248
+ "learning_rate": 0.0002443742997658538,
249
+ "loss": 0.5360498428344727,
250
+ "mean_token_accuracy": 0.8371847170591354,
251
+ "num_tokens": 1568798.0,
252
+ "step": 1100
253
+ },
254
+ {
255
+ "entropy": 0.5957039377093315,
256
+ "epoch": 2.9575289575289574,
257
+ "grad_norm": 0.5043798685073853,
258
+ "learning_rate": 0.00024065000945565205,
259
+ "loss": 0.5342231369018555,
260
+ "mean_token_accuracy": 0.8380735236406326,
261
+ "num_tokens": 1643449.0,
262
+ "step": 1150
263
+ },
264
+ {
265
+ "epoch": 3.0,
266
+ "eval_entropy": 0.5346978819861854,
267
+ "eval_loss": 0.6255015134811401,
268
+ "eval_mean_token_accuracy": 0.8202014476368108,
269
+ "eval_num_tokens": 1666479.0,
270
+ "eval_runtime": 161.6098,
271
+ "eval_samples_per_second": 9.56,
272
+ "eval_steps_per_second": 1.2,
273
+ "step": 1167
274
+ },
275
+ {
276
+ "entropy": 0.5212209137401196,
277
+ "epoch": 3.0849420849420848,
278
+ "grad_norm": 0.7095440626144409,
279
+ "learning_rate": 0.00023671403410632178,
280
+ "loss": 0.45311901092529294,
281
+ "mean_token_accuracy": 0.856362871449403,
282
+ "num_tokens": 1713536.0,
283
+ "step": 1200
284
+ },
285
+ {
286
+ "entropy": 0.4701593083143234,
287
+ "epoch": 3.213642213642214,
288
+ "grad_norm": 0.6408083438873291,
289
+ "learning_rate": 0.0002325742957216607,
290
+ "loss": 0.39916397094726563,
291
+ "mean_token_accuracy": 0.8698061722517013,
292
+ "num_tokens": 1785609.0,
293
+ "step": 1250
294
+ },
295
+ {
296
+ "entropy": 0.4766591975092888,
297
+ "epoch": 3.3423423423423424,
298
+ "grad_norm": 0.6415093541145325,
299
+ "learning_rate": 0.0002282391264227552,
300
+ "loss": 0.4116698455810547,
301
+ "mean_token_accuracy": 0.8679435575008392,
302
+ "num_tokens": 1858651.0,
303
+ "step": 1300
304
+ },
305
+ {
306
+ "entropy": 0.4937947469949722,
307
+ "epoch": 3.471042471042471,
308
+ "grad_norm": 0.6549825072288513,
309
+ "learning_rate": 0.00022371725167778054,
310
+ "loss": 0.4296376037597656,
311
+ "mean_token_accuracy": 0.8609357953071595,
312
+ "num_tokens": 1928692.0,
313
+ "step": 1350
314
+ },
315
+ {
316
+ "entropy": 0.4920153194665909,
317
+ "epoch": 3.5997425997425996,
318
+ "grad_norm": 0.6452126502990723,
319
+ "learning_rate": 0.00021901777274010406,
320
+ "loss": 0.4307489013671875,
321
+ "mean_token_accuracy": 0.8606827831268311,
322
+ "num_tokens": 1998668.0,
323
+ "step": 1400
324
+ },
325
+ {
326
+ "entropy": 0.490042342543602,
327
+ "epoch": 3.7284427284427286,
328
+ "grad_norm": 0.5727734565734863,
329
+ "learning_rate": 0.0002141501483300395,
330
+ "loss": 0.4295254135131836,
331
+ "mean_token_accuracy": 0.8616545403003693,
332
+ "num_tokens": 2072809.0,
333
+ "step": 1450
334
+ },
335
+ {
336
+ "entropy": 0.49857193052768706,
337
+ "epoch": 3.857142857142857,
338
+ "grad_norm": 0.7732954025268555,
339
+ "learning_rate": 0.00020912417559712133,
340
+ "loss": 0.4289303207397461,
341
+ "mean_token_accuracy": 0.8616443765163422,
342
+ "num_tokens": 2142475.0,
343
+ "step": 1500
344
+ },
345
+ {
346
+ "entropy": 0.4746784272789955,
347
+ "epoch": 3.985842985842986,
348
+ "grad_norm": 0.5791187882423401,
349
+ "learning_rate": 0.00020394997040121726,
350
+ "loss": 0.4180263900756836,
351
+ "mean_token_accuracy": 0.866080379486084,
352
+ "num_tokens": 2214044.0,
353
+ "step": 1550
354
+ },
355
+ {
356
+ "epoch": 4.0,
357
+ "eval_entropy": 0.4793926059585257,
358
+ "eval_loss": 0.627627968788147,
359
+ "eval_mean_token_accuracy": 0.8233870095813397,
360
+ "eval_num_tokens": 2221972.0,
361
+ "eval_runtime": 162.0225,
362
+ "eval_samples_per_second": 9.536,
363
+ "eval_steps_per_second": 1.197,
364
+ "step": 1556
365
+ },
366
+ {
367
+ "entropy": 0.38195489000792454,
368
+ "epoch": 4.113256113256114,
369
+ "grad_norm": 0.6002617478370667,
370
+ "learning_rate": 0.0001986379469521669,
371
+ "loss": 0.30819049835205076,
372
+ "mean_token_accuracy": 0.8977848634575353,
373
+ "num_tokens": 2282164.0,
374
+ "step": 1600
375
+ },
376
+ {
377
+ "entropy": 0.3655787402391434,
378
+ "epoch": 4.241956241956242,
379
+ "grad_norm": 0.7100041508674622,
380
+ "learning_rate": 0.00019319879684892634,
381
+ "loss": 0.29959835052490236,
382
+ "mean_token_accuracy": 0.8991208010911942,
383
+ "num_tokens": 2353213.0,
384
+ "step": 1650
385
+ },
386
+ {
387
+ "entropy": 0.3821141055226326,
388
+ "epoch": 4.370656370656371,
389
+ "grad_norm": 0.5848307013511658,
390
+ "learning_rate": 0.00018764346756040715,
391
+ "loss": 0.313802490234375,
392
+ "mean_token_accuracy": 0.895167955160141,
393
+ "num_tokens": 2425068.0,
394
+ "step": 1700
395
+ },
396
+ {
397
+ "entropy": 0.37083797007799146,
398
+ "epoch": 4.499356499356499,
399
+ "grad_norm": 0.6447024941444397,
400
+ "learning_rate": 0.00018198314039132143,
401
+ "loss": 0.30583988189697264,
402
+ "mean_token_accuracy": 0.8961733293533325,
403
+ "num_tokens": 2498321.0,
404
+ "step": 1750
405
+ },
406
+ {
407
+ "entropy": 0.3791545969247818,
408
+ "epoch": 4.628056628056628,
409
+ "grad_norm": 0.6575382351875305,
410
+ "learning_rate": 0.00017622920797738184,
411
+ "loss": 0.3088031005859375,
412
+ "mean_token_accuracy": 0.8960050916671753,
413
+ "num_tokens": 2570321.0,
414
+ "step": 1800
415
+ },
416
+ {
417
+ "entropy": 0.3946831756830215,
418
+ "epoch": 4.756756756756757,
419
+ "grad_norm": 0.5351552963256836,
420
+ "learning_rate": 0.00017039325135515207,
421
+ "loss": 0.3229162979125977,
422
+ "mean_token_accuracy": 0.8920552498102188,
423
+ "num_tokens": 2642851.0,
424
+ "step": 1850
425
+ },
426
+ {
427
+ "entropy": 0.37298239797353744,
428
+ "epoch": 4.885456885456885,
429
+ "grad_norm": 0.7624587416648865,
430
+ "learning_rate": 0.00016448701665269964,
431
+ "loss": 0.3067934799194336,
432
+ "mean_token_accuracy": 0.8951873427629471,
433
+ "num_tokens": 2715629.0,
434
+ "step": 1900
435
+ },
436
+ {
437
+ "epoch": 5.0,
438
+ "eval_entropy": 0.3738150204887095,
439
+ "eval_loss": 0.7131896615028381,
440
+ "eval_mean_token_accuracy": 0.8207752468045225,
441
+ "eval_num_tokens": 2777465.0,
442
+ "eval_runtime": 162.1244,
443
+ "eval_samples_per_second": 9.53,
444
+ "eval_steps_per_second": 1.197,
445
+ "step": 1945
446
+ },
447
+ {
448
+ "entropy": 0.377820266617669,
449
+ "epoch": 5.012870012870013,
450
+ "grad_norm": 0.41990190744400024,
451
+ "learning_rate": 0.00015852239144796624,
452
+ "loss": 0.3058685111999512,
453
+ "mean_token_accuracy": 0.8964343480389527,
454
+ "num_tokens": 2784896.0,
455
+ "step": 1950
456
+ },
457
+ {
458
+ "entropy": 0.2714502356946468,
459
+ "epoch": 5.141570141570142,
460
+ "grad_norm": 0.422568678855896,
461
+ "learning_rate": 0.00015251138084243995,
462
+ "loss": 0.2093442153930664,
463
+ "mean_token_accuracy": 0.9311346983909607,
464
+ "num_tokens": 2854374.0,
465
+ "step": 2000
466
+ },
467
+ {
468
+ "entropy": 0.268475965410471,
469
+ "epoch": 5.27027027027027,
470
+ "grad_norm": 0.6637414693832397,
471
+ "learning_rate": 0.0001464660832982852,
472
+ "loss": 0.20736080169677734,
473
+ "mean_token_accuracy": 0.9289199805259705,
474
+ "num_tokens": 2927362.0,
475
+ "step": 2050
476
+ },
477
+ {
478
+ "entropy": 0.2644876340031624,
479
+ "epoch": 5.398970398970399,
480
+ "grad_norm": 0.47317707538604736,
481
+ "learning_rate": 0.00014039866628756467,
482
+ "loss": 0.20464908599853515,
483
+ "mean_token_accuracy": 0.9300856202840805,
484
+ "num_tokens": 3000143.0,
485
+ "step": 2100
486
+ },
487
+ {
488
+ "entropy": 0.2675253136456013,
489
+ "epoch": 5.527670527670527,
490
+ "grad_norm": 0.5253982543945312,
491
+ "learning_rate": 0.00013432134180256338,
492
+ "loss": 0.21154335021972656,
493
+ "mean_token_accuracy": 0.9283734840154648,
494
+ "num_tokens": 3072561.0,
495
+ "step": 2150
496
+ },
497
+ {
498
+ "entropy": 0.27213907435536383,
499
+ "epoch": 5.656370656370656,
500
+ "grad_norm": 0.46738553047180176,
501
+ "learning_rate": 0.00012824634177650664,
502
+ "loss": 0.21339216232299804,
503
+ "mean_token_accuracy": 0.9272083270549775,
504
+ "num_tokens": 3144831.0,
505
+ "step": 2200
506
+ },
507
+ {
508
+ "entropy": 0.2785488124191761,
509
+ "epoch": 5.785070785070785,
510
+ "grad_norm": 0.4469502866268158,
511
+ "learning_rate": 0.00012218589346414205,
512
+ "loss": 0.21601097106933595,
513
+ "mean_token_accuracy": 0.9255663657188415,
514
+ "num_tokens": 3215960.0,
515
+ "step": 2250
516
+ },
517
+ {
518
+ "entropy": 0.2699935150146484,
519
+ "epoch": 5.913770913770914,
520
+ "grad_norm": 0.7359778881072998,
521
+ "learning_rate": 0.00011615219483173828,
522
+ "loss": 0.20725584030151367,
523
+ "mean_token_accuracy": 0.9286630594730377,
524
+ "num_tokens": 3287499.0,
525
+ "step": 2300
526
+ },
527
+ {
528
+ "epoch": 6.0,
529
+ "eval_entropy": 0.26417383682174783,
530
+ "eval_loss": 0.8880229592323303,
531
+ "eval_mean_token_accuracy": 0.8159987201395723,
532
+ "eval_num_tokens": 3332958.0,
533
+ "eval_runtime": 162.0991,
534
+ "eval_samples_per_second": 9.531,
535
+ "eval_steps_per_second": 1.197,
536
+ "step": 2334
537
+ }
538
+ ],
539
+ "logging_steps": 50,
540
+ "max_steps": 3890,
541
+ "num_input_tokens_seen": 0,
542
+ "num_train_epochs": 10,
543
+ "save_steps": 500,
544
+ "stateful_callbacks": {
545
+ "TrainerControl": {
546
+ "args": {
547
+ "should_epoch_stop": false,
548
+ "should_evaluate": false,
549
+ "should_log": false,
550
+ "should_save": true,
551
+ "should_training_stop": false
552
+ },
553
+ "attributes": {}
554
+ }
555
+ },
556
+ "total_flos": 5.587061113467187e+17,
557
+ "train_batch_size": 8,
558
+ "trial_name": null,
559
+ "trial_params": null
560
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/README.md ADDED
@@ -0,0 +1,209 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen3-14B-Base
3
+ library_name: peft
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - base_model:adapter:Qwen/Qwen3-14B-Base
7
+ - lora
8
+ - sft
9
+ - transformers
10
+ - trl
11
+ ---
12
+
13
+ # Model Card for Model ID
14
+
15
+ <!-- Provide a quick summary of what the model is/does. -->
16
+
17
+
18
+
19
+ ## Model Details
20
+
21
+ ### Model Description
22
+
23
+ <!-- Provide a longer summary of what this model is. -->
24
+
25
+
26
+
27
+ - **Developed by:** [More Information Needed]
28
+ - **Funded by [optional]:** [More Information Needed]
29
+ - **Shared by [optional]:** [More Information Needed]
30
+ - **Model type:** [More Information Needed]
31
+ - **Language(s) (NLP):** [More Information Needed]
32
+ - **License:** [More Information Needed]
33
+ - **Finetuned from model [optional]:** [More Information Needed]
34
+
35
+ ### Model Sources [optional]
36
+
37
+ <!-- Provide the basic links for the model. -->
38
+
39
+ - **Repository:** [More Information Needed]
40
+ - **Paper [optional]:** [More Information Needed]
41
+ - **Demo [optional]:** [More Information Needed]
42
+
43
+ ## Uses
44
+
45
+ <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
46
+
47
+ ### Direct Use
48
+
49
+ <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
50
+
51
+ [More Information Needed]
52
+
53
+ ### Downstream Use [optional]
54
+
55
+ <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
56
+
57
+ [More Information Needed]
58
+
59
+ ### Out-of-Scope Use
60
+
61
+ <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
62
+
63
+ [More Information Needed]
64
+
65
+ ## Bias, Risks, and Limitations
66
+
67
+ <!-- This section is meant to convey both technical and sociotechnical limitations. -->
68
+
69
+ [More Information Needed]
70
+
71
+ ### Recommendations
72
+
73
+ <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
74
+
75
+ Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
76
+
77
+ ## How to Get Started with the Model
78
+
79
+ Use the code below to get started with the model.
80
+
81
+ [More Information Needed]
82
+
83
+ ## Training Details
84
+
85
+ ### Training Data
86
+
87
+ <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
88
+
89
+ [More Information Needed]
90
+
91
+ ### Training Procedure
92
+
93
+ <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
94
+
95
+ #### Preprocessing [optional]
96
+
97
+ [More Information Needed]
98
+
99
+
100
+ #### Training Hyperparameters
101
+
102
+ - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
103
+
104
+ #### Speeds, Sizes, Times [optional]
105
+
106
+ <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
107
+
108
+ [More Information Needed]
109
+
110
+ ## Evaluation
111
+
112
+ <!-- This section describes the evaluation protocols and provides the results. -->
113
+
114
+ ### Testing Data, Factors & Metrics
115
+
116
+ #### Testing Data
117
+
118
+ <!-- This should link to a Dataset Card if possible. -->
119
+
120
+ [More Information Needed]
121
+
122
+ #### Factors
123
+
124
+ <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
125
+
126
+ [More Information Needed]
127
+
128
+ #### Metrics
129
+
130
+ <!-- These are the evaluation metrics being used, ideally with a description of why. -->
131
+
132
+ [More Information Needed]
133
+
134
+ ### Results
135
+
136
+ [More Information Needed]
137
+
138
+ #### Summary
139
+
140
+
141
+
142
+ ## Model Examination [optional]
143
+
144
+ <!-- Relevant interpretability work for the model goes here -->
145
+
146
+ [More Information Needed]
147
+
148
+ ## Environmental Impact
149
+
150
+ <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
151
+
152
+ Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
153
+
154
+ - **Hardware Type:** [More Information Needed]
155
+ - **Hours used:** [More Information Needed]
156
+ - **Cloud Provider:** [More Information Needed]
157
+ - **Compute Region:** [More Information Needed]
158
+ - **Carbon Emitted:** [More Information Needed]
159
+
160
+ ## Technical Specifications [optional]
161
+
162
+ ### Model Architecture and Objective
163
+
164
+ [More Information Needed]
165
+
166
+ ### Compute Infrastructure
167
+
168
+ [More Information Needed]
169
+
170
+ #### Hardware
171
+
172
+ [More Information Needed]
173
+
174
+ #### Software
175
+
176
+ [More Information Needed]
177
+
178
+ ## Citation [optional]
179
+
180
+ <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
181
+
182
+ **BibTeX:**
183
+
184
+ [More Information Needed]
185
+
186
+ **APA:**
187
+
188
+ [More Information Needed]
189
+
190
+ ## Glossary [optional]
191
+
192
+ <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
193
+
194
+ [More Information Needed]
195
+
196
+ ## More Information [optional]
197
+
198
+ [More Information Needed]
199
+
200
+ ## Model Card Authors [optional]
201
+
202
+ [More Information Needed]
203
+
204
+ ## Model Card Contact
205
+
206
+ [More Information Needed]
207
+ ### Framework versions
208
+
209
+ - PEFT 0.19.1
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/adapter_config.json ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "Qwen/Qwen3-14B-Base",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "layer_replication": null,
16
+ "layers_pattern": null,
17
+ "layers_to_transform": null,
18
+ "loftq_config": {},
19
+ "lora_alpha": 256,
20
+ "lora_bias": false,
21
+ "lora_dropout": 0.00237968804112545,
22
+ "lora_ga_config": null,
23
+ "megatron_config": null,
24
+ "megatron_core": "megatron.core",
25
+ "modules_to_save": null,
26
+ "peft_type": "LORA",
27
+ "peft_version": "0.19.1",
28
+ "qalora_group_size": 16,
29
+ "r": 128,
30
+ "rank_pattern": {},
31
+ "revision": null,
32
+ "target_modules": [
33
+ "gate_proj",
34
+ "q_proj",
35
+ "up_proj",
36
+ "o_proj",
37
+ "v_proj",
38
+ "k_proj",
39
+ "down_proj"
40
+ ],
41
+ "target_parameters": null,
42
+ "task_type": "CAUSAL_LM",
43
+ "trainable_token_indices": null,
44
+ "use_bdlora": null,
45
+ "use_dora": false,
46
+ "use_qalora": false,
47
+ "use_rslora": false
48
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/chat_template.jinja ADDED
@@ -0,0 +1,85 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- if tools %}
2
+ {{- '<|im_start|>system\n' }}
3
+ {%- if messages[0].role == 'system' %}
4
+ {{- messages[0].content + '\n\n' }}
5
+ {%- endif %}
6
+ {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
7
+ {%- for tool in tools %}
8
+ {{- "\n" }}
9
+ {{- tool | tojson }}
10
+ {%- endfor %}
11
+ {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
12
+ {%- else %}
13
+ {%- if messages[0].role == 'system' %}
14
+ {{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
15
+ {%- endif %}
16
+ {%- endif %}
17
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
18
+ {%- for message in messages[::-1] %}
19
+ {%- set index = (messages|length - 1) - loop.index0 %}
20
+ {%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
21
+ {%- set ns.multi_step_tool = false %}
22
+ {%- set ns.last_query_index = index %}
23
+ {%- endif %}
24
+ {%- endfor %}
25
+ {%- for message in messages %}
26
+ {%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
27
+ {{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
28
+ {%- elif message.role == "assistant" %}
29
+ {%- set content = message.content %}
30
+ {%- set reasoning_content = '' %}
31
+ {%- if message.reasoning_content is defined and message.reasoning_content is not none %}
32
+ {%- set reasoning_content = message.reasoning_content %}
33
+ {%- else %}
34
+ {%- if '</think>' in message.content %}
35
+ {%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
36
+ {%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
37
+ {%- endif %}
38
+ {%- endif %}
39
+ {%- if loop.index0 > ns.last_query_index %}
40
+ {%- if loop.last or (not loop.last and reasoning_content) %}
41
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
42
+ {%- else %}
43
+ {{- '<|im_start|>' + message.role + '\n' + content }}
44
+ {%- endif %}
45
+ {%- else %}
46
+ {{- '<|im_start|>' + message.role + '\n' + content }}
47
+ {%- endif %}
48
+ {%- if message.tool_calls %}
49
+ {%- for tool_call in message.tool_calls %}
50
+ {%- if (loop.first and content) or (not loop.first) %}
51
+ {{- '\n' }}
52
+ {%- endif %}
53
+ {%- if tool_call.function %}
54
+ {%- set tool_call = tool_call.function %}
55
+ {%- endif %}
56
+ {{- '<tool_call>\n{"name": "' }}
57
+ {{- tool_call.name }}
58
+ {{- '", "arguments": ' }}
59
+ {%- if tool_call.arguments is string %}
60
+ {{- tool_call.arguments }}
61
+ {%- else %}
62
+ {{- tool_call.arguments | tojson }}
63
+ {%- endif %}
64
+ {{- '}\n</tool_call>' }}
65
+ {%- endfor %}
66
+ {%- endif %}
67
+ {{- '<|im_end|>\n' }}
68
+ {%- elif message.role == "tool" %}
69
+ {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
70
+ {{- '<|im_start|>user' }}
71
+ {%- endif %}
72
+ {{- '\n<tool_response>\n' }}
73
+ {{- message.content }}
74
+ {{- '\n</tool_response>' }}
75
+ {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
76
+ {{- '<|im_end|>\n' }}
77
+ {%- endif %}
78
+ {%- endif %}
79
+ {%- endfor %}
80
+ {%- if add_generation_prompt %}
81
+ {{- '<|im_start|>assistant\n' }}
82
+ {%- if enable_thinking is defined and enable_thinking is false %}
83
+ {{- '<think>\n\n</think>\n\n' }}
84
+ {%- endif %}
85
+ {%- endif %}
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/tokenizer_config.json ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "backend": "tokenizers",
4
+ "bos_token": null,
5
+ "clean_up_tokenization_spaces": false,
6
+ "eos_token": "<|endoftext|>",
7
+ "errors": "replace",
8
+ "extra_special_tokens": [
9
+ "<|im_start|>",
10
+ "<|im_end|>",
11
+ "<|object_ref_start|>",
12
+ "<|object_ref_end|>",
13
+ "<|box_start|>",
14
+ "<|box_end|>",
15
+ "<|quad_start|>",
16
+ "<|quad_end|>",
17
+ "<|vision_start|>",
18
+ "<|vision_end|>",
19
+ "<|vision_pad|>",
20
+ "<|image_pad|>",
21
+ "<|video_pad|>"
22
+ ],
23
+ "is_local": false,
24
+ "model_max_length": 131072,
25
+ "pad_token": "<|endoftext|>",
26
+ "split_special_tokens": false,
27
+ "tokenizer_class": "Qwen2Tokenizer",
28
+ "unk_token": null
29
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/trainer_state.json ADDED
@@ -0,0 +1,651 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "best_global_step": null,
3
+ "best_metric": null,
4
+ "best_model_checkpoint": null,
5
+ "epoch": 7.0,
6
+ "eval_steps": 500,
7
+ "global_step": 2723,
8
+ "is_hyper_param_search": false,
9
+ "is_local_process_zero": true,
10
+ "is_world_process_zero": true,
11
+ "log_history": [
12
+ {
13
+ "entropy": 1.6110110306739807,
14
+ "epoch": 0.1287001287001287,
15
+ "grad_norm": 0.8611119389533997,
16
+ "learning_rate": 3.4130257962133866e-05,
17
+ "loss": 1.53956298828125,
18
+ "mean_token_accuracy": 0.6652342769503593,
19
+ "num_tokens": 73407.0,
20
+ "step": 50
21
+ },
22
+ {
23
+ "entropy": 0.8662820833921433,
24
+ "epoch": 0.2574002574002574,
25
+ "grad_norm": 0.7405035495758057,
26
+ "learning_rate": 6.895705180104598e-05,
27
+ "loss": 0.7929539489746094,
28
+ "mean_token_accuracy": 0.7786347842216492,
29
+ "num_tokens": 143994.0,
30
+ "step": 100
31
+ },
32
+ {
33
+ "entropy": 0.7912819278240204,
34
+ "epoch": 0.3861003861003861,
35
+ "grad_norm": 0.6105485558509827,
36
+ "learning_rate": 0.00010378384563995809,
37
+ "loss": 0.7236511993408203,
38
+ "mean_token_accuracy": 0.7925467795133591,
39
+ "num_tokens": 216171.0,
40
+ "step": 150
41
+ },
42
+ {
43
+ "entropy": 0.7777709531784057,
44
+ "epoch": 0.5148005148005148,
45
+ "grad_norm": 0.5781793594360352,
46
+ "learning_rate": 0.0001386106394788702,
47
+ "loss": 0.7024919891357422,
48
+ "mean_token_accuracy": 0.7969876372814179,
49
+ "num_tokens": 284702.0,
50
+ "step": 200
51
+ },
52
+ {
53
+ "entropy": 0.7632609683275223,
54
+ "epoch": 0.6435006435006435,
55
+ "grad_norm": 0.6553444862365723,
56
+ "learning_rate": 0.00017343743331778232,
57
+ "loss": 0.7000718688964844,
58
+ "mean_token_accuracy": 0.7994742071628571,
59
+ "num_tokens": 356393.0,
60
+ "step": 250
61
+ },
62
+ {
63
+ "entropy": 0.754165632724762,
64
+ "epoch": 0.7722007722007722,
65
+ "grad_norm": 0.634222149848938,
66
+ "learning_rate": 0.00020826422715669444,
67
+ "loss": 0.6909049987792969,
68
+ "mean_token_accuracy": 0.8016809666156769,
69
+ "num_tokens": 426916.0,
70
+ "step": 300
71
+ },
72
+ {
73
+ "entropy": 0.7530711203813553,
74
+ "epoch": 0.9009009009009009,
75
+ "grad_norm": 0.7020156383514404,
76
+ "learning_rate": 0.00024309102099560653,
77
+ "loss": 0.69554931640625,
78
+ "mean_token_accuracy": 0.8002807641029358,
79
+ "num_tokens": 499603.0,
80
+ "step": 350
81
+ },
82
+ {
83
+ "epoch": 1.0,
84
+ "eval_entropy": 0.5728044164242204,
85
+ "eval_loss": 0.6785851716995239,
86
+ "eval_mean_token_accuracy": 0.798433246993527,
87
+ "eval_num_tokens": 555493.0,
88
+ "eval_runtime": 162.4129,
89
+ "eval_samples_per_second": 9.513,
90
+ "eval_steps_per_second": 1.194,
91
+ "step": 389
92
+ },
93
+ {
94
+ "entropy": 0.7489313946829902,
95
+ "epoch": 1.0283140283140284,
96
+ "grad_norm": 0.7532815933227539,
97
+ "learning_rate": 0.0002709470016827303,
98
+ "loss": 0.6856581878662109,
99
+ "mean_token_accuracy": 0.8023027079273956,
100
+ "num_tokens": 570813.0,
101
+ "step": 400
102
+ },
103
+ {
104
+ "entropy": 0.7238409864902496,
105
+ "epoch": 1.157014157014157,
106
+ "grad_norm": 1.2016215324401855,
107
+ "learning_rate": 0.0002707561443541359,
108
+ "loss": 0.6699818420410156,
109
+ "mean_token_accuracy": 0.8070302194356919,
110
+ "num_tokens": 642956.0,
111
+ "step": 450
112
+ },
113
+ {
114
+ "entropy": 0.7388938587903976,
115
+ "epoch": 1.2857142857142856,
116
+ "grad_norm": 0.7279272079467773,
117
+ "learning_rate": 0.0002702930068622498,
118
+ "loss": 0.6728517150878907,
119
+ "mean_token_accuracy": 0.8049580943584442,
120
+ "num_tokens": 714498.0,
121
+ "step": 500
122
+ },
123
+ {
124
+ "entropy": 0.7284485149383545,
125
+ "epoch": 1.4144144144144144,
126
+ "grad_norm": 0.8563987016677856,
127
+ "learning_rate": 0.0002695585213716931,
128
+ "loss": 0.6657986450195312,
129
+ "mean_token_accuracy": 0.8085588800907135,
130
+ "num_tokens": 785595.0,
131
+ "step": 550
132
+ },
133
+ {
134
+ "entropy": 0.6960959500074386,
135
+ "epoch": 1.5431145431145432,
136
+ "grad_norm": 0.5145474672317505,
137
+ "learning_rate": 0.0002685541661937683,
138
+ "loss": 0.6358638763427734,
139
+ "mean_token_accuracy": 0.8131621342897415,
140
+ "num_tokens": 857551.0,
141
+ "step": 600
142
+ },
143
+ {
144
+ "entropy": 0.7220414417982102,
145
+ "epoch": 1.6718146718146718,
146
+ "grad_norm": 0.9381059408187866,
147
+ "learning_rate": 0.00026728196281103746,
148
+ "loss": 0.6531407928466797,
149
+ "mean_token_accuracy": 0.811619822382927,
150
+ "num_tokens": 926864.0,
151
+ "step": 650
152
+ },
153
+ {
154
+ "entropy": 0.6955459499359131,
155
+ "epoch": 1.8005148005148004,
156
+ "grad_norm": 0.6115108728408813,
157
+ "learning_rate": 0.0002657444718086503,
158
+ "loss": 0.6269588088989257,
159
+ "mean_token_accuracy": 0.8155373805761337,
160
+ "num_tokens": 998824.0,
161
+ "step": 700
162
+ },
163
+ {
164
+ "entropy": 0.6929812705516816,
165
+ "epoch": 1.9292149292149292,
166
+ "grad_norm": 0.749489426612854,
167
+ "learning_rate": 0.0002639447877206115,
168
+ "loss": 0.629054069519043,
169
+ "mean_token_accuracy": 0.8183021235466004,
170
+ "num_tokens": 1069332.0,
171
+ "step": 750
172
+ },
173
+ {
174
+ "epoch": 2.0,
175
+ "eval_entropy": 0.5833694102223387,
176
+ "eval_loss": 0.6328718662261963,
177
+ "eval_mean_token_accuracy": 0.8159044071571114,
178
+ "eval_num_tokens": 1110986.0,
179
+ "eval_runtime": 161.885,
180
+ "eval_samples_per_second": 9.544,
181
+ "eval_steps_per_second": 1.198,
182
+ "step": 778
183
+ },
184
+ {
185
+ "entropy": 0.647241060480927,
186
+ "epoch": 2.056628056628057,
187
+ "grad_norm": 0.7070767879486084,
188
+ "learning_rate": 0.00026188653280135975,
189
+ "loss": 0.5823922348022461,
190
+ "mean_token_accuracy": 0.8265365301960647,
191
+ "num_tokens": 1141195.0,
192
+ "step": 800
193
+ },
194
+ {
195
+ "entropy": 0.5995378407835961,
196
+ "epoch": 2.1853281853281854,
197
+ "grad_norm": 0.8090486526489258,
198
+ "learning_rate": 0.0002595738497351955,
199
+ "loss": 0.5325597763061524,
200
+ "mean_token_accuracy": 0.8369336777925491,
201
+ "num_tokens": 1210708.0,
202
+ "step": 850
203
+ },
204
+ {
205
+ "entropy": 0.6043145382404327,
206
+ "epoch": 2.314028314028314,
207
+ "grad_norm": 0.8279913067817688,
208
+ "learning_rate": 0.00025701139329823054,
209
+ "loss": 0.5414446258544922,
210
+ "mean_token_accuracy": 0.8361396533250809,
211
+ "num_tokens": 1283441.0,
212
+ "step": 900
213
+ },
214
+ {
215
+ "entropy": 0.5953224584460258,
216
+ "epoch": 2.4427284427284426,
217
+ "grad_norm": 0.6075023412704468,
218
+ "learning_rate": 0.00025420432098964183,
219
+ "loss": 0.536654167175293,
220
+ "mean_token_accuracy": 0.8340093129873276,
221
+ "num_tokens": 1356418.0,
222
+ "step": 950
223
+ },
224
+ {
225
+ "entropy": 0.5998479858040809,
226
+ "epoch": 2.571428571428571,
227
+ "grad_norm": 1.0311471223831177,
228
+ "learning_rate": 0.0002511582826510862,
229
+ "loss": 0.5372924423217773,
230
+ "mean_token_accuracy": 0.8366045409440994,
231
+ "num_tokens": 1427797.0,
232
+ "step": 1000
233
+ },
234
+ {
235
+ "entropy": 0.5980158120393753,
236
+ "epoch": 2.7001287001287,
237
+ "grad_norm": 0.5971426367759705,
238
+ "learning_rate": 0.0002478794090951689,
239
+ "loss": 0.5392885208129883,
240
+ "mean_token_accuracy": 0.8347727072238922,
241
+ "num_tokens": 1498082.0,
242
+ "step": 1050
243
+ },
244
+ {
245
+ "entropy": 0.5990680930018425,
246
+ "epoch": 2.828828828828829,
247
+ "grad_norm": 0.5662627220153809,
248
+ "learning_rate": 0.0002443742997658538,
249
+ "loss": 0.5360498428344727,
250
+ "mean_token_accuracy": 0.8371847170591354,
251
+ "num_tokens": 1568798.0,
252
+ "step": 1100
253
+ },
254
+ {
255
+ "entropy": 0.5957039377093315,
256
+ "epoch": 2.9575289575289574,
257
+ "grad_norm": 0.5043798685073853,
258
+ "learning_rate": 0.00024065000945565205,
259
+ "loss": 0.5342231369018555,
260
+ "mean_token_accuracy": 0.8380735236406326,
261
+ "num_tokens": 1643449.0,
262
+ "step": 1150
263
+ },
264
+ {
265
+ "epoch": 3.0,
266
+ "eval_entropy": 0.5346978819861854,
267
+ "eval_loss": 0.6255015134811401,
268
+ "eval_mean_token_accuracy": 0.8202014476368108,
269
+ "eval_num_tokens": 1666479.0,
270
+ "eval_runtime": 161.6098,
271
+ "eval_samples_per_second": 9.56,
272
+ "eval_steps_per_second": 1.2,
273
+ "step": 1167
274
+ },
275
+ {
276
+ "entropy": 0.5212209137401196,
277
+ "epoch": 3.0849420849420848,
278
+ "grad_norm": 0.7095440626144409,
279
+ "learning_rate": 0.00023671403410632178,
280
+ "loss": 0.45311901092529294,
281
+ "mean_token_accuracy": 0.856362871449403,
282
+ "num_tokens": 1713536.0,
283
+ "step": 1200
284
+ },
285
+ {
286
+ "entropy": 0.4701593083143234,
287
+ "epoch": 3.213642213642214,
288
+ "grad_norm": 0.6408083438873291,
289
+ "learning_rate": 0.0002325742957216607,
290
+ "loss": 0.39916397094726563,
291
+ "mean_token_accuracy": 0.8698061722517013,
292
+ "num_tokens": 1785609.0,
293
+ "step": 1250
294
+ },
295
+ {
296
+ "entropy": 0.4766591975092888,
297
+ "epoch": 3.3423423423423424,
298
+ "grad_norm": 0.6415093541145325,
299
+ "learning_rate": 0.0002282391264227552,
300
+ "loss": 0.4116698455810547,
301
+ "mean_token_accuracy": 0.8679435575008392,
302
+ "num_tokens": 1858651.0,
303
+ "step": 1300
304
+ },
305
+ {
306
+ "entropy": 0.4937947469949722,
307
+ "epoch": 3.471042471042471,
308
+ "grad_norm": 0.6549825072288513,
309
+ "learning_rate": 0.00022371725167778054,
310
+ "loss": 0.4296376037597656,
311
+ "mean_token_accuracy": 0.8609357953071595,
312
+ "num_tokens": 1928692.0,
313
+ "step": 1350
314
+ },
315
+ {
316
+ "entropy": 0.4920153194665909,
317
+ "epoch": 3.5997425997425996,
318
+ "grad_norm": 0.6452126502990723,
319
+ "learning_rate": 0.00021901777274010406,
320
+ "loss": 0.4307489013671875,
321
+ "mean_token_accuracy": 0.8606827831268311,
322
+ "num_tokens": 1998668.0,
323
+ "step": 1400
324
+ },
325
+ {
326
+ "entropy": 0.490042342543602,
327
+ "epoch": 3.7284427284427286,
328
+ "grad_norm": 0.5727734565734863,
329
+ "learning_rate": 0.0002141501483300395,
330
+ "loss": 0.4295254135131836,
331
+ "mean_token_accuracy": 0.8616545403003693,
332
+ "num_tokens": 2072809.0,
333
+ "step": 1450
334
+ },
335
+ {
336
+ "entropy": 0.49857193052768706,
337
+ "epoch": 3.857142857142857,
338
+ "grad_norm": 0.7732954025268555,
339
+ "learning_rate": 0.00020912417559712133,
340
+ "loss": 0.4289303207397461,
341
+ "mean_token_accuracy": 0.8616443765163422,
342
+ "num_tokens": 2142475.0,
343
+ "step": 1500
344
+ },
345
+ {
346
+ "entropy": 0.4746784272789955,
347
+ "epoch": 3.985842985842986,
348
+ "grad_norm": 0.5791187882423401,
349
+ "learning_rate": 0.00020394997040121726,
350
+ "loss": 0.4180263900756836,
351
+ "mean_token_accuracy": 0.866080379486084,
352
+ "num_tokens": 2214044.0,
353
+ "step": 1550
354
+ },
355
+ {
356
+ "epoch": 4.0,
357
+ "eval_entropy": 0.4793926059585257,
358
+ "eval_loss": 0.627627968788147,
359
+ "eval_mean_token_accuracy": 0.8233870095813397,
360
+ "eval_num_tokens": 2221972.0,
361
+ "eval_runtime": 162.0225,
362
+ "eval_samples_per_second": 9.536,
363
+ "eval_steps_per_second": 1.197,
364
+ "step": 1556
365
+ },
366
+ {
367
+ "entropy": 0.38195489000792454,
368
+ "epoch": 4.113256113256114,
369
+ "grad_norm": 0.6002617478370667,
370
+ "learning_rate": 0.0001986379469521669,
371
+ "loss": 0.30819049835205076,
372
+ "mean_token_accuracy": 0.8977848634575353,
373
+ "num_tokens": 2282164.0,
374
+ "step": 1600
375
+ },
376
+ {
377
+ "entropy": 0.3655787402391434,
378
+ "epoch": 4.241956241956242,
379
+ "grad_norm": 0.7100041508674622,
380
+ "learning_rate": 0.00019319879684892634,
381
+ "loss": 0.29959835052490236,
382
+ "mean_token_accuracy": 0.8991208010911942,
383
+ "num_tokens": 2353213.0,
384
+ "step": 1650
385
+ },
386
+ {
387
+ "entropy": 0.3821141055226326,
388
+ "epoch": 4.370656370656371,
389
+ "grad_norm": 0.5848307013511658,
390
+ "learning_rate": 0.00018764346756040715,
391
+ "loss": 0.313802490234375,
392
+ "mean_token_accuracy": 0.895167955160141,
393
+ "num_tokens": 2425068.0,
394
+ "step": 1700
395
+ },
396
+ {
397
+ "entropy": 0.37083797007799146,
398
+ "epoch": 4.499356499356499,
399
+ "grad_norm": 0.6447024941444397,
400
+ "learning_rate": 0.00018198314039132143,
401
+ "loss": 0.30583988189697264,
402
+ "mean_token_accuracy": 0.8961733293533325,
403
+ "num_tokens": 2498321.0,
404
+ "step": 1750
405
+ },
406
+ {
407
+ "entropy": 0.3791545969247818,
408
+ "epoch": 4.628056628056628,
409
+ "grad_norm": 0.6575382351875305,
410
+ "learning_rate": 0.00017622920797738184,
411
+ "loss": 0.3088031005859375,
412
+ "mean_token_accuracy": 0.8960050916671753,
413
+ "num_tokens": 2570321.0,
414
+ "step": 1800
415
+ },
416
+ {
417
+ "entropy": 0.3946831756830215,
418
+ "epoch": 4.756756756756757,
419
+ "grad_norm": 0.5351552963256836,
420
+ "learning_rate": 0.00017039325135515207,
421
+ "loss": 0.3229162979125977,
422
+ "mean_token_accuracy": 0.8920552498102188,
423
+ "num_tokens": 2642851.0,
424
+ "step": 1850
425
+ },
426
+ {
427
+ "entropy": 0.37298239797353744,
428
+ "epoch": 4.885456885456885,
429
+ "grad_norm": 0.7624587416648865,
430
+ "learning_rate": 0.00016448701665269964,
431
+ "loss": 0.3067934799194336,
432
+ "mean_token_accuracy": 0.8951873427629471,
433
+ "num_tokens": 2715629.0,
434
+ "step": 1900
435
+ },
436
+ {
437
+ "epoch": 5.0,
438
+ "eval_entropy": 0.3738150204887095,
439
+ "eval_loss": 0.7131896615028381,
440
+ "eval_mean_token_accuracy": 0.8207752468045225,
441
+ "eval_num_tokens": 2777465.0,
442
+ "eval_runtime": 162.1244,
443
+ "eval_samples_per_second": 9.53,
444
+ "eval_steps_per_second": 1.197,
445
+ "step": 1945
446
+ },
447
+ {
448
+ "entropy": 0.377820266617669,
449
+ "epoch": 5.012870012870013,
450
+ "grad_norm": 0.41990190744400024,
451
+ "learning_rate": 0.00015852239144796624,
452
+ "loss": 0.3058685111999512,
453
+ "mean_token_accuracy": 0.8964343480389527,
454
+ "num_tokens": 2784896.0,
455
+ "step": 1950
456
+ },
457
+ {
458
+ "entropy": 0.2714502356946468,
459
+ "epoch": 5.141570141570142,
460
+ "grad_norm": 0.422568678855896,
461
+ "learning_rate": 0.00015251138084243995,
462
+ "loss": 0.2093442153930664,
463
+ "mean_token_accuracy": 0.9311346983909607,
464
+ "num_tokens": 2854374.0,
465
+ "step": 2000
466
+ },
467
+ {
468
+ "entropy": 0.268475965410471,
469
+ "epoch": 5.27027027027027,
470
+ "grad_norm": 0.6637414693832397,
471
+ "learning_rate": 0.0001464660832982852,
472
+ "loss": 0.20736080169677734,
473
+ "mean_token_accuracy": 0.9289199805259705,
474
+ "num_tokens": 2927362.0,
475
+ "step": 2050
476
+ },
477
+ {
478
+ "entropy": 0.2644876340031624,
479
+ "epoch": 5.398970398970399,
480
+ "grad_norm": 0.47317707538604736,
481
+ "learning_rate": 0.00014039866628756467,
482
+ "loss": 0.20464908599853515,
483
+ "mean_token_accuracy": 0.9300856202840805,
484
+ "num_tokens": 3000143.0,
485
+ "step": 2100
486
+ },
487
+ {
488
+ "entropy": 0.2675253136456013,
489
+ "epoch": 5.527670527670527,
490
+ "grad_norm": 0.5253982543945312,
491
+ "learning_rate": 0.00013432134180256338,
492
+ "loss": 0.21154335021972656,
493
+ "mean_token_accuracy": 0.9283734840154648,
494
+ "num_tokens": 3072561.0,
495
+ "step": 2150
496
+ },
497
+ {
498
+ "entropy": 0.27213907435536383,
499
+ "epoch": 5.656370656370656,
500
+ "grad_norm": 0.46738553047180176,
501
+ "learning_rate": 0.00012824634177650664,
502
+ "loss": 0.21339216232299804,
503
+ "mean_token_accuracy": 0.9272083270549775,
504
+ "num_tokens": 3144831.0,
505
+ "step": 2200
506
+ },
507
+ {
508
+ "entropy": 0.2785488124191761,
509
+ "epoch": 5.785070785070785,
510
+ "grad_norm": 0.4469502866268158,
511
+ "learning_rate": 0.00012218589346414205,
512
+ "loss": 0.21601097106933595,
513
+ "mean_token_accuracy": 0.9255663657188415,
514
+ "num_tokens": 3215960.0,
515
+ "step": 2250
516
+ },
517
+ {
518
+ "entropy": 0.2699935150146484,
519
+ "epoch": 5.913770913770914,
520
+ "grad_norm": 0.7359778881072998,
521
+ "learning_rate": 0.00011615219483173828,
522
+ "loss": 0.20725584030151367,
523
+ "mean_token_accuracy": 0.9286630594730377,
524
+ "num_tokens": 3287499.0,
525
+ "step": 2300
526
+ },
527
+ {
528
+ "epoch": 6.0,
529
+ "eval_entropy": 0.26417383682174783,
530
+ "eval_loss": 0.8880229592323303,
531
+ "eval_mean_token_accuracy": 0.8159987201395723,
532
+ "eval_num_tokens": 3332958.0,
533
+ "eval_runtime": 162.0991,
534
+ "eval_samples_per_second": 9.531,
535
+ "eval_steps_per_second": 1.197,
536
+ "step": 2334
537
+ },
538
+ {
539
+ "entropy": 0.24814540704693458,
540
+ "epoch": 6.041184041184041,
541
+ "grad_norm": 0.4953760802745819,
542
+ "learning_rate": 0.00011015739000603316,
543
+ "loss": 0.18749794006347656,
544
+ "mean_token_accuracy": 0.9370789509831052,
545
+ "num_tokens": 3356879.0,
546
+ "step": 2350
547
+ },
548
+ {
549
+ "entropy": 0.19976271741092205,
550
+ "epoch": 6.1698841698841695,
551
+ "grad_norm": 0.4834803342819214,
552
+ "learning_rate": 0.00010421354483154553,
553
+ "loss": 0.14283526420593262,
554
+ "mean_token_accuracy": 0.9521516615152359,
555
+ "num_tokens": 3427587.0,
556
+ "step": 2400
557
+ },
558
+ {
559
+ "entropy": 0.2060488449037075,
560
+ "epoch": 6.298584298584299,
561
+ "grad_norm": 0.4888673722743988,
562
+ "learning_rate": 9.8332622585447e-05,
563
+ "loss": 0.14414511680603026,
564
+ "mean_token_accuracy": 0.9510996866226197,
565
+ "num_tokens": 3498688.0,
566
+ "step": 2450
567
+ },
568
+ {
569
+ "entropy": 0.2059111550450325,
570
+ "epoch": 6.427284427284428,
571
+ "grad_norm": 0.4064404368400574,
572
+ "learning_rate": 9.252645989887253e-05,
573
+ "loss": 0.14820143699645996,
574
+ "mean_token_accuracy": 0.9507584601640702,
575
+ "num_tokens": 3566137.0,
576
+ "step": 2500
577
+ },
578
+ {
579
+ "entropy": 0.19700154662132263,
580
+ "epoch": 6.555984555984556,
581
+ "grad_norm": 0.467965304851532,
582
+ "learning_rate": 8.680674293313417e-05,
583
+ "loss": 0.14303470611572267,
584
+ "mean_token_accuracy": 0.9515972435474396,
585
+ "num_tokens": 3639573.0,
586
+ "step": 2550
587
+ },
588
+ {
589
+ "entropy": 0.20180423602461814,
590
+ "epoch": 6.684684684684685,
591
+ "grad_norm": 0.36836138367652893,
592
+ "learning_rate": 8.118498385878736e-05,
593
+ "loss": 0.14280882835388184,
594
+ "mean_token_accuracy": 0.9515993863344192,
595
+ "num_tokens": 3710433.0,
596
+ "step": 2600
597
+ },
598
+ {
599
+ "entropy": 0.20024395987391472,
600
+ "epoch": 6.813384813384813,
601
+ "grad_norm": 0.38375866413116455,
602
+ "learning_rate": 7.567249768489171e-05,
603
+ "loss": 0.1427844524383545,
604
+ "mean_token_accuracy": 0.9524166631698608,
605
+ "num_tokens": 3781550.0,
606
+ "step": 2650
607
+ },
608
+ {
609
+ "entropy": 0.19561587080359458,
610
+ "epoch": 6.942084942084942,
611
+ "grad_norm": 0.41185441613197327,
612
+ "learning_rate": 7.028037948510187e-05,
613
+ "loss": 0.13993803024291993,
614
+ "mean_token_accuracy": 0.9522478264570237,
615
+ "num_tokens": 3854952.0,
616
+ "step": 2700
617
+ },
618
+ {
619
+ "epoch": 7.0,
620
+ "eval_entropy": 0.19501976062034823,
621
+ "eval_loss": 1.0653952360153198,
622
+ "eval_mean_token_accuracy": 0.8205490803595671,
623
+ "eval_num_tokens": 3888451.0,
624
+ "eval_runtime": 161.8533,
625
+ "eval_samples_per_second": 9.546,
626
+ "eval_steps_per_second": 1.199,
627
+ "step": 2723
628
+ }
629
+ ],
630
+ "logging_steps": 50,
631
+ "max_steps": 3890,
632
+ "num_input_tokens_seen": 0,
633
+ "num_train_epochs": 10,
634
+ "save_steps": 500,
635
+ "stateful_callbacks": {
636
+ "TrainerControl": {
637
+ "args": {
638
+ "should_epoch_stop": false,
639
+ "should_evaluate": false,
640
+ "should_log": false,
641
+ "should_save": true,
642
+ "should_training_stop": false
643
+ },
644
+ "attributes": {}
645
+ }
646
+ },
647
+ "total_flos": 6.516671077296845e+17,
648
+ "train_batch_size": 8,
649
+ "trial_name": null,
650
+ "trial_params": null
651
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/README.md ADDED
@@ -0,0 +1,209 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen3-14B-Base
3
+ library_name: peft
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - base_model:adapter:Qwen/Qwen3-14B-Base
7
+ - lora
8
+ - sft
9
+ - transformers
10
+ - trl
11
+ ---
12
+
13
+ # Model Card for Model ID
14
+
15
+ <!-- Provide a quick summary of what the model is/does. -->
16
+
17
+
18
+
19
+ ## Model Details
20
+
21
+ ### Model Description
22
+
23
+ <!-- Provide a longer summary of what this model is. -->
24
+
25
+
26
+
27
+ - **Developed by:** [More Information Needed]
28
+ - **Funded by [optional]:** [More Information Needed]
29
+ - **Shared by [optional]:** [More Information Needed]
30
+ - **Model type:** [More Information Needed]
31
+ - **Language(s) (NLP):** [More Information Needed]
32
+ - **License:** [More Information Needed]
33
+ - **Finetuned from model [optional]:** [More Information Needed]
34
+
35
+ ### Model Sources [optional]
36
+
37
+ <!-- Provide the basic links for the model. -->
38
+
39
+ - **Repository:** [More Information Needed]
40
+ - **Paper [optional]:** [More Information Needed]
41
+ - **Demo [optional]:** [More Information Needed]
42
+
43
+ ## Uses
44
+
45
+ <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
46
+
47
+ ### Direct Use
48
+
49
+ <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
50
+
51
+ [More Information Needed]
52
+
53
+ ### Downstream Use [optional]
54
+
55
+ <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
56
+
57
+ [More Information Needed]
58
+
59
+ ### Out-of-Scope Use
60
+
61
+ <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
62
+
63
+ [More Information Needed]
64
+
65
+ ## Bias, Risks, and Limitations
66
+
67
+ <!-- This section is meant to convey both technical and sociotechnical limitations. -->
68
+
69
+ [More Information Needed]
70
+
71
+ ### Recommendations
72
+
73
+ <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
74
+
75
+ Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
76
+
77
+ ## How to Get Started with the Model
78
+
79
+ Use the code below to get started with the model.
80
+
81
+ [More Information Needed]
82
+
83
+ ## Training Details
84
+
85
+ ### Training Data
86
+
87
+ <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
88
+
89
+ [More Information Needed]
90
+
91
+ ### Training Procedure
92
+
93
+ <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
94
+
95
+ #### Preprocessing [optional]
96
+
97
+ [More Information Needed]
98
+
99
+
100
+ #### Training Hyperparameters
101
+
102
+ - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
103
+
104
+ #### Speeds, Sizes, Times [optional]
105
+
106
+ <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
107
+
108
+ [More Information Needed]
109
+
110
+ ## Evaluation
111
+
112
+ <!-- This section describes the evaluation protocols and provides the results. -->
113
+
114
+ ### Testing Data, Factors & Metrics
115
+
116
+ #### Testing Data
117
+
118
+ <!-- This should link to a Dataset Card if possible. -->
119
+
120
+ [More Information Needed]
121
+
122
+ #### Factors
123
+
124
+ <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
125
+
126
+ [More Information Needed]
127
+
128
+ #### Metrics
129
+
130
+ <!-- These are the evaluation metrics being used, ideally with a description of why. -->
131
+
132
+ [More Information Needed]
133
+
134
+ ### Results
135
+
136
+ [More Information Needed]
137
+
138
+ #### Summary
139
+
140
+
141
+
142
+ ## Model Examination [optional]
143
+
144
+ <!-- Relevant interpretability work for the model goes here -->
145
+
146
+ [More Information Needed]
147
+
148
+ ## Environmental Impact
149
+
150
+ <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
151
+
152
+ Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
153
+
154
+ - **Hardware Type:** [More Information Needed]
155
+ - **Hours used:** [More Information Needed]
156
+ - **Cloud Provider:** [More Information Needed]
157
+ - **Compute Region:** [More Information Needed]
158
+ - **Carbon Emitted:** [More Information Needed]
159
+
160
+ ## Technical Specifications [optional]
161
+
162
+ ### Model Architecture and Objective
163
+
164
+ [More Information Needed]
165
+
166
+ ### Compute Infrastructure
167
+
168
+ [More Information Needed]
169
+
170
+ #### Hardware
171
+
172
+ [More Information Needed]
173
+
174
+ #### Software
175
+
176
+ [More Information Needed]
177
+
178
+ ## Citation [optional]
179
+
180
+ <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
181
+
182
+ **BibTeX:**
183
+
184
+ [More Information Needed]
185
+
186
+ **APA:**
187
+
188
+ [More Information Needed]
189
+
190
+ ## Glossary [optional]
191
+
192
+ <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
193
+
194
+ [More Information Needed]
195
+
196
+ ## More Information [optional]
197
+
198
+ [More Information Needed]
199
+
200
+ ## Model Card Authors [optional]
201
+
202
+ [More Information Needed]
203
+
204
+ ## Model Card Contact
205
+
206
+ [More Information Needed]
207
+ ### Framework versions
208
+
209
+ - PEFT 0.19.1
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/adapter_config.json ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "Qwen/Qwen3-14B-Base",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "layer_replication": null,
16
+ "layers_pattern": null,
17
+ "layers_to_transform": null,
18
+ "loftq_config": {},
19
+ "lora_alpha": 256,
20
+ "lora_bias": false,
21
+ "lora_dropout": 0.00237968804112545,
22
+ "lora_ga_config": null,
23
+ "megatron_config": null,
24
+ "megatron_core": "megatron.core",
25
+ "modules_to_save": null,
26
+ "peft_type": "LORA",
27
+ "peft_version": "0.19.1",
28
+ "qalora_group_size": 16,
29
+ "r": 128,
30
+ "rank_pattern": {},
31
+ "revision": null,
32
+ "target_modules": [
33
+ "gate_proj",
34
+ "q_proj",
35
+ "up_proj",
36
+ "o_proj",
37
+ "v_proj",
38
+ "k_proj",
39
+ "down_proj"
40
+ ],
41
+ "target_parameters": null,
42
+ "task_type": "CAUSAL_LM",
43
+ "trainable_token_indices": null,
44
+ "use_bdlora": null,
45
+ "use_dora": false,
46
+ "use_qalora": false,
47
+ "use_rslora": false
48
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/chat_template.jinja ADDED
@@ -0,0 +1,85 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- if tools %}
2
+ {{- '<|im_start|>system\n' }}
3
+ {%- if messages[0].role == 'system' %}
4
+ {{- messages[0].content + '\n\n' }}
5
+ {%- endif %}
6
+ {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
7
+ {%- for tool in tools %}
8
+ {{- "\n" }}
9
+ {{- tool | tojson }}
10
+ {%- endfor %}
11
+ {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
12
+ {%- else %}
13
+ {%- if messages[0].role == 'system' %}
14
+ {{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
15
+ {%- endif %}
16
+ {%- endif %}
17
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
18
+ {%- for message in messages[::-1] %}
19
+ {%- set index = (messages|length - 1) - loop.index0 %}
20
+ {%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
21
+ {%- set ns.multi_step_tool = false %}
22
+ {%- set ns.last_query_index = index %}
23
+ {%- endif %}
24
+ {%- endfor %}
25
+ {%- for message in messages %}
26
+ {%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
27
+ {{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
28
+ {%- elif message.role == "assistant" %}
29
+ {%- set content = message.content %}
30
+ {%- set reasoning_content = '' %}
31
+ {%- if message.reasoning_content is defined and message.reasoning_content is not none %}
32
+ {%- set reasoning_content = message.reasoning_content %}
33
+ {%- else %}
34
+ {%- if '</think>' in message.content %}
35
+ {%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
36
+ {%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
37
+ {%- endif %}
38
+ {%- endif %}
39
+ {%- if loop.index0 > ns.last_query_index %}
40
+ {%- if loop.last or (not loop.last and reasoning_content) %}
41
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
42
+ {%- else %}
43
+ {{- '<|im_start|>' + message.role + '\n' + content }}
44
+ {%- endif %}
45
+ {%- else %}
46
+ {{- '<|im_start|>' + message.role + '\n' + content }}
47
+ {%- endif %}
48
+ {%- if message.tool_calls %}
49
+ {%- for tool_call in message.tool_calls %}
50
+ {%- if (loop.first and content) or (not loop.first) %}
51
+ {{- '\n' }}
52
+ {%- endif %}
53
+ {%- if tool_call.function %}
54
+ {%- set tool_call = tool_call.function %}
55
+ {%- endif %}
56
+ {{- '<tool_call>\n{"name": "' }}
57
+ {{- tool_call.name }}
58
+ {{- '", "arguments": ' }}
59
+ {%- if tool_call.arguments is string %}
60
+ {{- tool_call.arguments }}
61
+ {%- else %}
62
+ {{- tool_call.arguments | tojson }}
63
+ {%- endif %}
64
+ {{- '}\n</tool_call>' }}
65
+ {%- endfor %}
66
+ {%- endif %}
67
+ {{- '<|im_end|>\n' }}
68
+ {%- elif message.role == "tool" %}
69
+ {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
70
+ {{- '<|im_start|>user' }}
71
+ {%- endif %}
72
+ {{- '\n<tool_response>\n' }}
73
+ {{- message.content }}
74
+ {{- '\n</tool_response>' }}
75
+ {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
76
+ {{- '<|im_end|>\n' }}
77
+ {%- endif %}
78
+ {%- endif %}
79
+ {%- endfor %}
80
+ {%- if add_generation_prompt %}
81
+ {{- '<|im_start|>assistant\n' }}
82
+ {%- if enable_thinking is defined and enable_thinking is false %}
83
+ {{- '<think>\n\n</think>\n\n' }}
84
+ {%- endif %}
85
+ {%- endif %}
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/tokenizer_config.json ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "backend": "tokenizers",
4
+ "bos_token": null,
5
+ "clean_up_tokenization_spaces": false,
6
+ "eos_token": "<|endoftext|>",
7
+ "errors": "replace",
8
+ "extra_special_tokens": [
9
+ "<|im_start|>",
10
+ "<|im_end|>",
11
+ "<|object_ref_start|>",
12
+ "<|object_ref_end|>",
13
+ "<|box_start|>",
14
+ "<|box_end|>",
15
+ "<|quad_start|>",
16
+ "<|quad_end|>",
17
+ "<|vision_start|>",
18
+ "<|vision_end|>",
19
+ "<|vision_pad|>",
20
+ "<|image_pad|>",
21
+ "<|video_pad|>"
22
+ ],
23
+ "is_local": false,
24
+ "model_max_length": 131072,
25
+ "pad_token": "<|endoftext|>",
26
+ "split_special_tokens": false,
27
+ "tokenizer_class": "Qwen2Tokenizer",
28
+ "unk_token": null
29
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/trainer_state.json ADDED
@@ -0,0 +1,742 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "best_global_step": null,
3
+ "best_metric": null,
4
+ "best_model_checkpoint": null,
5
+ "epoch": 8.0,
6
+ "eval_steps": 500,
7
+ "global_step": 3112,
8
+ "is_hyper_param_search": false,
9
+ "is_local_process_zero": true,
10
+ "is_world_process_zero": true,
11
+ "log_history": [
12
+ {
13
+ "entropy": 1.6110110306739807,
14
+ "epoch": 0.1287001287001287,
15
+ "grad_norm": 0.8611119389533997,
16
+ "learning_rate": 3.4130257962133866e-05,
17
+ "loss": 1.53956298828125,
18
+ "mean_token_accuracy": 0.6652342769503593,
19
+ "num_tokens": 73407.0,
20
+ "step": 50
21
+ },
22
+ {
23
+ "entropy": 0.8662820833921433,
24
+ "epoch": 0.2574002574002574,
25
+ "grad_norm": 0.7405035495758057,
26
+ "learning_rate": 6.895705180104598e-05,
27
+ "loss": 0.7929539489746094,
28
+ "mean_token_accuracy": 0.7786347842216492,
29
+ "num_tokens": 143994.0,
30
+ "step": 100
31
+ },
32
+ {
33
+ "entropy": 0.7912819278240204,
34
+ "epoch": 0.3861003861003861,
35
+ "grad_norm": 0.6105485558509827,
36
+ "learning_rate": 0.00010378384563995809,
37
+ "loss": 0.7236511993408203,
38
+ "mean_token_accuracy": 0.7925467795133591,
39
+ "num_tokens": 216171.0,
40
+ "step": 150
41
+ },
42
+ {
43
+ "entropy": 0.7777709531784057,
44
+ "epoch": 0.5148005148005148,
45
+ "grad_norm": 0.5781793594360352,
46
+ "learning_rate": 0.0001386106394788702,
47
+ "loss": 0.7024919891357422,
48
+ "mean_token_accuracy": 0.7969876372814179,
49
+ "num_tokens": 284702.0,
50
+ "step": 200
51
+ },
52
+ {
53
+ "entropy": 0.7632609683275223,
54
+ "epoch": 0.6435006435006435,
55
+ "grad_norm": 0.6553444862365723,
56
+ "learning_rate": 0.00017343743331778232,
57
+ "loss": 0.7000718688964844,
58
+ "mean_token_accuracy": 0.7994742071628571,
59
+ "num_tokens": 356393.0,
60
+ "step": 250
61
+ },
62
+ {
63
+ "entropy": 0.754165632724762,
64
+ "epoch": 0.7722007722007722,
65
+ "grad_norm": 0.634222149848938,
66
+ "learning_rate": 0.00020826422715669444,
67
+ "loss": 0.6909049987792969,
68
+ "mean_token_accuracy": 0.8016809666156769,
69
+ "num_tokens": 426916.0,
70
+ "step": 300
71
+ },
72
+ {
73
+ "entropy": 0.7530711203813553,
74
+ "epoch": 0.9009009009009009,
75
+ "grad_norm": 0.7020156383514404,
76
+ "learning_rate": 0.00024309102099560653,
77
+ "loss": 0.69554931640625,
78
+ "mean_token_accuracy": 0.8002807641029358,
79
+ "num_tokens": 499603.0,
80
+ "step": 350
81
+ },
82
+ {
83
+ "epoch": 1.0,
84
+ "eval_entropy": 0.5728044164242204,
85
+ "eval_loss": 0.6785851716995239,
86
+ "eval_mean_token_accuracy": 0.798433246993527,
87
+ "eval_num_tokens": 555493.0,
88
+ "eval_runtime": 162.4129,
89
+ "eval_samples_per_second": 9.513,
90
+ "eval_steps_per_second": 1.194,
91
+ "step": 389
92
+ },
93
+ {
94
+ "entropy": 0.7489313946829902,
95
+ "epoch": 1.0283140283140284,
96
+ "grad_norm": 0.7532815933227539,
97
+ "learning_rate": 0.0002709470016827303,
98
+ "loss": 0.6856581878662109,
99
+ "mean_token_accuracy": 0.8023027079273956,
100
+ "num_tokens": 570813.0,
101
+ "step": 400
102
+ },
103
+ {
104
+ "entropy": 0.7238409864902496,
105
+ "epoch": 1.157014157014157,
106
+ "grad_norm": 1.2016215324401855,
107
+ "learning_rate": 0.0002707561443541359,
108
+ "loss": 0.6699818420410156,
109
+ "mean_token_accuracy": 0.8070302194356919,
110
+ "num_tokens": 642956.0,
111
+ "step": 450
112
+ },
113
+ {
114
+ "entropy": 0.7388938587903976,
115
+ "epoch": 1.2857142857142856,
116
+ "grad_norm": 0.7279272079467773,
117
+ "learning_rate": 0.0002702930068622498,
118
+ "loss": 0.6728517150878907,
119
+ "mean_token_accuracy": 0.8049580943584442,
120
+ "num_tokens": 714498.0,
121
+ "step": 500
122
+ },
123
+ {
124
+ "entropy": 0.7284485149383545,
125
+ "epoch": 1.4144144144144144,
126
+ "grad_norm": 0.8563987016677856,
127
+ "learning_rate": 0.0002695585213716931,
128
+ "loss": 0.6657986450195312,
129
+ "mean_token_accuracy": 0.8085588800907135,
130
+ "num_tokens": 785595.0,
131
+ "step": 550
132
+ },
133
+ {
134
+ "entropy": 0.6960959500074386,
135
+ "epoch": 1.5431145431145432,
136
+ "grad_norm": 0.5145474672317505,
137
+ "learning_rate": 0.0002685541661937683,
138
+ "loss": 0.6358638763427734,
139
+ "mean_token_accuracy": 0.8131621342897415,
140
+ "num_tokens": 857551.0,
141
+ "step": 600
142
+ },
143
+ {
144
+ "entropy": 0.7220414417982102,
145
+ "epoch": 1.6718146718146718,
146
+ "grad_norm": 0.9381059408187866,
147
+ "learning_rate": 0.00026728196281103746,
148
+ "loss": 0.6531407928466797,
149
+ "mean_token_accuracy": 0.811619822382927,
150
+ "num_tokens": 926864.0,
151
+ "step": 650
152
+ },
153
+ {
154
+ "entropy": 0.6955459499359131,
155
+ "epoch": 1.8005148005148004,
156
+ "grad_norm": 0.6115108728408813,
157
+ "learning_rate": 0.0002657444718086503,
158
+ "loss": 0.6269588088989257,
159
+ "mean_token_accuracy": 0.8155373805761337,
160
+ "num_tokens": 998824.0,
161
+ "step": 700
162
+ },
163
+ {
164
+ "entropy": 0.6929812705516816,
165
+ "epoch": 1.9292149292149292,
166
+ "grad_norm": 0.749489426612854,
167
+ "learning_rate": 0.0002639447877206115,
168
+ "loss": 0.629054069519043,
169
+ "mean_token_accuracy": 0.8183021235466004,
170
+ "num_tokens": 1069332.0,
171
+ "step": 750
172
+ },
173
+ {
174
+ "epoch": 2.0,
175
+ "eval_entropy": 0.5833694102223387,
176
+ "eval_loss": 0.6328718662261963,
177
+ "eval_mean_token_accuracy": 0.8159044071571114,
178
+ "eval_num_tokens": 1110986.0,
179
+ "eval_runtime": 161.885,
180
+ "eval_samples_per_second": 9.544,
181
+ "eval_steps_per_second": 1.198,
182
+ "step": 778
183
+ },
184
+ {
185
+ "entropy": 0.647241060480927,
186
+ "epoch": 2.056628056628057,
187
+ "grad_norm": 0.7070767879486084,
188
+ "learning_rate": 0.00026188653280135975,
189
+ "loss": 0.5823922348022461,
190
+ "mean_token_accuracy": 0.8265365301960647,
191
+ "num_tokens": 1141195.0,
192
+ "step": 800
193
+ },
194
+ {
195
+ "entropy": 0.5995378407835961,
196
+ "epoch": 2.1853281853281854,
197
+ "grad_norm": 0.8090486526489258,
198
+ "learning_rate": 0.0002595738497351955,
199
+ "loss": 0.5325597763061524,
200
+ "mean_token_accuracy": 0.8369336777925491,
201
+ "num_tokens": 1210708.0,
202
+ "step": 850
203
+ },
204
+ {
205
+ "entropy": 0.6043145382404327,
206
+ "epoch": 2.314028314028314,
207
+ "grad_norm": 0.8279913067817688,
208
+ "learning_rate": 0.00025701139329823054,
209
+ "loss": 0.5414446258544922,
210
+ "mean_token_accuracy": 0.8361396533250809,
211
+ "num_tokens": 1283441.0,
212
+ "step": 900
213
+ },
214
+ {
215
+ "entropy": 0.5953224584460258,
216
+ "epoch": 2.4427284427284426,
217
+ "grad_norm": 0.6075023412704468,
218
+ "learning_rate": 0.00025420432098964183,
219
+ "loss": 0.536654167175293,
220
+ "mean_token_accuracy": 0.8340093129873276,
221
+ "num_tokens": 1356418.0,
222
+ "step": 950
223
+ },
224
+ {
225
+ "entropy": 0.5998479858040809,
226
+ "epoch": 2.571428571428571,
227
+ "grad_norm": 1.0311471223831177,
228
+ "learning_rate": 0.0002511582826510862,
229
+ "loss": 0.5372924423217773,
230
+ "mean_token_accuracy": 0.8366045409440994,
231
+ "num_tokens": 1427797.0,
232
+ "step": 1000
233
+ },
234
+ {
235
+ "entropy": 0.5980158120393753,
236
+ "epoch": 2.7001287001287,
237
+ "grad_norm": 0.5971426367759705,
238
+ "learning_rate": 0.0002478794090951689,
239
+ "loss": 0.5392885208129883,
240
+ "mean_token_accuracy": 0.8347727072238922,
241
+ "num_tokens": 1498082.0,
242
+ "step": 1050
243
+ },
244
+ {
245
+ "entropy": 0.5990680930018425,
246
+ "epoch": 2.828828828828829,
247
+ "grad_norm": 0.5662627220153809,
248
+ "learning_rate": 0.0002443742997658538,
249
+ "loss": 0.5360498428344727,
250
+ "mean_token_accuracy": 0.8371847170591354,
251
+ "num_tokens": 1568798.0,
252
+ "step": 1100
253
+ },
254
+ {
255
+ "entropy": 0.5957039377093315,
256
+ "epoch": 2.9575289575289574,
257
+ "grad_norm": 0.5043798685073853,
258
+ "learning_rate": 0.00024065000945565205,
259
+ "loss": 0.5342231369018555,
260
+ "mean_token_accuracy": 0.8380735236406326,
261
+ "num_tokens": 1643449.0,
262
+ "step": 1150
263
+ },
264
+ {
265
+ "epoch": 3.0,
266
+ "eval_entropy": 0.5346978819861854,
267
+ "eval_loss": 0.6255015134811401,
268
+ "eval_mean_token_accuracy": 0.8202014476368108,
269
+ "eval_num_tokens": 1666479.0,
270
+ "eval_runtime": 161.6098,
271
+ "eval_samples_per_second": 9.56,
272
+ "eval_steps_per_second": 1.2,
273
+ "step": 1167
274
+ },
275
+ {
276
+ "entropy": 0.5212209137401196,
277
+ "epoch": 3.0849420849420848,
278
+ "grad_norm": 0.7095440626144409,
279
+ "learning_rate": 0.00023671403410632178,
280
+ "loss": 0.45311901092529294,
281
+ "mean_token_accuracy": 0.856362871449403,
282
+ "num_tokens": 1713536.0,
283
+ "step": 1200
284
+ },
285
+ {
286
+ "entropy": 0.4701593083143234,
287
+ "epoch": 3.213642213642214,
288
+ "grad_norm": 0.6408083438873291,
289
+ "learning_rate": 0.0002325742957216607,
290
+ "loss": 0.39916397094726563,
291
+ "mean_token_accuracy": 0.8698061722517013,
292
+ "num_tokens": 1785609.0,
293
+ "step": 1250
294
+ },
295
+ {
296
+ "entropy": 0.4766591975092888,
297
+ "epoch": 3.3423423423423424,
298
+ "grad_norm": 0.6415093541145325,
299
+ "learning_rate": 0.0002282391264227552,
300
+ "loss": 0.4116698455810547,
301
+ "mean_token_accuracy": 0.8679435575008392,
302
+ "num_tokens": 1858651.0,
303
+ "step": 1300
304
+ },
305
+ {
306
+ "entropy": 0.4937947469949722,
307
+ "epoch": 3.471042471042471,
308
+ "grad_norm": 0.6549825072288513,
309
+ "learning_rate": 0.00022371725167778054,
310
+ "loss": 0.4296376037597656,
311
+ "mean_token_accuracy": 0.8609357953071595,
312
+ "num_tokens": 1928692.0,
313
+ "step": 1350
314
+ },
315
+ {
316
+ "entropy": 0.4920153194665909,
317
+ "epoch": 3.5997425997425996,
318
+ "grad_norm": 0.6452126502990723,
319
+ "learning_rate": 0.00021901777274010406,
320
+ "loss": 0.4307489013671875,
321
+ "mean_token_accuracy": 0.8606827831268311,
322
+ "num_tokens": 1998668.0,
323
+ "step": 1400
324
+ },
325
+ {
326
+ "entropy": 0.490042342543602,
327
+ "epoch": 3.7284427284427286,
328
+ "grad_norm": 0.5727734565734863,
329
+ "learning_rate": 0.0002141501483300395,
330
+ "loss": 0.4295254135131836,
331
+ "mean_token_accuracy": 0.8616545403003693,
332
+ "num_tokens": 2072809.0,
333
+ "step": 1450
334
+ },
335
+ {
336
+ "entropy": 0.49857193052768706,
337
+ "epoch": 3.857142857142857,
338
+ "grad_norm": 0.7732954025268555,
339
+ "learning_rate": 0.00020912417559712133,
340
+ "loss": 0.4289303207397461,
341
+ "mean_token_accuracy": 0.8616443765163422,
342
+ "num_tokens": 2142475.0,
343
+ "step": 1500
344
+ },
345
+ {
346
+ "entropy": 0.4746784272789955,
347
+ "epoch": 3.985842985842986,
348
+ "grad_norm": 0.5791187882423401,
349
+ "learning_rate": 0.00020394997040121726,
350
+ "loss": 0.4180263900756836,
351
+ "mean_token_accuracy": 0.866080379486084,
352
+ "num_tokens": 2214044.0,
353
+ "step": 1550
354
+ },
355
+ {
356
+ "epoch": 4.0,
357
+ "eval_entropy": 0.4793926059585257,
358
+ "eval_loss": 0.627627968788147,
359
+ "eval_mean_token_accuracy": 0.8233870095813397,
360
+ "eval_num_tokens": 2221972.0,
361
+ "eval_runtime": 162.0225,
362
+ "eval_samples_per_second": 9.536,
363
+ "eval_steps_per_second": 1.197,
364
+ "step": 1556
365
+ },
366
+ {
367
+ "entropy": 0.38195489000792454,
368
+ "epoch": 4.113256113256114,
369
+ "grad_norm": 0.6002617478370667,
370
+ "learning_rate": 0.0001986379469521669,
371
+ "loss": 0.30819049835205076,
372
+ "mean_token_accuracy": 0.8977848634575353,
373
+ "num_tokens": 2282164.0,
374
+ "step": 1600
375
+ },
376
+ {
377
+ "entropy": 0.3655787402391434,
378
+ "epoch": 4.241956241956242,
379
+ "grad_norm": 0.7100041508674622,
380
+ "learning_rate": 0.00019319879684892634,
381
+ "loss": 0.29959835052490236,
382
+ "mean_token_accuracy": 0.8991208010911942,
383
+ "num_tokens": 2353213.0,
384
+ "step": 1650
385
+ },
386
+ {
387
+ "entropy": 0.3821141055226326,
388
+ "epoch": 4.370656370656371,
389
+ "grad_norm": 0.5848307013511658,
390
+ "learning_rate": 0.00018764346756040715,
391
+ "loss": 0.313802490234375,
392
+ "mean_token_accuracy": 0.895167955160141,
393
+ "num_tokens": 2425068.0,
394
+ "step": 1700
395
+ },
396
+ {
397
+ "entropy": 0.37083797007799146,
398
+ "epoch": 4.499356499356499,
399
+ "grad_norm": 0.6447024941444397,
400
+ "learning_rate": 0.00018198314039132143,
401
+ "loss": 0.30583988189697264,
402
+ "mean_token_accuracy": 0.8961733293533325,
403
+ "num_tokens": 2498321.0,
404
+ "step": 1750
405
+ },
406
+ {
407
+ "entropy": 0.3791545969247818,
408
+ "epoch": 4.628056628056628,
409
+ "grad_norm": 0.6575382351875305,
410
+ "learning_rate": 0.00017622920797738184,
411
+ "loss": 0.3088031005859375,
412
+ "mean_token_accuracy": 0.8960050916671753,
413
+ "num_tokens": 2570321.0,
414
+ "step": 1800
415
+ },
416
+ {
417
+ "entropy": 0.3946831756830215,
418
+ "epoch": 4.756756756756757,
419
+ "grad_norm": 0.5351552963256836,
420
+ "learning_rate": 0.00017039325135515207,
421
+ "loss": 0.3229162979125977,
422
+ "mean_token_accuracy": 0.8920552498102188,
423
+ "num_tokens": 2642851.0,
424
+ "step": 1850
425
+ },
426
+ {
427
+ "entropy": 0.37298239797353744,
428
+ "epoch": 4.885456885456885,
429
+ "grad_norm": 0.7624587416648865,
430
+ "learning_rate": 0.00016448701665269964,
431
+ "loss": 0.3067934799194336,
432
+ "mean_token_accuracy": 0.8951873427629471,
433
+ "num_tokens": 2715629.0,
434
+ "step": 1900
435
+ },
436
+ {
437
+ "epoch": 5.0,
438
+ "eval_entropy": 0.3738150204887095,
439
+ "eval_loss": 0.7131896615028381,
440
+ "eval_mean_token_accuracy": 0.8207752468045225,
441
+ "eval_num_tokens": 2777465.0,
442
+ "eval_runtime": 162.1244,
443
+ "eval_samples_per_second": 9.53,
444
+ "eval_steps_per_second": 1.197,
445
+ "step": 1945
446
+ },
447
+ {
448
+ "entropy": 0.377820266617669,
449
+ "epoch": 5.012870012870013,
450
+ "grad_norm": 0.41990190744400024,
451
+ "learning_rate": 0.00015852239144796624,
452
+ "loss": 0.3058685111999512,
453
+ "mean_token_accuracy": 0.8964343480389527,
454
+ "num_tokens": 2784896.0,
455
+ "step": 1950
456
+ },
457
+ {
458
+ "entropy": 0.2714502356946468,
459
+ "epoch": 5.141570141570142,
460
+ "grad_norm": 0.422568678855896,
461
+ "learning_rate": 0.00015251138084243995,
462
+ "loss": 0.2093442153930664,
463
+ "mean_token_accuracy": 0.9311346983909607,
464
+ "num_tokens": 2854374.0,
465
+ "step": 2000
466
+ },
467
+ {
468
+ "entropy": 0.268475965410471,
469
+ "epoch": 5.27027027027027,
470
+ "grad_norm": 0.6637414693832397,
471
+ "learning_rate": 0.0001464660832982852,
472
+ "loss": 0.20736080169677734,
473
+ "mean_token_accuracy": 0.9289199805259705,
474
+ "num_tokens": 2927362.0,
475
+ "step": 2050
476
+ },
477
+ {
478
+ "entropy": 0.2644876340031624,
479
+ "epoch": 5.398970398970399,
480
+ "grad_norm": 0.47317707538604736,
481
+ "learning_rate": 0.00014039866628756467,
482
+ "loss": 0.20464908599853515,
483
+ "mean_token_accuracy": 0.9300856202840805,
484
+ "num_tokens": 3000143.0,
485
+ "step": 2100
486
+ },
487
+ {
488
+ "entropy": 0.2675253136456013,
489
+ "epoch": 5.527670527670527,
490
+ "grad_norm": 0.5253982543945312,
491
+ "learning_rate": 0.00013432134180256338,
492
+ "loss": 0.21154335021972656,
493
+ "mean_token_accuracy": 0.9283734840154648,
494
+ "num_tokens": 3072561.0,
495
+ "step": 2150
496
+ },
497
+ {
498
+ "entropy": 0.27213907435536383,
499
+ "epoch": 5.656370656370656,
500
+ "grad_norm": 0.46738553047180176,
501
+ "learning_rate": 0.00012824634177650664,
502
+ "loss": 0.21339216232299804,
503
+ "mean_token_accuracy": 0.9272083270549775,
504
+ "num_tokens": 3144831.0,
505
+ "step": 2200
506
+ },
507
+ {
508
+ "entropy": 0.2785488124191761,
509
+ "epoch": 5.785070785070785,
510
+ "grad_norm": 0.4469502866268158,
511
+ "learning_rate": 0.00012218589346414205,
512
+ "loss": 0.21601097106933595,
513
+ "mean_token_accuracy": 0.9255663657188415,
514
+ "num_tokens": 3215960.0,
515
+ "step": 2250
516
+ },
517
+ {
518
+ "entropy": 0.2699935150146484,
519
+ "epoch": 5.913770913770914,
520
+ "grad_norm": 0.7359778881072998,
521
+ "learning_rate": 0.00011615219483173828,
522
+ "loss": 0.20725584030151367,
523
+ "mean_token_accuracy": 0.9286630594730377,
524
+ "num_tokens": 3287499.0,
525
+ "step": 2300
526
+ },
527
+ {
528
+ "epoch": 6.0,
529
+ "eval_entropy": 0.26417383682174783,
530
+ "eval_loss": 0.8880229592323303,
531
+ "eval_mean_token_accuracy": 0.8159987201395723,
532
+ "eval_num_tokens": 3332958.0,
533
+ "eval_runtime": 162.0991,
534
+ "eval_samples_per_second": 9.531,
535
+ "eval_steps_per_second": 1.197,
536
+ "step": 2334
537
+ },
538
+ {
539
+ "entropy": 0.24814540704693458,
540
+ "epoch": 6.041184041184041,
541
+ "grad_norm": 0.4953760802745819,
542
+ "learning_rate": 0.00011015739000603316,
543
+ "loss": 0.18749794006347656,
544
+ "mean_token_accuracy": 0.9370789509831052,
545
+ "num_tokens": 3356879.0,
546
+ "step": 2350
547
+ },
548
+ {
549
+ "entropy": 0.19976271741092205,
550
+ "epoch": 6.1698841698841695,
551
+ "grad_norm": 0.4834803342819214,
552
+ "learning_rate": 0.00010421354483154553,
553
+ "loss": 0.14283526420593262,
554
+ "mean_token_accuracy": 0.9521516615152359,
555
+ "num_tokens": 3427587.0,
556
+ "step": 2400
557
+ },
558
+ {
559
+ "entropy": 0.2060488449037075,
560
+ "epoch": 6.298584298584299,
561
+ "grad_norm": 0.4888673722743988,
562
+ "learning_rate": 9.8332622585447e-05,
563
+ "loss": 0.14414511680603026,
564
+ "mean_token_accuracy": 0.9510996866226197,
565
+ "num_tokens": 3498688.0,
566
+ "step": 2450
567
+ },
568
+ {
569
+ "entropy": 0.2059111550450325,
570
+ "epoch": 6.427284427284428,
571
+ "grad_norm": 0.4064404368400574,
572
+ "learning_rate": 9.252645989887253e-05,
573
+ "loss": 0.14820143699645996,
574
+ "mean_token_accuracy": 0.9507584601640702,
575
+ "num_tokens": 3566137.0,
576
+ "step": 2500
577
+ },
578
+ {
579
+ "entropy": 0.19700154662132263,
580
+ "epoch": 6.555984555984556,
581
+ "grad_norm": 0.467965304851532,
582
+ "learning_rate": 8.680674293313417e-05,
583
+ "loss": 0.14303470611572267,
584
+ "mean_token_accuracy": 0.9515972435474396,
585
+ "num_tokens": 3639573.0,
586
+ "step": 2550
587
+ },
588
+ {
589
+ "entropy": 0.20180423602461814,
590
+ "epoch": 6.684684684684685,
591
+ "grad_norm": 0.36836138367652893,
592
+ "learning_rate": 8.118498385878736e-05,
593
+ "loss": 0.14280882835388184,
594
+ "mean_token_accuracy": 0.9515993863344192,
595
+ "num_tokens": 3710433.0,
596
+ "step": 2600
597
+ },
598
+ {
599
+ "entropy": 0.20024395987391472,
600
+ "epoch": 6.813384813384813,
601
+ "grad_norm": 0.38375866413116455,
602
+ "learning_rate": 7.567249768489171e-05,
603
+ "loss": 0.1427844524383545,
604
+ "mean_token_accuracy": 0.9524166631698608,
605
+ "num_tokens": 3781550.0,
606
+ "step": 2650
607
+ },
608
+ {
609
+ "entropy": 0.19561587080359458,
610
+ "epoch": 6.942084942084942,
611
+ "grad_norm": 0.41185441613197327,
612
+ "learning_rate": 7.028037948510187e-05,
613
+ "loss": 0.13993803024291993,
614
+ "mean_token_accuracy": 0.9522478264570237,
615
+ "num_tokens": 3854952.0,
616
+ "step": 2700
617
+ },
618
+ {
619
+ "epoch": 7.0,
620
+ "eval_entropy": 0.19501976062034823,
621
+ "eval_loss": 1.0653952360153198,
622
+ "eval_mean_token_accuracy": 0.8205490803595671,
623
+ "eval_num_tokens": 3888451.0,
624
+ "eval_runtime": 161.8533,
625
+ "eval_samples_per_second": 9.546,
626
+ "eval_steps_per_second": 1.199,
627
+ "step": 2723
628
+ },
629
+ {
630
+ "entropy": 0.17789882526855277,
631
+ "epoch": 7.06949806949807,
632
+ "grad_norm": 0.41413992643356323,
633
+ "learning_rate": 6.50194820664261e-05,
634
+ "loss": 0.12078390121459961,
635
+ "mean_token_accuracy": 0.9589925727458916,
636
+ "num_tokens": 3928354.0,
637
+ "step": 2750
638
+ },
639
+ {
640
+ "entropy": 0.16781829454004765,
641
+ "epoch": 7.198198198198198,
642
+ "grad_norm": 0.25806066393852234,
643
+ "learning_rate": 5.990039412559906e-05,
644
+ "loss": 0.10963023185729981,
645
+ "mean_token_accuracy": 0.9617267113924026,
646
+ "num_tokens": 4000113.0,
647
+ "step": 2800
648
+ },
649
+ {
650
+ "entropy": 0.1649068508297205,
651
+ "epoch": 7.326898326898327,
652
+ "grad_norm": 0.27411890029907227,
653
+ "learning_rate": 5.493341893703393e-05,
654
+ "loss": 0.11152458190917969,
655
+ "mean_token_accuracy": 0.9620639663934708,
656
+ "num_tokens": 4071032.0,
657
+ "step": 2850
658
+ },
659
+ {
660
+ "entropy": 0.161333369910717,
661
+ "epoch": 7.455598455598455,
662
+ "grad_norm": 0.24944494664669037,
663
+ "learning_rate": 5.0128553615248396e-05,
664
+ "loss": 0.1094522476196289,
665
+ "mean_token_accuracy": 0.962428919672966,
666
+ "num_tokens": 4143616.0,
667
+ "step": 2900
668
+ },
669
+ {
670
+ "entropy": 0.15613057143986225,
671
+ "epoch": 7.584298584298584,
672
+ "grad_norm": 0.1455036848783493,
673
+ "learning_rate": 4.549546899350423e-05,
674
+ "loss": 0.11092090606689453,
675
+ "mean_token_accuracy": 0.9620462411642074,
676
+ "num_tokens": 4215664.0,
677
+ "step": 2950
678
+ },
679
+ {
680
+ "entropy": 0.1631234459578991,
681
+ "epoch": 7.712998712998713,
682
+ "grad_norm": 0.2129560261964798,
683
+ "learning_rate": 4.104349015915862e-05,
684
+ "loss": 0.1141857624053955,
685
+ "mean_token_accuracy": 0.9613765001296997,
686
+ "num_tokens": 4286387.0,
687
+ "step": 3000
688
+ },
689
+ {
690
+ "entropy": 0.1680422095954418,
691
+ "epoch": 7.841698841698841,
692
+ "grad_norm": 0.24886097013950348,
693
+ "learning_rate": 3.678157768490372e-05,
694
+ "loss": 0.11513191223144531,
695
+ "mean_token_accuracy": 0.9615794748067856,
696
+ "num_tokens": 4355875.0,
697
+ "step": 3050
698
+ },
699
+ {
700
+ "entropy": 0.16354035697877406,
701
+ "epoch": 7.97039897039897,
702
+ "grad_norm": 0.27600204944610596,
703
+ "learning_rate": 3.27183095936714e-05,
704
+ "loss": 0.1118631362915039,
705
+ "mean_token_accuracy": 0.9623224419355393,
706
+ "num_tokens": 4427088.0,
707
+ "step": 3100
708
+ },
709
+ {
710
+ "epoch": 8.0,
711
+ "eval_entropy": 0.16586771200305409,
712
+ "eval_loss": 1.204746961593628,
713
+ "eval_mean_token_accuracy": 0.8229975042883882,
714
+ "eval_num_tokens": 4443944.0,
715
+ "eval_runtime": 162.0251,
716
+ "eval_samples_per_second": 9.536,
717
+ "eval_steps_per_second": 1.197,
718
+ "step": 3112
719
+ }
720
+ ],
721
+ "logging_steps": 50,
722
+ "max_steps": 3890,
723
+ "num_input_tokens_seen": 0,
724
+ "num_train_epochs": 10,
725
+ "save_steps": 500,
726
+ "stateful_callbacks": {
727
+ "TrainerControl": {
728
+ "args": {
729
+ "should_epoch_stop": false,
730
+ "should_evaluate": false,
731
+ "should_log": false,
732
+ "should_save": true,
733
+ "should_training_stop": false
734
+ },
735
+ "attributes": {}
736
+ }
737
+ },
738
+ "total_flos": 7.445001770940518e+17,
739
+ "train_batch_size": 8,
740
+ "trial_name": null,
741
+ "trial_params": null
742
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/README.md ADDED
@@ -0,0 +1,209 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen3-14B-Base
3
+ library_name: peft
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - base_model:adapter:Qwen/Qwen3-14B-Base
7
+ - lora
8
+ - sft
9
+ - transformers
10
+ - trl
11
+ ---
12
+
13
+ # Model Card for Model ID
14
+
15
+ <!-- Provide a quick summary of what the model is/does. -->
16
+
17
+
18
+
19
+ ## Model Details
20
+
21
+ ### Model Description
22
+
23
+ <!-- Provide a longer summary of what this model is. -->
24
+
25
+
26
+
27
+ - **Developed by:** [More Information Needed]
28
+ - **Funded by [optional]:** [More Information Needed]
29
+ - **Shared by [optional]:** [More Information Needed]
30
+ - **Model type:** [More Information Needed]
31
+ - **Language(s) (NLP):** [More Information Needed]
32
+ - **License:** [More Information Needed]
33
+ - **Finetuned from model [optional]:** [More Information Needed]
34
+
35
+ ### Model Sources [optional]
36
+
37
+ <!-- Provide the basic links for the model. -->
38
+
39
+ - **Repository:** [More Information Needed]
40
+ - **Paper [optional]:** [More Information Needed]
41
+ - **Demo [optional]:** [More Information Needed]
42
+
43
+ ## Uses
44
+
45
+ <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
46
+
47
+ ### Direct Use
48
+
49
+ <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
50
+
51
+ [More Information Needed]
52
+
53
+ ### Downstream Use [optional]
54
+
55
+ <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
56
+
57
+ [More Information Needed]
58
+
59
+ ### Out-of-Scope Use
60
+
61
+ <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
62
+
63
+ [More Information Needed]
64
+
65
+ ## Bias, Risks, and Limitations
66
+
67
+ <!-- This section is meant to convey both technical and sociotechnical limitations. -->
68
+
69
+ [More Information Needed]
70
+
71
+ ### Recommendations
72
+
73
+ <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
74
+
75
+ Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
76
+
77
+ ## How to Get Started with the Model
78
+
79
+ Use the code below to get started with the model.
80
+
81
+ [More Information Needed]
82
+
83
+ ## Training Details
84
+
85
+ ### Training Data
86
+
87
+ <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
88
+
89
+ [More Information Needed]
90
+
91
+ ### Training Procedure
92
+
93
+ <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
94
+
95
+ #### Preprocessing [optional]
96
+
97
+ [More Information Needed]
98
+
99
+
100
+ #### Training Hyperparameters
101
+
102
+ - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
103
+
104
+ #### Speeds, Sizes, Times [optional]
105
+
106
+ <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
107
+
108
+ [More Information Needed]
109
+
110
+ ## Evaluation
111
+
112
+ <!-- This section describes the evaluation protocols and provides the results. -->
113
+
114
+ ### Testing Data, Factors & Metrics
115
+
116
+ #### Testing Data
117
+
118
+ <!-- This should link to a Dataset Card if possible. -->
119
+
120
+ [More Information Needed]
121
+
122
+ #### Factors
123
+
124
+ <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
125
+
126
+ [More Information Needed]
127
+
128
+ #### Metrics
129
+
130
+ <!-- These are the evaluation metrics being used, ideally with a description of why. -->
131
+
132
+ [More Information Needed]
133
+
134
+ ### Results
135
+
136
+ [More Information Needed]
137
+
138
+ #### Summary
139
+
140
+
141
+
142
+ ## Model Examination [optional]
143
+
144
+ <!-- Relevant interpretability work for the model goes here -->
145
+
146
+ [More Information Needed]
147
+
148
+ ## Environmental Impact
149
+
150
+ <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
151
+
152
+ Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
153
+
154
+ - **Hardware Type:** [More Information Needed]
155
+ - **Hours used:** [More Information Needed]
156
+ - **Cloud Provider:** [More Information Needed]
157
+ - **Compute Region:** [More Information Needed]
158
+ - **Carbon Emitted:** [More Information Needed]
159
+
160
+ ## Technical Specifications [optional]
161
+
162
+ ### Model Architecture and Objective
163
+
164
+ [More Information Needed]
165
+
166
+ ### Compute Infrastructure
167
+
168
+ [More Information Needed]
169
+
170
+ #### Hardware
171
+
172
+ [More Information Needed]
173
+
174
+ #### Software
175
+
176
+ [More Information Needed]
177
+
178
+ ## Citation [optional]
179
+
180
+ <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
181
+
182
+ **BibTeX:**
183
+
184
+ [More Information Needed]
185
+
186
+ **APA:**
187
+
188
+ [More Information Needed]
189
+
190
+ ## Glossary [optional]
191
+
192
+ <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
193
+
194
+ [More Information Needed]
195
+
196
+ ## More Information [optional]
197
+
198
+ [More Information Needed]
199
+
200
+ ## Model Card Authors [optional]
201
+
202
+ [More Information Needed]
203
+
204
+ ## Model Card Contact
205
+
206
+ [More Information Needed]
207
+ ### Framework versions
208
+
209
+ - PEFT 0.19.1
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/adapter_config.json ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "Qwen/Qwen3-14B-Base",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "layer_replication": null,
16
+ "layers_pattern": null,
17
+ "layers_to_transform": null,
18
+ "loftq_config": {},
19
+ "lora_alpha": 256,
20
+ "lora_bias": false,
21
+ "lora_dropout": 0.00237968804112545,
22
+ "lora_ga_config": null,
23
+ "megatron_config": null,
24
+ "megatron_core": "megatron.core",
25
+ "modules_to_save": null,
26
+ "peft_type": "LORA",
27
+ "peft_version": "0.19.1",
28
+ "qalora_group_size": 16,
29
+ "r": 128,
30
+ "rank_pattern": {},
31
+ "revision": null,
32
+ "target_modules": [
33
+ "gate_proj",
34
+ "q_proj",
35
+ "up_proj",
36
+ "o_proj",
37
+ "v_proj",
38
+ "k_proj",
39
+ "down_proj"
40
+ ],
41
+ "target_parameters": null,
42
+ "task_type": "CAUSAL_LM",
43
+ "trainable_token_indices": null,
44
+ "use_bdlora": null,
45
+ "use_dora": false,
46
+ "use_qalora": false,
47
+ "use_rslora": false
48
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/chat_template.jinja ADDED
@@ -0,0 +1,85 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- if tools %}
2
+ {{- '<|im_start|>system\n' }}
3
+ {%- if messages[0].role == 'system' %}
4
+ {{- messages[0].content + '\n\n' }}
5
+ {%- endif %}
6
+ {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
7
+ {%- for tool in tools %}
8
+ {{- "\n" }}
9
+ {{- tool | tojson }}
10
+ {%- endfor %}
11
+ {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
12
+ {%- else %}
13
+ {%- if messages[0].role == 'system' %}
14
+ {{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
15
+ {%- endif %}
16
+ {%- endif %}
17
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
18
+ {%- for message in messages[::-1] %}
19
+ {%- set index = (messages|length - 1) - loop.index0 %}
20
+ {%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
21
+ {%- set ns.multi_step_tool = false %}
22
+ {%- set ns.last_query_index = index %}
23
+ {%- endif %}
24
+ {%- endfor %}
25
+ {%- for message in messages %}
26
+ {%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
27
+ {{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
28
+ {%- elif message.role == "assistant" %}
29
+ {%- set content = message.content %}
30
+ {%- set reasoning_content = '' %}
31
+ {%- if message.reasoning_content is defined and message.reasoning_content is not none %}
32
+ {%- set reasoning_content = message.reasoning_content %}
33
+ {%- else %}
34
+ {%- if '</think>' in message.content %}
35
+ {%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
36
+ {%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
37
+ {%- endif %}
38
+ {%- endif %}
39
+ {%- if loop.index0 > ns.last_query_index %}
40
+ {%- if loop.last or (not loop.last and reasoning_content) %}
41
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
42
+ {%- else %}
43
+ {{- '<|im_start|>' + message.role + '\n' + content }}
44
+ {%- endif %}
45
+ {%- else %}
46
+ {{- '<|im_start|>' + message.role + '\n' + content }}
47
+ {%- endif %}
48
+ {%- if message.tool_calls %}
49
+ {%- for tool_call in message.tool_calls %}
50
+ {%- if (loop.first and content) or (not loop.first) %}
51
+ {{- '\n' }}
52
+ {%- endif %}
53
+ {%- if tool_call.function %}
54
+ {%- set tool_call = tool_call.function %}
55
+ {%- endif %}
56
+ {{- '<tool_call>\n{"name": "' }}
57
+ {{- tool_call.name }}
58
+ {{- '", "arguments": ' }}
59
+ {%- if tool_call.arguments is string %}
60
+ {{- tool_call.arguments }}
61
+ {%- else %}
62
+ {{- tool_call.arguments | tojson }}
63
+ {%- endif %}
64
+ {{- '}\n</tool_call>' }}
65
+ {%- endfor %}
66
+ {%- endif %}
67
+ {{- '<|im_end|>\n' }}
68
+ {%- elif message.role == "tool" %}
69
+ {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
70
+ {{- '<|im_start|>user' }}
71
+ {%- endif %}
72
+ {{- '\n<tool_response>\n' }}
73
+ {{- message.content }}
74
+ {{- '\n</tool_response>' }}
75
+ {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
76
+ {{- '<|im_end|>\n' }}
77
+ {%- endif %}
78
+ {%- endif %}
79
+ {%- endfor %}
80
+ {%- if add_generation_prompt %}
81
+ {{- '<|im_start|>assistant\n' }}
82
+ {%- if enable_thinking is defined and enable_thinking is false %}
83
+ {{- '<think>\n\n</think>\n\n' }}
84
+ {%- endif %}
85
+ {%- endif %}
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/tokenizer_config.json ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "backend": "tokenizers",
4
+ "bos_token": null,
5
+ "clean_up_tokenization_spaces": false,
6
+ "eos_token": "<|endoftext|>",
7
+ "errors": "replace",
8
+ "extra_special_tokens": [
9
+ "<|im_start|>",
10
+ "<|im_end|>",
11
+ "<|object_ref_start|>",
12
+ "<|object_ref_end|>",
13
+ "<|box_start|>",
14
+ "<|box_end|>",
15
+ "<|quad_start|>",
16
+ "<|quad_end|>",
17
+ "<|vision_start|>",
18
+ "<|vision_end|>",
19
+ "<|vision_pad|>",
20
+ "<|image_pad|>",
21
+ "<|video_pad|>"
22
+ ],
23
+ "is_local": false,
24
+ "model_max_length": 131072,
25
+ "pad_token": "<|endoftext|>",
26
+ "split_special_tokens": false,
27
+ "tokenizer_class": "Qwen2Tokenizer",
28
+ "unk_token": null
29
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/trainer_state.json ADDED
@@ -0,0 +1,833 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "best_global_step": null,
3
+ "best_metric": null,
4
+ "best_model_checkpoint": null,
5
+ "epoch": 9.0,
6
+ "eval_steps": 500,
7
+ "global_step": 3501,
8
+ "is_hyper_param_search": false,
9
+ "is_local_process_zero": true,
10
+ "is_world_process_zero": true,
11
+ "log_history": [
12
+ {
13
+ "entropy": 1.6110110306739807,
14
+ "epoch": 0.1287001287001287,
15
+ "grad_norm": 0.8611119389533997,
16
+ "learning_rate": 3.4130257962133866e-05,
17
+ "loss": 1.53956298828125,
18
+ "mean_token_accuracy": 0.6652342769503593,
19
+ "num_tokens": 73407.0,
20
+ "step": 50
21
+ },
22
+ {
23
+ "entropy": 0.8662820833921433,
24
+ "epoch": 0.2574002574002574,
25
+ "grad_norm": 0.7405035495758057,
26
+ "learning_rate": 6.895705180104598e-05,
27
+ "loss": 0.7929539489746094,
28
+ "mean_token_accuracy": 0.7786347842216492,
29
+ "num_tokens": 143994.0,
30
+ "step": 100
31
+ },
32
+ {
33
+ "entropy": 0.7912819278240204,
34
+ "epoch": 0.3861003861003861,
35
+ "grad_norm": 0.6105485558509827,
36
+ "learning_rate": 0.00010378384563995809,
37
+ "loss": 0.7236511993408203,
38
+ "mean_token_accuracy": 0.7925467795133591,
39
+ "num_tokens": 216171.0,
40
+ "step": 150
41
+ },
42
+ {
43
+ "entropy": 0.7777709531784057,
44
+ "epoch": 0.5148005148005148,
45
+ "grad_norm": 0.5781793594360352,
46
+ "learning_rate": 0.0001386106394788702,
47
+ "loss": 0.7024919891357422,
48
+ "mean_token_accuracy": 0.7969876372814179,
49
+ "num_tokens": 284702.0,
50
+ "step": 200
51
+ },
52
+ {
53
+ "entropy": 0.7632609683275223,
54
+ "epoch": 0.6435006435006435,
55
+ "grad_norm": 0.6553444862365723,
56
+ "learning_rate": 0.00017343743331778232,
57
+ "loss": 0.7000718688964844,
58
+ "mean_token_accuracy": 0.7994742071628571,
59
+ "num_tokens": 356393.0,
60
+ "step": 250
61
+ },
62
+ {
63
+ "entropy": 0.754165632724762,
64
+ "epoch": 0.7722007722007722,
65
+ "grad_norm": 0.634222149848938,
66
+ "learning_rate": 0.00020826422715669444,
67
+ "loss": 0.6909049987792969,
68
+ "mean_token_accuracy": 0.8016809666156769,
69
+ "num_tokens": 426916.0,
70
+ "step": 300
71
+ },
72
+ {
73
+ "entropy": 0.7530711203813553,
74
+ "epoch": 0.9009009009009009,
75
+ "grad_norm": 0.7020156383514404,
76
+ "learning_rate": 0.00024309102099560653,
77
+ "loss": 0.69554931640625,
78
+ "mean_token_accuracy": 0.8002807641029358,
79
+ "num_tokens": 499603.0,
80
+ "step": 350
81
+ },
82
+ {
83
+ "epoch": 1.0,
84
+ "eval_entropy": 0.5728044164242204,
85
+ "eval_loss": 0.6785851716995239,
86
+ "eval_mean_token_accuracy": 0.798433246993527,
87
+ "eval_num_tokens": 555493.0,
88
+ "eval_runtime": 162.4129,
89
+ "eval_samples_per_second": 9.513,
90
+ "eval_steps_per_second": 1.194,
91
+ "step": 389
92
+ },
93
+ {
94
+ "entropy": 0.7489313946829902,
95
+ "epoch": 1.0283140283140284,
96
+ "grad_norm": 0.7532815933227539,
97
+ "learning_rate": 0.0002709470016827303,
98
+ "loss": 0.6856581878662109,
99
+ "mean_token_accuracy": 0.8023027079273956,
100
+ "num_tokens": 570813.0,
101
+ "step": 400
102
+ },
103
+ {
104
+ "entropy": 0.7238409864902496,
105
+ "epoch": 1.157014157014157,
106
+ "grad_norm": 1.2016215324401855,
107
+ "learning_rate": 0.0002707561443541359,
108
+ "loss": 0.6699818420410156,
109
+ "mean_token_accuracy": 0.8070302194356919,
110
+ "num_tokens": 642956.0,
111
+ "step": 450
112
+ },
113
+ {
114
+ "entropy": 0.7388938587903976,
115
+ "epoch": 1.2857142857142856,
116
+ "grad_norm": 0.7279272079467773,
117
+ "learning_rate": 0.0002702930068622498,
118
+ "loss": 0.6728517150878907,
119
+ "mean_token_accuracy": 0.8049580943584442,
120
+ "num_tokens": 714498.0,
121
+ "step": 500
122
+ },
123
+ {
124
+ "entropy": 0.7284485149383545,
125
+ "epoch": 1.4144144144144144,
126
+ "grad_norm": 0.8563987016677856,
127
+ "learning_rate": 0.0002695585213716931,
128
+ "loss": 0.6657986450195312,
129
+ "mean_token_accuracy": 0.8085588800907135,
130
+ "num_tokens": 785595.0,
131
+ "step": 550
132
+ },
133
+ {
134
+ "entropy": 0.6960959500074386,
135
+ "epoch": 1.5431145431145432,
136
+ "grad_norm": 0.5145474672317505,
137
+ "learning_rate": 0.0002685541661937683,
138
+ "loss": 0.6358638763427734,
139
+ "mean_token_accuracy": 0.8131621342897415,
140
+ "num_tokens": 857551.0,
141
+ "step": 600
142
+ },
143
+ {
144
+ "entropy": 0.7220414417982102,
145
+ "epoch": 1.6718146718146718,
146
+ "grad_norm": 0.9381059408187866,
147
+ "learning_rate": 0.00026728196281103746,
148
+ "loss": 0.6531407928466797,
149
+ "mean_token_accuracy": 0.811619822382927,
150
+ "num_tokens": 926864.0,
151
+ "step": 650
152
+ },
153
+ {
154
+ "entropy": 0.6955459499359131,
155
+ "epoch": 1.8005148005148004,
156
+ "grad_norm": 0.6115108728408813,
157
+ "learning_rate": 0.0002657444718086503,
158
+ "loss": 0.6269588088989257,
159
+ "mean_token_accuracy": 0.8155373805761337,
160
+ "num_tokens": 998824.0,
161
+ "step": 700
162
+ },
163
+ {
164
+ "entropy": 0.6929812705516816,
165
+ "epoch": 1.9292149292149292,
166
+ "grad_norm": 0.749489426612854,
167
+ "learning_rate": 0.0002639447877206115,
168
+ "loss": 0.629054069519043,
169
+ "mean_token_accuracy": 0.8183021235466004,
170
+ "num_tokens": 1069332.0,
171
+ "step": 750
172
+ },
173
+ {
174
+ "epoch": 2.0,
175
+ "eval_entropy": 0.5833694102223387,
176
+ "eval_loss": 0.6328718662261963,
177
+ "eval_mean_token_accuracy": 0.8159044071571114,
178
+ "eval_num_tokens": 1110986.0,
179
+ "eval_runtime": 161.885,
180
+ "eval_samples_per_second": 9.544,
181
+ "eval_steps_per_second": 1.198,
182
+ "step": 778
183
+ },
184
+ {
185
+ "entropy": 0.647241060480927,
186
+ "epoch": 2.056628056628057,
187
+ "grad_norm": 0.7070767879486084,
188
+ "learning_rate": 0.00026188653280135975,
189
+ "loss": 0.5823922348022461,
190
+ "mean_token_accuracy": 0.8265365301960647,
191
+ "num_tokens": 1141195.0,
192
+ "step": 800
193
+ },
194
+ {
195
+ "entropy": 0.5995378407835961,
196
+ "epoch": 2.1853281853281854,
197
+ "grad_norm": 0.8090486526489258,
198
+ "learning_rate": 0.0002595738497351955,
199
+ "loss": 0.5325597763061524,
200
+ "mean_token_accuracy": 0.8369336777925491,
201
+ "num_tokens": 1210708.0,
202
+ "step": 850
203
+ },
204
+ {
205
+ "entropy": 0.6043145382404327,
206
+ "epoch": 2.314028314028314,
207
+ "grad_norm": 0.8279913067817688,
208
+ "learning_rate": 0.00025701139329823054,
209
+ "loss": 0.5414446258544922,
210
+ "mean_token_accuracy": 0.8361396533250809,
211
+ "num_tokens": 1283441.0,
212
+ "step": 900
213
+ },
214
+ {
215
+ "entropy": 0.5953224584460258,
216
+ "epoch": 2.4427284427284426,
217
+ "grad_norm": 0.6075023412704468,
218
+ "learning_rate": 0.00025420432098964183,
219
+ "loss": 0.536654167175293,
220
+ "mean_token_accuracy": 0.8340093129873276,
221
+ "num_tokens": 1356418.0,
222
+ "step": 950
223
+ },
224
+ {
225
+ "entropy": 0.5998479858040809,
226
+ "epoch": 2.571428571428571,
227
+ "grad_norm": 1.0311471223831177,
228
+ "learning_rate": 0.0002511582826510862,
229
+ "loss": 0.5372924423217773,
230
+ "mean_token_accuracy": 0.8366045409440994,
231
+ "num_tokens": 1427797.0,
232
+ "step": 1000
233
+ },
234
+ {
235
+ "entropy": 0.5980158120393753,
236
+ "epoch": 2.7001287001287,
237
+ "grad_norm": 0.5971426367759705,
238
+ "learning_rate": 0.0002478794090951689,
239
+ "loss": 0.5392885208129883,
240
+ "mean_token_accuracy": 0.8347727072238922,
241
+ "num_tokens": 1498082.0,
242
+ "step": 1050
243
+ },
244
+ {
245
+ "entropy": 0.5990680930018425,
246
+ "epoch": 2.828828828828829,
247
+ "grad_norm": 0.5662627220153809,
248
+ "learning_rate": 0.0002443742997658538,
249
+ "loss": 0.5360498428344727,
250
+ "mean_token_accuracy": 0.8371847170591354,
251
+ "num_tokens": 1568798.0,
252
+ "step": 1100
253
+ },
254
+ {
255
+ "entropy": 0.5957039377093315,
256
+ "epoch": 2.9575289575289574,
257
+ "grad_norm": 0.5043798685073853,
258
+ "learning_rate": 0.00024065000945565205,
259
+ "loss": 0.5342231369018555,
260
+ "mean_token_accuracy": 0.8380735236406326,
261
+ "num_tokens": 1643449.0,
262
+ "step": 1150
263
+ },
264
+ {
265
+ "epoch": 3.0,
266
+ "eval_entropy": 0.5346978819861854,
267
+ "eval_loss": 0.6255015134811401,
268
+ "eval_mean_token_accuracy": 0.8202014476368108,
269
+ "eval_num_tokens": 1666479.0,
270
+ "eval_runtime": 161.6098,
271
+ "eval_samples_per_second": 9.56,
272
+ "eval_steps_per_second": 1.2,
273
+ "step": 1167
274
+ },
275
+ {
276
+ "entropy": 0.5212209137401196,
277
+ "epoch": 3.0849420849420848,
278
+ "grad_norm": 0.7095440626144409,
279
+ "learning_rate": 0.00023671403410632178,
280
+ "loss": 0.45311901092529294,
281
+ "mean_token_accuracy": 0.856362871449403,
282
+ "num_tokens": 1713536.0,
283
+ "step": 1200
284
+ },
285
+ {
286
+ "entropy": 0.4701593083143234,
287
+ "epoch": 3.213642213642214,
288
+ "grad_norm": 0.6408083438873291,
289
+ "learning_rate": 0.0002325742957216607,
290
+ "loss": 0.39916397094726563,
291
+ "mean_token_accuracy": 0.8698061722517013,
292
+ "num_tokens": 1785609.0,
293
+ "step": 1250
294
+ },
295
+ {
296
+ "entropy": 0.4766591975092888,
297
+ "epoch": 3.3423423423423424,
298
+ "grad_norm": 0.6415093541145325,
299
+ "learning_rate": 0.0002282391264227552,
300
+ "loss": 0.4116698455810547,
301
+ "mean_token_accuracy": 0.8679435575008392,
302
+ "num_tokens": 1858651.0,
303
+ "step": 1300
304
+ },
305
+ {
306
+ "entropy": 0.4937947469949722,
307
+ "epoch": 3.471042471042471,
308
+ "grad_norm": 0.6549825072288513,
309
+ "learning_rate": 0.00022371725167778054,
310
+ "loss": 0.4296376037597656,
311
+ "mean_token_accuracy": 0.8609357953071595,
312
+ "num_tokens": 1928692.0,
313
+ "step": 1350
314
+ },
315
+ {
316
+ "entropy": 0.4920153194665909,
317
+ "epoch": 3.5997425997425996,
318
+ "grad_norm": 0.6452126502990723,
319
+ "learning_rate": 0.00021901777274010406,
320
+ "loss": 0.4307489013671875,
321
+ "mean_token_accuracy": 0.8606827831268311,
322
+ "num_tokens": 1998668.0,
323
+ "step": 1400
324
+ },
325
+ {
326
+ "entropy": 0.490042342543602,
327
+ "epoch": 3.7284427284427286,
328
+ "grad_norm": 0.5727734565734863,
329
+ "learning_rate": 0.0002141501483300395,
330
+ "loss": 0.4295254135131836,
331
+ "mean_token_accuracy": 0.8616545403003693,
332
+ "num_tokens": 2072809.0,
333
+ "step": 1450
334
+ },
335
+ {
336
+ "entropy": 0.49857193052768706,
337
+ "epoch": 3.857142857142857,
338
+ "grad_norm": 0.7732954025268555,
339
+ "learning_rate": 0.00020912417559712133,
340
+ "loss": 0.4289303207397461,
341
+ "mean_token_accuracy": 0.8616443765163422,
342
+ "num_tokens": 2142475.0,
343
+ "step": 1500
344
+ },
345
+ {
346
+ "entropy": 0.4746784272789955,
347
+ "epoch": 3.985842985842986,
348
+ "grad_norm": 0.5791187882423401,
349
+ "learning_rate": 0.00020394997040121726,
350
+ "loss": 0.4180263900756836,
351
+ "mean_token_accuracy": 0.866080379486084,
352
+ "num_tokens": 2214044.0,
353
+ "step": 1550
354
+ },
355
+ {
356
+ "epoch": 4.0,
357
+ "eval_entropy": 0.4793926059585257,
358
+ "eval_loss": 0.627627968788147,
359
+ "eval_mean_token_accuracy": 0.8233870095813397,
360
+ "eval_num_tokens": 2221972.0,
361
+ "eval_runtime": 162.0225,
362
+ "eval_samples_per_second": 9.536,
363
+ "eval_steps_per_second": 1.197,
364
+ "step": 1556
365
+ },
366
+ {
367
+ "entropy": 0.38195489000792454,
368
+ "epoch": 4.113256113256114,
369
+ "grad_norm": 0.6002617478370667,
370
+ "learning_rate": 0.0001986379469521669,
371
+ "loss": 0.30819049835205076,
372
+ "mean_token_accuracy": 0.8977848634575353,
373
+ "num_tokens": 2282164.0,
374
+ "step": 1600
375
+ },
376
+ {
377
+ "entropy": 0.3655787402391434,
378
+ "epoch": 4.241956241956242,
379
+ "grad_norm": 0.7100041508674622,
380
+ "learning_rate": 0.00019319879684892634,
381
+ "loss": 0.29959835052490236,
382
+ "mean_token_accuracy": 0.8991208010911942,
383
+ "num_tokens": 2353213.0,
384
+ "step": 1650
385
+ },
386
+ {
387
+ "entropy": 0.3821141055226326,
388
+ "epoch": 4.370656370656371,
389
+ "grad_norm": 0.5848307013511658,
390
+ "learning_rate": 0.00018764346756040715,
391
+ "loss": 0.313802490234375,
392
+ "mean_token_accuracy": 0.895167955160141,
393
+ "num_tokens": 2425068.0,
394
+ "step": 1700
395
+ },
396
+ {
397
+ "entropy": 0.37083797007799146,
398
+ "epoch": 4.499356499356499,
399
+ "grad_norm": 0.6447024941444397,
400
+ "learning_rate": 0.00018198314039132143,
401
+ "loss": 0.30583988189697264,
402
+ "mean_token_accuracy": 0.8961733293533325,
403
+ "num_tokens": 2498321.0,
404
+ "step": 1750
405
+ },
406
+ {
407
+ "entropy": 0.3791545969247818,
408
+ "epoch": 4.628056628056628,
409
+ "grad_norm": 0.6575382351875305,
410
+ "learning_rate": 0.00017622920797738184,
411
+ "loss": 0.3088031005859375,
412
+ "mean_token_accuracy": 0.8960050916671753,
413
+ "num_tokens": 2570321.0,
414
+ "step": 1800
415
+ },
416
+ {
417
+ "entropy": 0.3946831756830215,
418
+ "epoch": 4.756756756756757,
419
+ "grad_norm": 0.5351552963256836,
420
+ "learning_rate": 0.00017039325135515207,
421
+ "loss": 0.3229162979125977,
422
+ "mean_token_accuracy": 0.8920552498102188,
423
+ "num_tokens": 2642851.0,
424
+ "step": 1850
425
+ },
426
+ {
427
+ "entropy": 0.37298239797353744,
428
+ "epoch": 4.885456885456885,
429
+ "grad_norm": 0.7624587416648865,
430
+ "learning_rate": 0.00016448701665269964,
431
+ "loss": 0.3067934799194336,
432
+ "mean_token_accuracy": 0.8951873427629471,
433
+ "num_tokens": 2715629.0,
434
+ "step": 1900
435
+ },
436
+ {
437
+ "epoch": 5.0,
438
+ "eval_entropy": 0.3738150204887095,
439
+ "eval_loss": 0.7131896615028381,
440
+ "eval_mean_token_accuracy": 0.8207752468045225,
441
+ "eval_num_tokens": 2777465.0,
442
+ "eval_runtime": 162.1244,
443
+ "eval_samples_per_second": 9.53,
444
+ "eval_steps_per_second": 1.197,
445
+ "step": 1945
446
+ },
447
+ {
448
+ "entropy": 0.377820266617669,
449
+ "epoch": 5.012870012870013,
450
+ "grad_norm": 0.41990190744400024,
451
+ "learning_rate": 0.00015852239144796624,
452
+ "loss": 0.3058685111999512,
453
+ "mean_token_accuracy": 0.8964343480389527,
454
+ "num_tokens": 2784896.0,
455
+ "step": 1950
456
+ },
457
+ {
458
+ "entropy": 0.2714502356946468,
459
+ "epoch": 5.141570141570142,
460
+ "grad_norm": 0.422568678855896,
461
+ "learning_rate": 0.00015251138084243995,
462
+ "loss": 0.2093442153930664,
463
+ "mean_token_accuracy": 0.9311346983909607,
464
+ "num_tokens": 2854374.0,
465
+ "step": 2000
466
+ },
467
+ {
468
+ "entropy": 0.268475965410471,
469
+ "epoch": 5.27027027027027,
470
+ "grad_norm": 0.6637414693832397,
471
+ "learning_rate": 0.0001464660832982852,
472
+ "loss": 0.20736080169677734,
473
+ "mean_token_accuracy": 0.9289199805259705,
474
+ "num_tokens": 2927362.0,
475
+ "step": 2050
476
+ },
477
+ {
478
+ "entropy": 0.2644876340031624,
479
+ "epoch": 5.398970398970399,
480
+ "grad_norm": 0.47317707538604736,
481
+ "learning_rate": 0.00014039866628756467,
482
+ "loss": 0.20464908599853515,
483
+ "mean_token_accuracy": 0.9300856202840805,
484
+ "num_tokens": 3000143.0,
485
+ "step": 2100
486
+ },
487
+ {
488
+ "entropy": 0.2675253136456013,
489
+ "epoch": 5.527670527670527,
490
+ "grad_norm": 0.5253982543945312,
491
+ "learning_rate": 0.00013432134180256338,
492
+ "loss": 0.21154335021972656,
493
+ "mean_token_accuracy": 0.9283734840154648,
494
+ "num_tokens": 3072561.0,
495
+ "step": 2150
496
+ },
497
+ {
498
+ "entropy": 0.27213907435536383,
499
+ "epoch": 5.656370656370656,
500
+ "grad_norm": 0.46738553047180176,
501
+ "learning_rate": 0.00012824634177650664,
502
+ "loss": 0.21339216232299804,
503
+ "mean_token_accuracy": 0.9272083270549775,
504
+ "num_tokens": 3144831.0,
505
+ "step": 2200
506
+ },
507
+ {
508
+ "entropy": 0.2785488124191761,
509
+ "epoch": 5.785070785070785,
510
+ "grad_norm": 0.4469502866268158,
511
+ "learning_rate": 0.00012218589346414205,
512
+ "loss": 0.21601097106933595,
513
+ "mean_token_accuracy": 0.9255663657188415,
514
+ "num_tokens": 3215960.0,
515
+ "step": 2250
516
+ },
517
+ {
518
+ "entropy": 0.2699935150146484,
519
+ "epoch": 5.913770913770914,
520
+ "grad_norm": 0.7359778881072998,
521
+ "learning_rate": 0.00011615219483173828,
522
+ "loss": 0.20725584030151367,
523
+ "mean_token_accuracy": 0.9286630594730377,
524
+ "num_tokens": 3287499.0,
525
+ "step": 2300
526
+ },
527
+ {
528
+ "epoch": 6.0,
529
+ "eval_entropy": 0.26417383682174783,
530
+ "eval_loss": 0.8880229592323303,
531
+ "eval_mean_token_accuracy": 0.8159987201395723,
532
+ "eval_num_tokens": 3332958.0,
533
+ "eval_runtime": 162.0991,
534
+ "eval_samples_per_second": 9.531,
535
+ "eval_steps_per_second": 1.197,
536
+ "step": 2334
537
+ },
538
+ {
539
+ "entropy": 0.24814540704693458,
540
+ "epoch": 6.041184041184041,
541
+ "grad_norm": 0.4953760802745819,
542
+ "learning_rate": 0.00011015739000603316,
543
+ "loss": 0.18749794006347656,
544
+ "mean_token_accuracy": 0.9370789509831052,
545
+ "num_tokens": 3356879.0,
546
+ "step": 2350
547
+ },
548
+ {
549
+ "entropy": 0.19976271741092205,
550
+ "epoch": 6.1698841698841695,
551
+ "grad_norm": 0.4834803342819214,
552
+ "learning_rate": 0.00010421354483154553,
553
+ "loss": 0.14283526420593262,
554
+ "mean_token_accuracy": 0.9521516615152359,
555
+ "num_tokens": 3427587.0,
556
+ "step": 2400
557
+ },
558
+ {
559
+ "entropy": 0.2060488449037075,
560
+ "epoch": 6.298584298584299,
561
+ "grad_norm": 0.4888673722743988,
562
+ "learning_rate": 9.8332622585447e-05,
563
+ "loss": 0.14414511680603026,
564
+ "mean_token_accuracy": 0.9510996866226197,
565
+ "num_tokens": 3498688.0,
566
+ "step": 2450
567
+ },
568
+ {
569
+ "entropy": 0.2059111550450325,
570
+ "epoch": 6.427284427284428,
571
+ "grad_norm": 0.4064404368400574,
572
+ "learning_rate": 9.252645989887253e-05,
573
+ "loss": 0.14820143699645996,
574
+ "mean_token_accuracy": 0.9507584601640702,
575
+ "num_tokens": 3566137.0,
576
+ "step": 2500
577
+ },
578
+ {
579
+ "entropy": 0.19700154662132263,
580
+ "epoch": 6.555984555984556,
581
+ "grad_norm": 0.467965304851532,
582
+ "learning_rate": 8.680674293313417e-05,
583
+ "loss": 0.14303470611572267,
584
+ "mean_token_accuracy": 0.9515972435474396,
585
+ "num_tokens": 3639573.0,
586
+ "step": 2550
587
+ },
588
+ {
589
+ "entropy": 0.20180423602461814,
590
+ "epoch": 6.684684684684685,
591
+ "grad_norm": 0.36836138367652893,
592
+ "learning_rate": 8.118498385878736e-05,
593
+ "loss": 0.14280882835388184,
594
+ "mean_token_accuracy": 0.9515993863344192,
595
+ "num_tokens": 3710433.0,
596
+ "step": 2600
597
+ },
598
+ {
599
+ "entropy": 0.20024395987391472,
600
+ "epoch": 6.813384813384813,
601
+ "grad_norm": 0.38375866413116455,
602
+ "learning_rate": 7.567249768489171e-05,
603
+ "loss": 0.1427844524383545,
604
+ "mean_token_accuracy": 0.9524166631698608,
605
+ "num_tokens": 3781550.0,
606
+ "step": 2650
607
+ },
608
+ {
609
+ "entropy": 0.19561587080359458,
610
+ "epoch": 6.942084942084942,
611
+ "grad_norm": 0.41185441613197327,
612
+ "learning_rate": 7.028037948510187e-05,
613
+ "loss": 0.13993803024291993,
614
+ "mean_token_accuracy": 0.9522478264570237,
615
+ "num_tokens": 3854952.0,
616
+ "step": 2700
617
+ },
618
+ {
619
+ "epoch": 7.0,
620
+ "eval_entropy": 0.19501976062034823,
621
+ "eval_loss": 1.0653952360153198,
622
+ "eval_mean_token_accuracy": 0.8205490803595671,
623
+ "eval_num_tokens": 3888451.0,
624
+ "eval_runtime": 161.8533,
625
+ "eval_samples_per_second": 9.546,
626
+ "eval_steps_per_second": 1.199,
627
+ "step": 2723
628
+ },
629
+ {
630
+ "entropy": 0.17789882526855277,
631
+ "epoch": 7.06949806949807,
632
+ "grad_norm": 0.41413992643356323,
633
+ "learning_rate": 6.50194820664261e-05,
634
+ "loss": 0.12078390121459961,
635
+ "mean_token_accuracy": 0.9589925727458916,
636
+ "num_tokens": 3928354.0,
637
+ "step": 2750
638
+ },
639
+ {
640
+ "entropy": 0.16781829454004765,
641
+ "epoch": 7.198198198198198,
642
+ "grad_norm": 0.25806066393852234,
643
+ "learning_rate": 5.990039412559906e-05,
644
+ "loss": 0.10963023185729981,
645
+ "mean_token_accuracy": 0.9617267113924026,
646
+ "num_tokens": 4000113.0,
647
+ "step": 2800
648
+ },
649
+ {
650
+ "entropy": 0.1649068508297205,
651
+ "epoch": 7.326898326898327,
652
+ "grad_norm": 0.27411890029907227,
653
+ "learning_rate": 5.493341893703393e-05,
654
+ "loss": 0.11152458190917969,
655
+ "mean_token_accuracy": 0.9620639663934708,
656
+ "num_tokens": 4071032.0,
657
+ "step": 2850
658
+ },
659
+ {
660
+ "entropy": 0.161333369910717,
661
+ "epoch": 7.455598455598455,
662
+ "grad_norm": 0.24944494664669037,
663
+ "learning_rate": 5.0128553615248396e-05,
664
+ "loss": 0.1094522476196289,
665
+ "mean_token_accuracy": 0.962428919672966,
666
+ "num_tokens": 4143616.0,
667
+ "step": 2900
668
+ },
669
+ {
670
+ "entropy": 0.15613057143986225,
671
+ "epoch": 7.584298584298584,
672
+ "grad_norm": 0.1455036848783493,
673
+ "learning_rate": 4.549546899350423e-05,
674
+ "loss": 0.11092090606689453,
675
+ "mean_token_accuracy": 0.9620462411642074,
676
+ "num_tokens": 4215664.0,
677
+ "step": 2950
678
+ },
679
+ {
680
+ "entropy": 0.1631234459578991,
681
+ "epoch": 7.712998712998713,
682
+ "grad_norm": 0.2129560261964798,
683
+ "learning_rate": 4.104349015915862e-05,
684
+ "loss": 0.1141857624053955,
685
+ "mean_token_accuracy": 0.9613765001296997,
686
+ "num_tokens": 4286387.0,
687
+ "step": 3000
688
+ },
689
+ {
690
+ "entropy": 0.1680422095954418,
691
+ "epoch": 7.841698841698841,
692
+ "grad_norm": 0.24886097013950348,
693
+ "learning_rate": 3.678157768490372e-05,
694
+ "loss": 0.11513191223144531,
695
+ "mean_token_accuracy": 0.9615794748067856,
696
+ "num_tokens": 4355875.0,
697
+ "step": 3050
698
+ },
699
+ {
700
+ "entropy": 0.16354035697877406,
701
+ "epoch": 7.97039897039897,
702
+ "grad_norm": 0.27600204944610596,
703
+ "learning_rate": 3.27183095936714e-05,
704
+ "loss": 0.1118631362915039,
705
+ "mean_token_accuracy": 0.9623224419355393,
706
+ "num_tokens": 4427088.0,
707
+ "step": 3100
708
+ },
709
+ {
710
+ "epoch": 8.0,
711
+ "eval_entropy": 0.16586771200305409,
712
+ "eval_loss": 1.204746961593628,
713
+ "eval_mean_token_accuracy": 0.8229975042883882,
714
+ "eval_num_tokens": 4443944.0,
715
+ "eval_runtime": 162.0251,
716
+ "eval_samples_per_second": 9.536,
717
+ "eval_steps_per_second": 1.197,
718
+ "step": 3112
719
+ },
720
+ {
721
+ "entropy": 0.1515902608934075,
722
+ "epoch": 8.097812097812097,
723
+ "grad_norm": 0.14122211933135986,
724
+ "learning_rate": 2.88618640935022e-05,
725
+ "loss": 0.09900871276855469,
726
+ "mean_token_accuracy": 0.9665110737386376,
727
+ "num_tokens": 4497867.0,
728
+ "step": 3150
729
+ },
730
+ {
731
+ "entropy": 0.14593622356653213,
732
+ "epoch": 8.226512226512227,
733
+ "grad_norm": 0.20527532696723938,
734
+ "learning_rate": 2.5220003117128462e-05,
735
+ "loss": 0.09842084884643555,
736
+ "mean_token_accuracy": 0.9655911487340927,
737
+ "num_tokens": 4568534.0,
738
+ "step": 3200
739
+ },
740
+ {
741
+ "entropy": 0.1482392605394125,
742
+ "epoch": 8.355212355212355,
743
+ "grad_norm": 0.13207173347473145,
744
+ "learning_rate": 2.1800056699401584e-05,
745
+ "loss": 0.09551989555358886,
746
+ "mean_token_accuracy": 0.9650829958915711,
747
+ "num_tokens": 4642530.0,
748
+ "step": 3250
749
+ },
750
+ {
751
+ "entropy": 0.15284131653606892,
752
+ "epoch": 8.483912483912484,
753
+ "grad_norm": 0.1777282953262329,
754
+ "learning_rate": 1.860890822400777e-05,
755
+ "loss": 0.10169261932373047,
756
+ "mean_token_accuracy": 0.9635212075710297,
757
+ "num_tokens": 4711573.0,
758
+ "step": 3300
759
+ },
760
+ {
761
+ "entropy": 0.15095721945166587,
762
+ "epoch": 8.612612612612612,
763
+ "grad_norm": 0.14988408982753754,
764
+ "learning_rate": 1.5652980569165692e-05,
765
+ "loss": 0.10045011520385742,
766
+ "mean_token_accuracy": 0.96439110994339,
767
+ "num_tokens": 4782666.0,
768
+ "step": 3350
769
+ },
770
+ {
771
+ "entropy": 0.15411154814064504,
772
+ "epoch": 8.741312741312742,
773
+ "grad_norm": 0.14055995643138885,
774
+ "learning_rate": 1.2938223180191691e-05,
775
+ "loss": 0.1034860897064209,
776
+ "mean_token_accuracy": 0.963447842001915,
777
+ "num_tokens": 4852180.0,
778
+ "step": 3400
779
+ },
780
+ {
781
+ "entropy": 0.14739766091108322,
782
+ "epoch": 8.87001287001287,
783
+ "grad_norm": 0.15041407942771912,
784
+ "learning_rate": 1.0470100094950792e-05,
785
+ "loss": 0.09690508842468262,
786
+ "mean_token_accuracy": 0.96561603307724,
787
+ "num_tokens": 4926402.0,
788
+ "step": 3450
789
+ },
790
+ {
791
+ "entropy": 0.1498453303426504,
792
+ "epoch": 8.998712998712998,
793
+ "grad_norm": 0.1293368935585022,
794
+ "learning_rate": 8.253578946296125e-06,
795
+ "loss": 0.09874271392822266,
796
+ "mean_token_accuracy": 0.9647125631570816,
797
+ "num_tokens": 4998841.0,
798
+ "step": 3500
799
+ },
800
+ {
801
+ "epoch": 9.0,
802
+ "eval_entropy": 0.151532097729211,
803
+ "eval_loss": 1.302620768547058,
804
+ "eval_mean_token_accuracy": 0.8230925079473516,
805
+ "eval_num_tokens": 4999437.0,
806
+ "eval_runtime": 161.5816,
807
+ "eval_samples_per_second": 9.562,
808
+ "eval_steps_per_second": 1.201,
809
+ "step": 3501
810
+ }
811
+ ],
812
+ "logging_steps": 50,
813
+ "max_steps": 3890,
814
+ "num_input_tokens_seen": 0,
815
+ "num_train_epochs": 10,
816
+ "save_steps": 500,
817
+ "stateful_callbacks": {
818
+ "TrainerControl": {
819
+ "args": {
820
+ "should_epoch_stop": false,
821
+ "should_evaluate": false,
822
+ "should_log": false,
823
+ "should_save": true,
824
+ "should_training_stop": false
825
+ },
826
+ "attributes": {}
827
+ }
828
+ },
829
+ "total_flos": 8.377780321673011e+17,
830
+ "train_batch_size": 8,
831
+ "trial_name": null,
832
+ "trial_params": null
833
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/README.md ADDED
@@ -0,0 +1,209 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen3-14B-Base
3
+ library_name: peft
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - base_model:adapter:Qwen/Qwen3-14B-Base
7
+ - lora
8
+ - sft
9
+ - transformers
10
+ - trl
11
+ ---
12
+
13
+ # Model Card for Model ID
14
+
15
+ <!-- Provide a quick summary of what the model is/does. -->
16
+
17
+
18
+
19
+ ## Model Details
20
+
21
+ ### Model Description
22
+
23
+ <!-- Provide a longer summary of what this model is. -->
24
+
25
+
26
+
27
+ - **Developed by:** [More Information Needed]
28
+ - **Funded by [optional]:** [More Information Needed]
29
+ - **Shared by [optional]:** [More Information Needed]
30
+ - **Model type:** [More Information Needed]
31
+ - **Language(s) (NLP):** [More Information Needed]
32
+ - **License:** [More Information Needed]
33
+ - **Finetuned from model [optional]:** [More Information Needed]
34
+
35
+ ### Model Sources [optional]
36
+
37
+ <!-- Provide the basic links for the model. -->
38
+
39
+ - **Repository:** [More Information Needed]
40
+ - **Paper [optional]:** [More Information Needed]
41
+ - **Demo [optional]:** [More Information Needed]
42
+
43
+ ## Uses
44
+
45
+ <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
46
+
47
+ ### Direct Use
48
+
49
+ <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
50
+
51
+ [More Information Needed]
52
+
53
+ ### Downstream Use [optional]
54
+
55
+ <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
56
+
57
+ [More Information Needed]
58
+
59
+ ### Out-of-Scope Use
60
+
61
+ <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
62
+
63
+ [More Information Needed]
64
+
65
+ ## Bias, Risks, and Limitations
66
+
67
+ <!-- This section is meant to convey both technical and sociotechnical limitations. -->
68
+
69
+ [More Information Needed]
70
+
71
+ ### Recommendations
72
+
73
+ <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
74
+
75
+ Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
76
+
77
+ ## How to Get Started with the Model
78
+
79
+ Use the code below to get started with the model.
80
+
81
+ [More Information Needed]
82
+
83
+ ## Training Details
84
+
85
+ ### Training Data
86
+
87
+ <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
88
+
89
+ [More Information Needed]
90
+
91
+ ### Training Procedure
92
+
93
+ <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
94
+
95
+ #### Preprocessing [optional]
96
+
97
+ [More Information Needed]
98
+
99
+
100
+ #### Training Hyperparameters
101
+
102
+ - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
103
+
104
+ #### Speeds, Sizes, Times [optional]
105
+
106
+ <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
107
+
108
+ [More Information Needed]
109
+
110
+ ## Evaluation
111
+
112
+ <!-- This section describes the evaluation protocols and provides the results. -->
113
+
114
+ ### Testing Data, Factors & Metrics
115
+
116
+ #### Testing Data
117
+
118
+ <!-- This should link to a Dataset Card if possible. -->
119
+
120
+ [More Information Needed]
121
+
122
+ #### Factors
123
+
124
+ <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
125
+
126
+ [More Information Needed]
127
+
128
+ #### Metrics
129
+
130
+ <!-- These are the evaluation metrics being used, ideally with a description of why. -->
131
+
132
+ [More Information Needed]
133
+
134
+ ### Results
135
+
136
+ [More Information Needed]
137
+
138
+ #### Summary
139
+
140
+
141
+
142
+ ## Model Examination [optional]
143
+
144
+ <!-- Relevant interpretability work for the model goes here -->
145
+
146
+ [More Information Needed]
147
+
148
+ ## Environmental Impact
149
+
150
+ <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
151
+
152
+ Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
153
+
154
+ - **Hardware Type:** [More Information Needed]
155
+ - **Hours used:** [More Information Needed]
156
+ - **Cloud Provider:** [More Information Needed]
157
+ - **Compute Region:** [More Information Needed]
158
+ - **Carbon Emitted:** [More Information Needed]
159
+
160
+ ## Technical Specifications [optional]
161
+
162
+ ### Model Architecture and Objective
163
+
164
+ [More Information Needed]
165
+
166
+ ### Compute Infrastructure
167
+
168
+ [More Information Needed]
169
+
170
+ #### Hardware
171
+
172
+ [More Information Needed]
173
+
174
+ #### Software
175
+
176
+ [More Information Needed]
177
+
178
+ ## Citation [optional]
179
+
180
+ <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
181
+
182
+ **BibTeX:**
183
+
184
+ [More Information Needed]
185
+
186
+ **APA:**
187
+
188
+ [More Information Needed]
189
+
190
+ ## Glossary [optional]
191
+
192
+ <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
193
+
194
+ [More Information Needed]
195
+
196
+ ## More Information [optional]
197
+
198
+ [More Information Needed]
199
+
200
+ ## Model Card Authors [optional]
201
+
202
+ [More Information Needed]
203
+
204
+ ## Model Card Contact
205
+
206
+ [More Information Needed]
207
+ ### Framework versions
208
+
209
+ - PEFT 0.19.1
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/adapter_config.json ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "Qwen/Qwen3-14B-Base",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "layer_replication": null,
16
+ "layers_pattern": null,
17
+ "layers_to_transform": null,
18
+ "loftq_config": {},
19
+ "lora_alpha": 256,
20
+ "lora_bias": false,
21
+ "lora_dropout": 0.00237968804112545,
22
+ "lora_ga_config": null,
23
+ "megatron_config": null,
24
+ "megatron_core": "megatron.core",
25
+ "modules_to_save": null,
26
+ "peft_type": "LORA",
27
+ "peft_version": "0.19.1",
28
+ "qalora_group_size": 16,
29
+ "r": 128,
30
+ "rank_pattern": {},
31
+ "revision": null,
32
+ "target_modules": [
33
+ "gate_proj",
34
+ "q_proj",
35
+ "up_proj",
36
+ "o_proj",
37
+ "v_proj",
38
+ "k_proj",
39
+ "down_proj"
40
+ ],
41
+ "target_parameters": null,
42
+ "task_type": "CAUSAL_LM",
43
+ "trainable_token_indices": null,
44
+ "use_bdlora": null,
45
+ "use_dora": false,
46
+ "use_qalora": false,
47
+ "use_rslora": false
48
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/chat_template.jinja ADDED
@@ -0,0 +1,85 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- if tools %}
2
+ {{- '<|im_start|>system\n' }}
3
+ {%- if messages[0].role == 'system' %}
4
+ {{- messages[0].content + '\n\n' }}
5
+ {%- endif %}
6
+ {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
7
+ {%- for tool in tools %}
8
+ {{- "\n" }}
9
+ {{- tool | tojson }}
10
+ {%- endfor %}
11
+ {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
12
+ {%- else %}
13
+ {%- if messages[0].role == 'system' %}
14
+ {{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
15
+ {%- endif %}
16
+ {%- endif %}
17
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
18
+ {%- for message in messages[::-1] %}
19
+ {%- set index = (messages|length - 1) - loop.index0 %}
20
+ {%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
21
+ {%- set ns.multi_step_tool = false %}
22
+ {%- set ns.last_query_index = index %}
23
+ {%- endif %}
24
+ {%- endfor %}
25
+ {%- for message in messages %}
26
+ {%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
27
+ {{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
28
+ {%- elif message.role == "assistant" %}
29
+ {%- set content = message.content %}
30
+ {%- set reasoning_content = '' %}
31
+ {%- if message.reasoning_content is defined and message.reasoning_content is not none %}
32
+ {%- set reasoning_content = message.reasoning_content %}
33
+ {%- else %}
34
+ {%- if '</think>' in message.content %}
35
+ {%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
36
+ {%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
37
+ {%- endif %}
38
+ {%- endif %}
39
+ {%- if loop.index0 > ns.last_query_index %}
40
+ {%- if loop.last or (not loop.last and reasoning_content) %}
41
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
42
+ {%- else %}
43
+ {{- '<|im_start|>' + message.role + '\n' + content }}
44
+ {%- endif %}
45
+ {%- else %}
46
+ {{- '<|im_start|>' + message.role + '\n' + content }}
47
+ {%- endif %}
48
+ {%- if message.tool_calls %}
49
+ {%- for tool_call in message.tool_calls %}
50
+ {%- if (loop.first and content) or (not loop.first) %}
51
+ {{- '\n' }}
52
+ {%- endif %}
53
+ {%- if tool_call.function %}
54
+ {%- set tool_call = tool_call.function %}
55
+ {%- endif %}
56
+ {{- '<tool_call>\n{"name": "' }}
57
+ {{- tool_call.name }}
58
+ {{- '", "arguments": ' }}
59
+ {%- if tool_call.arguments is string %}
60
+ {{- tool_call.arguments }}
61
+ {%- else %}
62
+ {{- tool_call.arguments | tojson }}
63
+ {%- endif %}
64
+ {{- '}\n</tool_call>' }}
65
+ {%- endfor %}
66
+ {%- endif %}
67
+ {{- '<|im_end|>\n' }}
68
+ {%- elif message.role == "tool" %}
69
+ {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
70
+ {{- '<|im_start|>user' }}
71
+ {%- endif %}
72
+ {{- '\n<tool_response>\n' }}
73
+ {{- message.content }}
74
+ {{- '\n</tool_response>' }}
75
+ {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
76
+ {{- '<|im_end|>\n' }}
77
+ {%- endif %}
78
+ {%- endif %}
79
+ {%- endfor %}
80
+ {%- if add_generation_prompt %}
81
+ {{- '<|im_start|>assistant\n' }}
82
+ {%- if enable_thinking is defined and enable_thinking is false %}
83
+ {{- '<think>\n\n</think>\n\n' }}
84
+ {%- endif %}
85
+ {%- endif %}
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/tokenizer_config.json ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "backend": "tokenizers",
4
+ "bos_token": null,
5
+ "clean_up_tokenization_spaces": false,
6
+ "eos_token": "<|endoftext|>",
7
+ "errors": "replace",
8
+ "extra_special_tokens": [
9
+ "<|im_start|>",
10
+ "<|im_end|>",
11
+ "<|object_ref_start|>",
12
+ "<|object_ref_end|>",
13
+ "<|box_start|>",
14
+ "<|box_end|>",
15
+ "<|quad_start|>",
16
+ "<|quad_end|>",
17
+ "<|vision_start|>",
18
+ "<|vision_end|>",
19
+ "<|vision_pad|>",
20
+ "<|image_pad|>",
21
+ "<|video_pad|>"
22
+ ],
23
+ "is_local": false,
24
+ "model_max_length": 131072,
25
+ "pad_token": "<|endoftext|>",
26
+ "split_special_tokens": false,
27
+ "tokenizer_class": "Qwen2Tokenizer",
28
+ "unk_token": null
29
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/trainer_state.json ADDED
@@ -0,0 +1,115 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "best_global_step": null,
3
+ "best_metric": null,
4
+ "best_model_checkpoint": null,
5
+ "epoch": 1.0,
6
+ "eval_steps": 500,
7
+ "global_step": 389,
8
+ "is_hyper_param_search": false,
9
+ "is_local_process_zero": true,
10
+ "is_world_process_zero": true,
11
+ "log_history": [
12
+ {
13
+ "entropy": 1.6110110306739807,
14
+ "epoch": 0.1287001287001287,
15
+ "grad_norm": 0.8611119389533997,
16
+ "learning_rate": 3.4130257962133866e-05,
17
+ "loss": 1.53956298828125,
18
+ "mean_token_accuracy": 0.6652342769503593,
19
+ "num_tokens": 73407.0,
20
+ "step": 50
21
+ },
22
+ {
23
+ "entropy": 0.8662820833921433,
24
+ "epoch": 0.2574002574002574,
25
+ "grad_norm": 0.7405035495758057,
26
+ "learning_rate": 6.895705180104598e-05,
27
+ "loss": 0.7929539489746094,
28
+ "mean_token_accuracy": 0.7786347842216492,
29
+ "num_tokens": 143994.0,
30
+ "step": 100
31
+ },
32
+ {
33
+ "entropy": 0.7912819278240204,
34
+ "epoch": 0.3861003861003861,
35
+ "grad_norm": 0.6105485558509827,
36
+ "learning_rate": 0.00010378384563995809,
37
+ "loss": 0.7236511993408203,
38
+ "mean_token_accuracy": 0.7925467795133591,
39
+ "num_tokens": 216171.0,
40
+ "step": 150
41
+ },
42
+ {
43
+ "entropy": 0.7777709531784057,
44
+ "epoch": 0.5148005148005148,
45
+ "grad_norm": 0.5781793594360352,
46
+ "learning_rate": 0.0001386106394788702,
47
+ "loss": 0.7024919891357422,
48
+ "mean_token_accuracy": 0.7969876372814179,
49
+ "num_tokens": 284702.0,
50
+ "step": 200
51
+ },
52
+ {
53
+ "entropy": 0.7632609683275223,
54
+ "epoch": 0.6435006435006435,
55
+ "grad_norm": 0.6553444862365723,
56
+ "learning_rate": 0.00017343743331778232,
57
+ "loss": 0.7000718688964844,
58
+ "mean_token_accuracy": 0.7994742071628571,
59
+ "num_tokens": 356393.0,
60
+ "step": 250
61
+ },
62
+ {
63
+ "entropy": 0.754165632724762,
64
+ "epoch": 0.7722007722007722,
65
+ "grad_norm": 0.634222149848938,
66
+ "learning_rate": 0.00020826422715669444,
67
+ "loss": 0.6909049987792969,
68
+ "mean_token_accuracy": 0.8016809666156769,
69
+ "num_tokens": 426916.0,
70
+ "step": 300
71
+ },
72
+ {
73
+ "entropy": 0.7530711203813553,
74
+ "epoch": 0.9009009009009009,
75
+ "grad_norm": 0.7020156383514404,
76
+ "learning_rate": 0.00024309102099560653,
77
+ "loss": 0.69554931640625,
78
+ "mean_token_accuracy": 0.8002807641029358,
79
+ "num_tokens": 499603.0,
80
+ "step": 350
81
+ },
82
+ {
83
+ "epoch": 1.0,
84
+ "eval_entropy": 0.5728044164242204,
85
+ "eval_loss": 0.6785851716995239,
86
+ "eval_mean_token_accuracy": 0.798433246993527,
87
+ "eval_num_tokens": 555493.0,
88
+ "eval_runtime": 162.4129,
89
+ "eval_samples_per_second": 9.513,
90
+ "eval_steps_per_second": 1.194,
91
+ "step": 389
92
+ }
93
+ ],
94
+ "logging_steps": 50,
95
+ "max_steps": 3890,
96
+ "num_input_tokens_seen": 0,
97
+ "num_train_epochs": 10,
98
+ "save_steps": 500,
99
+ "stateful_callbacks": {
100
+ "TrainerControl": {
101
+ "args": {
102
+ "should_epoch_stop": false,
103
+ "should_evaluate": false,
104
+ "should_log": false,
105
+ "should_save": true,
106
+ "should_training_stop": false
107
+ },
108
+ "attributes": {}
109
+ }
110
+ },
111
+ "total_flos": 9.306612280369152e+16,
112
+ "train_batch_size": 8,
113
+ "trial_name": null,
114
+ "trial_params": null
115
+ }
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3890/README.md ADDED
@@ -0,0 +1,209 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen3-14B-Base
3
+ library_name: peft
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - base_model:adapter:Qwen/Qwen3-14B-Base
7
+ - lora
8
+ - sft
9
+ - transformers
10
+ - trl
11
+ ---
12
+
13
+ # Model Card for Model ID
14
+
15
+ <!-- Provide a quick summary of what the model is/does. -->
16
+
17
+
18
+
19
+ ## Model Details
20
+
21
+ ### Model Description
22
+
23
+ <!-- Provide a longer summary of what this model is. -->
24
+
25
+
26
+
27
+ - **Developed by:** [More Information Needed]
28
+ - **Funded by [optional]:** [More Information Needed]
29
+ - **Shared by [optional]:** [More Information Needed]
30
+ - **Model type:** [More Information Needed]
31
+ - **Language(s) (NLP):** [More Information Needed]
32
+ - **License:** [More Information Needed]
33
+ - **Finetuned from model [optional]:** [More Information Needed]
34
+
35
+ ### Model Sources [optional]
36
+
37
+ <!-- Provide the basic links for the model. -->
38
+
39
+ - **Repository:** [More Information Needed]
40
+ - **Paper [optional]:** [More Information Needed]
41
+ - **Demo [optional]:** [More Information Needed]
42
+
43
+ ## Uses
44
+
45
+ <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
46
+
47
+ ### Direct Use
48
+
49
+ <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
50
+
51
+ [More Information Needed]
52
+
53
+ ### Downstream Use [optional]
54
+
55
+ <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
56
+
57
+ [More Information Needed]
58
+
59
+ ### Out-of-Scope Use
60
+
61
+ <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
62
+
63
+ [More Information Needed]
64
+
65
+ ## Bias, Risks, and Limitations
66
+
67
+ <!-- This section is meant to convey both technical and sociotechnical limitations. -->
68
+
69
+ [More Information Needed]
70
+
71
+ ### Recommendations
72
+
73
+ <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
74
+
75
+ Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
76
+
77
+ ## How to Get Started with the Model
78
+
79
+ Use the code below to get started with the model.
80
+
81
+ [More Information Needed]
82
+
83
+ ## Training Details
84
+
85
+ ### Training Data
86
+
87
+ <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
88
+
89
+ [More Information Needed]
90
+
91
+ ### Training Procedure
92
+
93
+ <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
94
+
95
+ #### Preprocessing [optional]
96
+
97
+ [More Information Needed]
98
+
99
+
100
+ #### Training Hyperparameters
101
+
102
+ - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
103
+
104
+ #### Speeds, Sizes, Times [optional]
105
+
106
+ <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
107
+
108
+ [More Information Needed]
109
+
110
+ ## Evaluation
111
+
112
+ <!-- This section describes the evaluation protocols and provides the results. -->
113
+
114
+ ### Testing Data, Factors & Metrics
115
+
116
+ #### Testing Data
117
+
118
+ <!-- This should link to a Dataset Card if possible. -->
119
+
120
+ [More Information Needed]
121
+
122
+ #### Factors
123
+
124
+ <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
125
+
126
+ [More Information Needed]
127
+
128
+ #### Metrics
129
+
130
+ <!-- These are the evaluation metrics being used, ideally with a description of why. -->
131
+
132
+ [More Information Needed]
133
+
134
+ ### Results
135
+
136
+ [More Information Needed]
137
+
138
+ #### Summary
139
+
140
+
141
+
142
+ ## Model Examination [optional]
143
+
144
+ <!-- Relevant interpretability work for the model goes here -->
145
+
146
+ [More Information Needed]
147
+
148
+ ## Environmental Impact
149
+
150
+ <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
151
+
152
+ Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
153
+
154
+ - **Hardware Type:** [More Information Needed]
155
+ - **Hours used:** [More Information Needed]
156
+ - **Cloud Provider:** [More Information Needed]
157
+ - **Compute Region:** [More Information Needed]
158
+ - **Carbon Emitted:** [More Information Needed]
159
+
160
+ ## Technical Specifications [optional]
161
+
162
+ ### Model Architecture and Objective
163
+
164
+ [More Information Needed]
165
+
166
+ ### Compute Infrastructure
167
+
168
+ [More Information Needed]
169
+
170
+ #### Hardware
171
+
172
+ [More Information Needed]
173
+
174
+ #### Software
175
+
176
+ [More Information Needed]
177
+
178
+ ## Citation [optional]
179
+
180
+ <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
181
+
182
+ **BibTeX:**
183
+
184
+ [More Information Needed]
185
+
186
+ **APA:**
187
+
188
+ [More Information Needed]
189
+
190
+ ## Glossary [optional]
191
+
192
+ <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
193
+
194
+ [More Information Needed]
195
+
196
+ ## More Information [optional]
197
+
198
+ [More Information Needed]
199
+
200
+ ## Model Card Authors [optional]
201
+
202
+ [More Information Needed]
203
+
204
+ ## Model Card Contact
205
+
206
+ [More Information Needed]
207
+ ### Framework versions
208
+
209
+ - PEFT 0.19.1
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3890/adapter_config.json ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "Qwen/Qwen3-14B-Base",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "layer_replication": null,
16
+ "layers_pattern": null,
17
+ "layers_to_transform": null,
18
+ "loftq_config": {},
19
+ "lora_alpha": 256,
20
+ "lora_bias": false,
21
+ "lora_dropout": 0.00237968804112545,
22
+ "lora_ga_config": null,
23
+ "megatron_config": null,
24
+ "megatron_core": "megatron.core",
25
+ "modules_to_save": null,
26
+ "peft_type": "LORA",
27
+ "peft_version": "0.19.1",
28
+ "qalora_group_size": 16,
29
+ "r": 128,
30
+ "rank_pattern": {},
31
+ "revision": null,
32
+ "target_modules": [
33
+ "gate_proj",
34
+ "q_proj",
35
+ "up_proj",
36
+ "o_proj",
37
+ "v_proj",
38
+ "k_proj",
39
+ "down_proj"
40
+ ],
41
+ "target_parameters": null,
42
+ "task_type": "CAUSAL_LM",
43
+ "trainable_token_indices": null,
44
+ "use_bdlora": null,
45
+ "use_dora": false,
46
+ "use_qalora": false,
47
+ "use_rslora": false
48
+ }
random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/README.md ADDED
@@ -0,0 +1,209 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen3-14B-Base
3
+ library_name: peft
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - base_model:adapter:Qwen/Qwen3-14B-Base
7
+ - lora
8
+ - sft
9
+ - transformers
10
+ - trl
11
+ ---
12
+
13
+ # Model Card for Model ID
14
+
15
+ <!-- Provide a quick summary of what the model is/does. -->
16
+
17
+
18
+
19
+ ## Model Details
20
+
21
+ ### Model Description
22
+
23
+ <!-- Provide a longer summary of what this model is. -->
24
+
25
+
26
+
27
+ - **Developed by:** [More Information Needed]
28
+ - **Funded by [optional]:** [More Information Needed]
29
+ - **Shared by [optional]:** [More Information Needed]
30
+ - **Model type:** [More Information Needed]
31
+ - **Language(s) (NLP):** [More Information Needed]
32
+ - **License:** [More Information Needed]
33
+ - **Finetuned from model [optional]:** [More Information Needed]
34
+
35
+ ### Model Sources [optional]
36
+
37
+ <!-- Provide the basic links for the model. -->
38
+
39
+ - **Repository:** [More Information Needed]
40
+ - **Paper [optional]:** [More Information Needed]
41
+ - **Demo [optional]:** [More Information Needed]
42
+
43
+ ## Uses
44
+
45
+ <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
46
+
47
+ ### Direct Use
48
+
49
+ <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
50
+
51
+ [More Information Needed]
52
+
53
+ ### Downstream Use [optional]
54
+
55
+ <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
56
+
57
+ [More Information Needed]
58
+
59
+ ### Out-of-Scope Use
60
+
61
+ <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
62
+
63
+ [More Information Needed]
64
+
65
+ ## Bias, Risks, and Limitations
66
+
67
+ <!-- This section is meant to convey both technical and sociotechnical limitations. -->
68
+
69
+ [More Information Needed]
70
+
71
+ ### Recommendations
72
+
73
+ <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
74
+
75
+ Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
76
+
77
+ ## How to Get Started with the Model
78
+
79
+ Use the code below to get started with the model.
80
+
81
+ [More Information Needed]
82
+
83
+ ## Training Details
84
+
85
+ ### Training Data
86
+
87
+ <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
88
+
89
+ [More Information Needed]
90
+
91
+ ### Training Procedure
92
+
93
+ <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
94
+
95
+ #### Preprocessing [optional]
96
+
97
+ [More Information Needed]
98
+
99
+
100
+ #### Training Hyperparameters
101
+
102
+ - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
103
+
104
+ #### Speeds, Sizes, Times [optional]
105
+
106
+ <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
107
+
108
+ [More Information Needed]
109
+
110
+ ## Evaluation
111
+
112
+ <!-- This section describes the evaluation protocols and provides the results. -->
113
+
114
+ ### Testing Data, Factors & Metrics
115
+
116
+ #### Testing Data
117
+
118
+ <!-- This should link to a Dataset Card if possible. -->
119
+
120
+ [More Information Needed]
121
+
122
+ #### Factors
123
+
124
+ <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
125
+
126
+ [More Information Needed]
127
+
128
+ #### Metrics
129
+
130
+ <!-- These are the evaluation metrics being used, ideally with a description of why. -->
131
+
132
+ [More Information Needed]
133
+
134
+ ### Results
135
+
136
+ [More Information Needed]
137
+
138
+ #### Summary
139
+
140
+
141
+
142
+ ## Model Examination [optional]
143
+
144
+ <!-- Relevant interpretability work for the model goes here -->
145
+
146
+ [More Information Needed]
147
+
148
+ ## Environmental Impact
149
+
150
+ <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
151
+
152
+ Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
153
+
154
+ - **Hardware Type:** [More Information Needed]
155
+ - **Hours used:** [More Information Needed]
156
+ - **Cloud Provider:** [More Information Needed]
157
+ - **Compute Region:** [More Information Needed]
158
+ - **Carbon Emitted:** [More Information Needed]
159
+
160
+ ## Technical Specifications [optional]
161
+
162
+ ### Model Architecture and Objective
163
+
164
+ [More Information Needed]
165
+
166
+ ### Compute Infrastructure
167
+
168
+ [More Information Needed]
169
+
170
+ #### Hardware
171
+
172
+ [More Information Needed]
173
+
174
+ #### Software
175
+
176
+ [More Information Needed]
177
+
178
+ ## Citation [optional]
179
+
180
+ <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
181
+
182
+ **BibTeX:**
183
+
184
+ [More Information Needed]
185
+
186
+ **APA:**
187
+
188
+ [More Information Needed]
189
+
190
+ ## Glossary [optional]
191
+
192
+ <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
193
+
194
+ [More Information Needed]
195
+
196
+ ## More Information [optional]
197
+
198
+ [More Information Needed]
199
+
200
+ ## Model Card Authors [optional]
201
+
202
+ [More Information Needed]
203
+
204
+ ## Model Card Contact
205
+
206
+ [More Information Needed]
207
+ ### Framework versions
208
+
209
+ - PEFT 0.19.1
random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/adapter_config.json ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "Qwen/Qwen3-14B-Base",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "layer_replication": null,
16
+ "layers_pattern": null,
17
+ "layers_to_transform": null,
18
+ "loftq_config": {},
19
+ "lora_alpha": 64,
20
+ "lora_bias": false,
21
+ "lora_dropout": 0.09831666542701797,
22
+ "lora_ga_config": null,
23
+ "megatron_config": null,
24
+ "megatron_core": "megatron.core",
25
+ "modules_to_save": null,
26
+ "peft_type": "LORA",
27
+ "peft_version": "0.19.1",
28
+ "qalora_group_size": 16,
29
+ "r": 32,
30
+ "rank_pattern": {},
31
+ "revision": null,
32
+ "target_modules": [
33
+ "gate_proj",
34
+ "o_proj",
35
+ "k_proj",
36
+ "q_proj",
37
+ "up_proj",
38
+ "v_proj",
39
+ "down_proj"
40
+ ],
41
+ "target_parameters": null,
42
+ "task_type": "CAUSAL_LM",
43
+ "trainable_token_indices": null,
44
+ "use_bdlora": null,
45
+ "use_dora": false,
46
+ "use_qalora": false,
47
+ "use_rslora": false
48
+ }
random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/chat_template.jinja ADDED
@@ -0,0 +1,85 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- if tools %}
2
+ {{- '<|im_start|>system\n' }}
3
+ {%- if messages[0].role == 'system' %}
4
+ {{- messages[0].content + '\n\n' }}
5
+ {%- endif %}
6
+ {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
7
+ {%- for tool in tools %}
8
+ {{- "\n" }}
9
+ {{- tool | tojson }}
10
+ {%- endfor %}
11
+ {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
12
+ {%- else %}
13
+ {%- if messages[0].role == 'system' %}
14
+ {{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
15
+ {%- endif %}
16
+ {%- endif %}
17
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
18
+ {%- for message in messages[::-1] %}
19
+ {%- set index = (messages|length - 1) - loop.index0 %}
20
+ {%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
21
+ {%- set ns.multi_step_tool = false %}
22
+ {%- set ns.last_query_index = index %}
23
+ {%- endif %}
24
+ {%- endfor %}
25
+ {%- for message in messages %}
26
+ {%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
27
+ {{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
28
+ {%- elif message.role == "assistant" %}
29
+ {%- set content = message.content %}
30
+ {%- set reasoning_content = '' %}
31
+ {%- if message.reasoning_content is defined and message.reasoning_content is not none %}
32
+ {%- set reasoning_content = message.reasoning_content %}
33
+ {%- else %}
34
+ {%- if '</think>' in message.content %}
35
+ {%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
36
+ {%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
37
+ {%- endif %}
38
+ {%- endif %}
39
+ {%- if loop.index0 > ns.last_query_index %}
40
+ {%- if loop.last or (not loop.last and reasoning_content) %}
41
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
42
+ {%- else %}
43
+ {{- '<|im_start|>' + message.role + '\n' + content }}
44
+ {%- endif %}
45
+ {%- else %}
46
+ {{- '<|im_start|>' + message.role + '\n' + content }}
47
+ {%- endif %}
48
+ {%- if message.tool_calls %}
49
+ {%- for tool_call in message.tool_calls %}
50
+ {%- if (loop.first and content) or (not loop.first) %}
51
+ {{- '\n' }}
52
+ {%- endif %}
53
+ {%- if tool_call.function %}
54
+ {%- set tool_call = tool_call.function %}
55
+ {%- endif %}
56
+ {{- '<tool_call>\n{"name": "' }}
57
+ {{- tool_call.name }}
58
+ {{- '", "arguments": ' }}
59
+ {%- if tool_call.arguments is string %}
60
+ {{- tool_call.arguments }}
61
+ {%- else %}
62
+ {{- tool_call.arguments | tojson }}
63
+ {%- endif %}
64
+ {{- '}\n</tool_call>' }}
65
+ {%- endfor %}
66
+ {%- endif %}
67
+ {{- '<|im_end|>\n' }}
68
+ {%- elif message.role == "tool" %}
69
+ {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
70
+ {{- '<|im_start|>user' }}
71
+ {%- endif %}
72
+ {{- '\n<tool_response>\n' }}
73
+ {{- message.content }}
74
+ {{- '\n</tool_response>' }}
75
+ {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
76
+ {{- '<|im_end|>\n' }}
77
+ {%- endif %}
78
+ {%- endif %}
79
+ {%- endfor %}
80
+ {%- if add_generation_prompt %}
81
+ {{- '<|im_start|>assistant\n' }}
82
+ {%- if enable_thinking is defined and enable_thinking is false %}
83
+ {{- '<think>\n\n</think>\n\n' }}
84
+ {%- endif %}
85
+ {%- endif %}
random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/tokenizer_config.json ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "backend": "tokenizers",
4
+ "bos_token": null,
5
+ "clean_up_tokenization_spaces": false,
6
+ "eos_token": "<|endoftext|>",
7
+ "errors": "replace",
8
+ "extra_special_tokens": [
9
+ "<|im_start|>",
10
+ "<|im_end|>",
11
+ "<|object_ref_start|>",
12
+ "<|object_ref_end|>",
13
+ "<|box_start|>",
14
+ "<|box_end|>",
15
+ "<|quad_start|>",
16
+ "<|quad_end|>",
17
+ "<|vision_start|>",
18
+ "<|vision_end|>",
19
+ "<|vision_pad|>",
20
+ "<|image_pad|>",
21
+ "<|video_pad|>"
22
+ ],
23
+ "is_local": false,
24
+ "model_max_length": 131072,
25
+ "pad_token": "<|endoftext|>",
26
+ "split_special_tokens": false,
27
+ "tokenizer_class": "Qwen2Tokenizer",
28
+ "unk_token": null
29
+ }
random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/trainer_state.json ADDED
@@ -0,0 +1,307 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "best_global_step": null,
3
+ "best_metric": null,
4
+ "best_model_checkpoint": null,
5
+ "epoch": 3.0,
6
+ "eval_steps": 500,
7
+ "global_step": 1224,
8
+ "is_hyper_param_search": false,
9
+ "is_local_process_zero": true,
10
+ "is_world_process_zero": true,
11
+ "log_history": [
12
+ {
13
+ "entropy": 1.6558652856945992,
14
+ "epoch": 0.12277470841006753,
15
+ "grad_norm": 0.8294193148612976,
16
+ "learning_rate": 2.604481765338724e-05,
17
+ "loss": 1.5730807495117187,
18
+ "mean_token_accuracy": 0.6558220008015633,
19
+ "num_tokens": 137019.0,
20
+ "step": 50
21
+ },
22
+ {
23
+ "entropy": 0.8554385647177696,
24
+ "epoch": 0.24554941682013506,
25
+ "grad_norm": 1.029894232749939,
26
+ "learning_rate": 5.2621162197659935e-05,
27
+ "loss": 0.7961322021484375,
28
+ "mean_token_accuracy": 0.7755533090233803,
29
+ "num_tokens": 267448.0,
30
+ "step": 100
31
+ },
32
+ {
33
+ "entropy": 0.7380270153284073,
34
+ "epoch": 0.3683241252302026,
35
+ "grad_norm": 0.6096898913383484,
36
+ "learning_rate": 7.919750674193263e-05,
37
+ "loss": 0.6843100738525391,
38
+ "mean_token_accuracy": 0.7998733684420586,
39
+ "num_tokens": 409071.0,
40
+ "step": 150
41
+ },
42
+ {
43
+ "entropy": 0.6932123881578446,
44
+ "epoch": 0.4910988336402701,
45
+ "grad_norm": 0.5934288501739502,
46
+ "learning_rate": 0.00010577385128620532,
47
+ "loss": 0.6458084106445312,
48
+ "mean_token_accuracy": 0.8083393195271492,
49
+ "num_tokens": 542207.0,
50
+ "step": 200
51
+ },
52
+ {
53
+ "entropy": 0.6840551143884659,
54
+ "epoch": 0.6138735420503376,
55
+ "grad_norm": 0.4665544331073761,
56
+ "learning_rate": 0.00013235019583047802,
57
+ "loss": 0.6333241653442383,
58
+ "mean_token_accuracy": 0.8122684139013291,
59
+ "num_tokens": 679619.0,
60
+ "step": 250
61
+ },
62
+ {
63
+ "entropy": 0.6649293206632138,
64
+ "epoch": 0.7366482504604052,
65
+ "grad_norm": 0.47893014550209045,
66
+ "learning_rate": 0.00015892654037475069,
67
+ "loss": 0.6107040786743164,
68
+ "mean_token_accuracy": 0.8165940269827843,
69
+ "num_tokens": 815397.0,
70
+ "step": 300
71
+ },
72
+ {
73
+ "entropy": 0.6492742404341698,
74
+ "epoch": 0.8594229588704727,
75
+ "grad_norm": 0.6059293746948242,
76
+ "learning_rate": 0.0001855028849190234,
77
+ "loss": 0.597186050415039,
78
+ "mean_token_accuracy": 0.8217281407117843,
79
+ "num_tokens": 949315.0,
80
+ "step": 350
81
+ },
82
+ {
83
+ "entropy": 0.642182088047266,
84
+ "epoch": 0.9821976672805403,
85
+ "grad_norm": 0.3793066143989563,
86
+ "learning_rate": 0.0002120792294632961,
87
+ "loss": 0.5949800491333008,
88
+ "mean_token_accuracy": 0.8229166463017463,
89
+ "num_tokens": 1082903.0,
90
+ "step": 400
91
+ },
92
+ {
93
+ "epoch": 1.0,
94
+ "eval_entropy": 0.6541631272860936,
95
+ "eval_loss": 0.5903413891792297,
96
+ "eval_mean_token_accuracy": 0.8252546640804835,
97
+ "eval_num_tokens": 1101768.0,
98
+ "eval_runtime": 108.9056,
99
+ "eval_samples_per_second": 12.818,
100
+ "eval_steps_per_second": 1.607,
101
+ "step": 408
102
+ },
103
+ {
104
+ "entropy": 0.6077755300829253,
105
+ "epoch": 1.1031307550644567,
106
+ "grad_norm": 0.5028847455978394,
107
+ "learning_rate": 0.00021679626884558217,
108
+ "loss": 0.5617346954345703,
109
+ "mean_token_accuracy": 0.8301527442665875,
110
+ "num_tokens": 1223073.0,
111
+ "step": 450
112
+ },
113
+ {
114
+ "entropy": 0.5860125370323658,
115
+ "epoch": 1.2259054634745243,
116
+ "grad_norm": 0.38494572043418884,
117
+ "learning_rate": 0.00021653451093163906,
118
+ "loss": 0.5437137985229492,
119
+ "mean_token_accuracy": 0.8346565261483192,
120
+ "num_tokens": 1360978.0,
121
+ "step": 500
122
+ },
123
+ {
124
+ "entropy": 0.5945393888652325,
125
+ "epoch": 1.3486801718845918,
126
+ "grad_norm": 0.5123576521873474,
127
+ "learning_rate": 0.00021607496224450087,
128
+ "loss": 0.54191650390625,
129
+ "mean_token_accuracy": 0.8335164493322372,
130
+ "num_tokens": 1492506.0,
131
+ "step": 550
132
+ },
133
+ {
134
+ "entropy": 0.6083934807777405,
135
+ "epoch": 1.4714548802946594,
136
+ "grad_norm": 0.46661439538002014,
137
+ "learning_rate": 0.000215418463597734,
138
+ "loss": 0.5550478744506836,
139
+ "mean_token_accuracy": 0.8311341696977615,
140
+ "num_tokens": 1622275.0,
141
+ "step": 600
142
+ },
143
+ {
144
+ "entropy": 0.5862340961396694,
145
+ "epoch": 1.5942295887047269,
146
+ "grad_norm": 0.49572765827178955,
147
+ "learning_rate": 0.00021456621615453177,
148
+ "loss": 0.5297146606445312,
149
+ "mean_token_accuracy": 0.8360210624337197,
150
+ "num_tokens": 1756450.0,
151
+ "step": 650
152
+ },
153
+ {
154
+ "entropy": 0.5832360745966434,
155
+ "epoch": 1.7170042971147943,
156
+ "grad_norm": 0.35520103573799133,
157
+ "learning_rate": 0.0002135197792300053,
158
+ "loss": 0.5287040328979492,
159
+ "mean_token_accuracy": 0.8368027776479721,
160
+ "num_tokens": 1890634.0,
161
+ "step": 700
162
+ },
163
+ {
164
+ "entropy": 0.5792283065617084,
165
+ "epoch": 1.839779005524862,
166
+ "grad_norm": 0.3548352122306824,
167
+ "learning_rate": 0.00021228106743818178,
168
+ "loss": 0.5251237487792969,
169
+ "mean_token_accuracy": 0.8396679371595382,
170
+ "num_tokens": 2028046.0,
171
+ "step": 750
172
+ },
173
+ {
174
+ "entropy": 0.5694268302619457,
175
+ "epoch": 1.9625537139349294,
176
+ "grad_norm": 0.3476282060146332,
177
+ "learning_rate": 0.00021085234718892933,
178
+ "loss": 0.5189918899536132,
179
+ "mean_token_accuracy": 0.8398650133609772,
180
+ "num_tokens": 2163703.0,
181
+ "step": 800
182
+ },
183
+ {
184
+ "epoch": 2.0,
185
+ "eval_entropy": 0.5552395452771868,
186
+ "eval_loss": 0.5418588519096375,
187
+ "eval_mean_token_accuracy": 0.8379638079234532,
188
+ "eval_num_tokens": 2203536.0,
189
+ "eval_runtime": 108.6649,
190
+ "eval_samples_per_second": 12.847,
191
+ "eval_steps_per_second": 1.61,
192
+ "step": 816
193
+ },
194
+ {
195
+ "entropy": 0.5162206395023365,
196
+ "epoch": 2.0834868017188457,
197
+ "grad_norm": 0.3772931396961212,
198
+ "learning_rate": 0.0002092362325412188,
199
+ "loss": 0.4619992446899414,
200
+ "mean_token_accuracy": 0.8528054483650904,
201
+ "num_tokens": 2298231.0,
202
+ "step": 850
203
+ },
204
+ {
205
+ "entropy": 0.5109021583199501,
206
+ "epoch": 2.2062615101289134,
207
+ "grad_norm": 0.4661090672016144,
208
+ "learning_rate": 0.000207435680420309,
209
+ "loss": 0.4571444702148437,
210
+ "mean_token_accuracy": 0.8552933797240257,
211
+ "num_tokens": 2427608.0,
212
+ "step": 900
213
+ },
214
+ {
215
+ "entropy": 0.5012516237795352,
216
+ "epoch": 2.329036218538981,
217
+ "grad_norm": 0.5193169713020325,
218
+ "learning_rate": 0.0002054539852076065,
219
+ "loss": 0.45432735443115235,
220
+ "mean_token_accuracy": 0.8564087572693825,
221
+ "num_tokens": 2565982.0,
222
+ "step": 950
223
+ },
224
+ {
225
+ "entropy": 0.5173932652175427,
226
+ "epoch": 2.4518109269490487,
227
+ "grad_norm": 0.4666334390640259,
228
+ "learning_rate": 0.00020329477271309812,
229
+ "loss": 0.4616986083984375,
230
+ "mean_token_accuracy": 0.8532203987240792,
231
+ "num_tokens": 2695545.0,
232
+ "step": 1000
233
+ },
234
+ {
235
+ "entropy": 0.5169526914507151,
236
+ "epoch": 2.574585635359116,
237
+ "grad_norm": 0.47921115159988403,
238
+ "learning_rate": 0.0002009619935413857,
239
+ "loss": 0.45737281799316404,
240
+ "mean_token_accuracy": 0.8535552659630775,
241
+ "num_tokens": 2829781.0,
242
+ "step": 1050
243
+ },
244
+ {
245
+ "entropy": 0.50618562489748,
246
+ "epoch": 2.6973603437691835,
247
+ "grad_norm": 0.3407684862613678,
248
+ "learning_rate": 0.00019845991586345935,
249
+ "loss": 0.45972068786621095,
250
+ "mean_token_accuracy": 0.8532172521948814,
251
+ "num_tokens": 2970162.0,
252
+ "step": 1100
253
+ },
254
+ {
255
+ "entropy": 0.5158357314765454,
256
+ "epoch": 2.820135052179251,
257
+ "grad_norm": 0.47901391983032227,
258
+ "learning_rate": 0.00019579311760743563,
259
+ "loss": 0.46119583129882813,
260
+ "mean_token_accuracy": 0.8536588314175606,
261
+ "num_tokens": 3103775.0,
262
+ "step": 1150
263
+ },
264
+ {
265
+ "entropy": 0.5069115920364857,
266
+ "epoch": 2.942909760589319,
267
+ "grad_norm": 0.4287651479244232,
268
+ "learning_rate": 0.00019296647808254838,
269
+ "loss": 0.45447597503662107,
270
+ "mean_token_accuracy": 0.856622197329998,
271
+ "num_tokens": 3241634.0,
272
+ "step": 1200
273
+ },
274
+ {
275
+ "epoch": 3.0,
276
+ "eval_entropy": 0.5097703012398311,
277
+ "eval_loss": 0.5291272401809692,
278
+ "eval_mean_token_accuracy": 0.8438761404582432,
279
+ "eval_num_tokens": 3305304.0,
280
+ "eval_runtime": 108.6946,
281
+ "eval_samples_per_second": 12.843,
282
+ "eval_steps_per_second": 1.61,
283
+ "step": 1224
284
+ }
285
+ ],
286
+ "logging_steps": 50,
287
+ "max_steps": 4080,
288
+ "num_input_tokens_seen": 0,
289
+ "num_train_epochs": 10,
290
+ "save_steps": 500,
291
+ "stateful_callbacks": {
292
+ "TrainerControl": {
293
+ "args": {
294
+ "should_epoch_stop": false,
295
+ "should_evaluate": false,
296
+ "should_log": false,
297
+ "should_save": true,
298
+ "should_training_stop": false
299
+ },
300
+ "attributes": {}
301
+ }
302
+ },
303
+ "total_flos": 5.376308378051789e+17,
304
+ "train_batch_size": 4,
305
+ "trial_name": null,
306
+ "trial_params": null
307
+ }
random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1632/README.md ADDED
@@ -0,0 +1,209 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen3-14B-Base
3
+ library_name: peft
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - base_model:adapter:Qwen/Qwen3-14B-Base
7
+ - lora
8
+ - sft
9
+ - transformers
10
+ - trl
11
+ ---
12
+
13
+ # Model Card for Model ID
14
+
15
+ <!-- Provide a quick summary of what the model is/does. -->
16
+
17
+
18
+
19
+ ## Model Details
20
+
21
+ ### Model Description
22
+
23
+ <!-- Provide a longer summary of what this model is. -->
24
+
25
+
26
+
27
+ - **Developed by:** [More Information Needed]
28
+ - **Funded by [optional]:** [More Information Needed]
29
+ - **Shared by [optional]:** [More Information Needed]
30
+ - **Model type:** [More Information Needed]
31
+ - **Language(s) (NLP):** [More Information Needed]
32
+ - **License:** [More Information Needed]
33
+ - **Finetuned from model [optional]:** [More Information Needed]
34
+
35
+ ### Model Sources [optional]
36
+
37
+ <!-- Provide the basic links for the model. -->
38
+
39
+ - **Repository:** [More Information Needed]
40
+ - **Paper [optional]:** [More Information Needed]
41
+ - **Demo [optional]:** [More Information Needed]
42
+
43
+ ## Uses
44
+
45
+ <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
46
+
47
+ ### Direct Use
48
+
49
+ <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
50
+
51
+ [More Information Needed]
52
+
53
+ ### Downstream Use [optional]
54
+
55
+ <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
56
+
57
+ [More Information Needed]
58
+
59
+ ### Out-of-Scope Use
60
+
61
+ <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
62
+
63
+ [More Information Needed]
64
+
65
+ ## Bias, Risks, and Limitations
66
+
67
+ <!-- This section is meant to convey both technical and sociotechnical limitations. -->
68
+
69
+ [More Information Needed]
70
+
71
+ ### Recommendations
72
+
73
+ <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
74
+
75
+ Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
76
+
77
+ ## How to Get Started with the Model
78
+
79
+ Use the code below to get started with the model.
80
+
81
+ [More Information Needed]
82
+
83
+ ## Training Details
84
+
85
+ ### Training Data
86
+
87
+ <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
88
+
89
+ [More Information Needed]
90
+
91
+ ### Training Procedure
92
+
93
+ <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
94
+
95
+ #### Preprocessing [optional]
96
+
97
+ [More Information Needed]
98
+
99
+
100
+ #### Training Hyperparameters
101
+
102
+ - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
103
+
104
+ #### Speeds, Sizes, Times [optional]
105
+
106
+ <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
107
+
108
+ [More Information Needed]
109
+
110
+ ## Evaluation
111
+
112
+ <!-- This section describes the evaluation protocols and provides the results. -->
113
+
114
+ ### Testing Data, Factors & Metrics
115
+
116
+ #### Testing Data
117
+
118
+ <!-- This should link to a Dataset Card if possible. -->
119
+
120
+ [More Information Needed]
121
+
122
+ #### Factors
123
+
124
+ <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
125
+
126
+ [More Information Needed]
127
+
128
+ #### Metrics
129
+
130
+ <!-- These are the evaluation metrics being used, ideally with a description of why. -->
131
+
132
+ [More Information Needed]
133
+
134
+ ### Results
135
+
136
+ [More Information Needed]
137
+
138
+ #### Summary
139
+
140
+
141
+
142
+ ## Model Examination [optional]
143
+
144
+ <!-- Relevant interpretability work for the model goes here -->
145
+
146
+ [More Information Needed]
147
+
148
+ ## Environmental Impact
149
+
150
+ <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
151
+
152
+ Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
153
+
154
+ - **Hardware Type:** [More Information Needed]
155
+ - **Hours used:** [More Information Needed]
156
+ - **Cloud Provider:** [More Information Needed]
157
+ - **Compute Region:** [More Information Needed]
158
+ - **Carbon Emitted:** [More Information Needed]
159
+
160
+ ## Technical Specifications [optional]
161
+
162
+ ### Model Architecture and Objective
163
+
164
+ [More Information Needed]
165
+
166
+ ### Compute Infrastructure
167
+
168
+ [More Information Needed]
169
+
170
+ #### Hardware
171
+
172
+ [More Information Needed]
173
+
174
+ #### Software
175
+
176
+ [More Information Needed]
177
+
178
+ ## Citation [optional]
179
+
180
+ <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
181
+
182
+ **BibTeX:**
183
+
184
+ [More Information Needed]
185
+
186
+ **APA:**
187
+
188
+ [More Information Needed]
189
+
190
+ ## Glossary [optional]
191
+
192
+ <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
193
+
194
+ [More Information Needed]
195
+
196
+ ## More Information [optional]
197
+
198
+ [More Information Needed]
199
+
200
+ ## Model Card Authors [optional]
201
+
202
+ [More Information Needed]
203
+
204
+ ## Model Card Contact
205
+
206
+ [More Information Needed]
207
+ ### Framework versions
208
+
209
+ - PEFT 0.19.1