Add files using upload-large-folder tool
Browse filesThis view is limited to 50 files because it contains too many changes. See raw diff
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test1/README.md +58 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/README.md +58 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/README.md +209 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/adapter_config.json +48 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/chat_template.jinja +85 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/tokenizer_config.json +29 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/trainer_state.json +297 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/README.md +209 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/adapter_config.json +48 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/chat_template.jinja +85 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/tokenizer_config.json +29 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/trainer_state.json +388 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/README.md +209 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/adapter_config.json +48 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/chat_template.jinja +85 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/tokenizer_config.json +29 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/trainer_state.json +469 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/README.md +209 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/adapter_config.json +48 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/chat_template.jinja +85 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/tokenizer_config.json +29 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/trainer_state.json +560 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/README.md +209 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/adapter_config.json +48 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/chat_template.jinja +85 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/tokenizer_config.json +29 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/trainer_state.json +651 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/README.md +209 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/adapter_config.json +48 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/chat_template.jinja +85 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/tokenizer_config.json +29 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/trainer_state.json +742 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/README.md +209 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/adapter_config.json +48 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/chat_template.jinja +85 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/tokenizer_config.json +29 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/trainer_state.json +833 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/README.md +209 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/adapter_config.json +48 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/chat_template.jinja +85 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/tokenizer_config.json +29 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/trainer_state.json +115 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3890/README.md +209 -0
- productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3890/adapter_config.json +48 -0
- random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/README.md +209 -0
- random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/adapter_config.json +48 -0
- random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/chat_template.jinja +85 -0
- random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/tokenizer_config.json +29 -0
- random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/trainer_state.json +307 -0
- random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1632/README.md +209 -0
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test1/README.md
ADDED
|
@@ -0,0 +1,58 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen3-14B-Base
|
| 3 |
+
library_name: transformers
|
| 4 |
+
model_name: Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test1
|
| 5 |
+
tags:
|
| 6 |
+
- generated_from_trainer
|
| 7 |
+
- trl
|
| 8 |
+
- sft
|
| 9 |
+
licence: license
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
# Model Card for Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test1
|
| 13 |
+
|
| 14 |
+
This model is a fine-tuned version of [Qwen/Qwen3-14B-Base](https://huggingface.co/Qwen/Qwen3-14B-Base).
|
| 15 |
+
It has been trained using [TRL](https://github.com/huggingface/trl).
|
| 16 |
+
|
| 17 |
+
## Quick start
|
| 18 |
+
|
| 19 |
+
```python
|
| 20 |
+
from transformers import pipeline
|
| 21 |
+
|
| 22 |
+
question = "If you had a time machine, but could only go to the past or the future once and never return, which would you choose and why?"
|
| 23 |
+
generator = pipeline("text-generation", model="None", device="cuda")
|
| 24 |
+
output = generator([{"role": "user", "content": question}], max_new_tokens=128, return_full_text=False)[0]
|
| 25 |
+
print(output["generated_text"])
|
| 26 |
+
```
|
| 27 |
+
|
| 28 |
+
## Training procedure
|
| 29 |
+
|
| 30 |
+
[<img src="https://raw.githubusercontent.com/wandb/assets/main/wandb-github-badge-28.svg" alt="Visualize in Weights & Biases" width="150" height="24"/>](https://wandb.ai/katriin-kukk/Cross_lingual_morphological_generalization/runs/tjg90wvc)
|
| 31 |
+
|
| 32 |
+
|
| 33 |
+
|
| 34 |
+
This model was trained with SFT.
|
| 35 |
+
|
| 36 |
+
### Framework versions
|
| 37 |
+
|
| 38 |
+
- TRL: 0.29.0
|
| 39 |
+
- Transformers: 5.5.4
|
| 40 |
+
- Pytorch: 2.10.0
|
| 41 |
+
- Datasets: 4.6.1
|
| 42 |
+
- Tokenizers: 0.22.2
|
| 43 |
+
|
| 44 |
+
## Citations
|
| 45 |
+
|
| 46 |
+
|
| 47 |
+
|
| 48 |
+
Cite TRL as:
|
| 49 |
+
|
| 50 |
+
```bibtex
|
| 51 |
+
@software{vonwerra2020trl,
|
| 52 |
+
title = {{TRL: Transformers Reinforcement Learning}},
|
| 53 |
+
author = {von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and Beeching, Edward and Thrush, Tristan and Lambert, Nathan and Huang, Shengyi and Rasul, Kashif and Gallouédec, Quentin},
|
| 54 |
+
license = {Apache-2.0},
|
| 55 |
+
url = {https://github.com/huggingface/trl},
|
| 56 |
+
year = {2020}
|
| 57 |
+
}
|
| 58 |
+
```
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/README.md
ADDED
|
@@ -0,0 +1,58 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen3-14B-Base
|
| 3 |
+
library_name: transformers
|
| 4 |
+
model_name: Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2
|
| 5 |
+
tags:
|
| 6 |
+
- generated_from_trainer
|
| 7 |
+
- trl
|
| 8 |
+
- sft
|
| 9 |
+
licence: license
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
# Model Card for Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2
|
| 13 |
+
|
| 14 |
+
This model is a fine-tuned version of [Qwen/Qwen3-14B-Base](https://huggingface.co/Qwen/Qwen3-14B-Base).
|
| 15 |
+
It has been trained using [TRL](https://github.com/huggingface/trl).
|
| 16 |
+
|
| 17 |
+
## Quick start
|
| 18 |
+
|
| 19 |
+
```python
|
| 20 |
+
from transformers import pipeline
|
| 21 |
+
|
| 22 |
+
question = "If you had a time machine, but could only go to the past or the future once and never return, which would you choose and why?"
|
| 23 |
+
generator = pipeline("text-generation", model="None", device="cuda")
|
| 24 |
+
output = generator([{"role": "user", "content": question}], max_new_tokens=128, return_full_text=False)[0]
|
| 25 |
+
print(output["generated_text"])
|
| 26 |
+
```
|
| 27 |
+
|
| 28 |
+
## Training procedure
|
| 29 |
+
|
| 30 |
+
[<img src="https://raw.githubusercontent.com/wandb/assets/main/wandb-github-badge-28.svg" alt="Visualize in Weights & Biases" width="150" height="24"/>](https://wandb.ai/katriin-kukk/Cross_lingual_morphological_generalization/runs/3ulga1iu)
|
| 31 |
+
|
| 32 |
+
|
| 33 |
+
|
| 34 |
+
This model was trained with SFT.
|
| 35 |
+
|
| 36 |
+
### Framework versions
|
| 37 |
+
|
| 38 |
+
- TRL: 0.29.0
|
| 39 |
+
- Transformers: 5.5.4
|
| 40 |
+
- Pytorch: 2.10.0
|
| 41 |
+
- Datasets: 4.6.1
|
| 42 |
+
- Tokenizers: 0.22.2
|
| 43 |
+
|
| 44 |
+
## Citations
|
| 45 |
+
|
| 46 |
+
|
| 47 |
+
|
| 48 |
+
Cite TRL as:
|
| 49 |
+
|
| 50 |
+
```bibtex
|
| 51 |
+
@software{vonwerra2020trl,
|
| 52 |
+
title = {{TRL: Transformers Reinforcement Learning}},
|
| 53 |
+
author = {von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and Beeching, Edward and Thrush, Tristan and Lambert, Nathan and Huang, Shengyi and Rasul, Kashif and Gallouédec, Quentin},
|
| 54 |
+
license = {Apache-2.0},
|
| 55 |
+
url = {https://github.com/huggingface/trl},
|
| 56 |
+
year = {2020}
|
| 57 |
+
}
|
| 58 |
+
```
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/README.md
ADDED
|
@@ -0,0 +1,209 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen3-14B-Base
|
| 3 |
+
library_name: peft
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:Qwen/Qwen3-14B-Base
|
| 7 |
+
- lora
|
| 8 |
+
- sft
|
| 9 |
+
- transformers
|
| 10 |
+
- trl
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Model Card for Model ID
|
| 14 |
+
|
| 15 |
+
<!-- Provide a quick summary of what the model is/does. -->
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
## Model Details
|
| 20 |
+
|
| 21 |
+
### Model Description
|
| 22 |
+
|
| 23 |
+
<!-- Provide a longer summary of what this model is. -->
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
- **Developed by:** [More Information Needed]
|
| 28 |
+
- **Funded by [optional]:** [More Information Needed]
|
| 29 |
+
- **Shared by [optional]:** [More Information Needed]
|
| 30 |
+
- **Model type:** [More Information Needed]
|
| 31 |
+
- **Language(s) (NLP):** [More Information Needed]
|
| 32 |
+
- **License:** [More Information Needed]
|
| 33 |
+
- **Finetuned from model [optional]:** [More Information Needed]
|
| 34 |
+
|
| 35 |
+
### Model Sources [optional]
|
| 36 |
+
|
| 37 |
+
<!-- Provide the basic links for the model. -->
|
| 38 |
+
|
| 39 |
+
- **Repository:** [More Information Needed]
|
| 40 |
+
- **Paper [optional]:** [More Information Needed]
|
| 41 |
+
- **Demo [optional]:** [More Information Needed]
|
| 42 |
+
|
| 43 |
+
## Uses
|
| 44 |
+
|
| 45 |
+
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
| 46 |
+
|
| 47 |
+
### Direct Use
|
| 48 |
+
|
| 49 |
+
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
| 50 |
+
|
| 51 |
+
[More Information Needed]
|
| 52 |
+
|
| 53 |
+
### Downstream Use [optional]
|
| 54 |
+
|
| 55 |
+
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
| 56 |
+
|
| 57 |
+
[More Information Needed]
|
| 58 |
+
|
| 59 |
+
### Out-of-Scope Use
|
| 60 |
+
|
| 61 |
+
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
| 62 |
+
|
| 63 |
+
[More Information Needed]
|
| 64 |
+
|
| 65 |
+
## Bias, Risks, and Limitations
|
| 66 |
+
|
| 67 |
+
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
| 68 |
+
|
| 69 |
+
[More Information Needed]
|
| 70 |
+
|
| 71 |
+
### Recommendations
|
| 72 |
+
|
| 73 |
+
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
| 74 |
+
|
| 75 |
+
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
| 76 |
+
|
| 77 |
+
## How to Get Started with the Model
|
| 78 |
+
|
| 79 |
+
Use the code below to get started with the model.
|
| 80 |
+
|
| 81 |
+
[More Information Needed]
|
| 82 |
+
|
| 83 |
+
## Training Details
|
| 84 |
+
|
| 85 |
+
### Training Data
|
| 86 |
+
|
| 87 |
+
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
| 88 |
+
|
| 89 |
+
[More Information Needed]
|
| 90 |
+
|
| 91 |
+
### Training Procedure
|
| 92 |
+
|
| 93 |
+
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
| 94 |
+
|
| 95 |
+
#### Preprocessing [optional]
|
| 96 |
+
|
| 97 |
+
[More Information Needed]
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
#### Training Hyperparameters
|
| 101 |
+
|
| 102 |
+
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
| 103 |
+
|
| 104 |
+
#### Speeds, Sizes, Times [optional]
|
| 105 |
+
|
| 106 |
+
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
| 107 |
+
|
| 108 |
+
[More Information Needed]
|
| 109 |
+
|
| 110 |
+
## Evaluation
|
| 111 |
+
|
| 112 |
+
<!-- This section describes the evaluation protocols and provides the results. -->
|
| 113 |
+
|
| 114 |
+
### Testing Data, Factors & Metrics
|
| 115 |
+
|
| 116 |
+
#### Testing Data
|
| 117 |
+
|
| 118 |
+
<!-- This should link to a Dataset Card if possible. -->
|
| 119 |
+
|
| 120 |
+
[More Information Needed]
|
| 121 |
+
|
| 122 |
+
#### Factors
|
| 123 |
+
|
| 124 |
+
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
| 125 |
+
|
| 126 |
+
[More Information Needed]
|
| 127 |
+
|
| 128 |
+
#### Metrics
|
| 129 |
+
|
| 130 |
+
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
| 131 |
+
|
| 132 |
+
[More Information Needed]
|
| 133 |
+
|
| 134 |
+
### Results
|
| 135 |
+
|
| 136 |
+
[More Information Needed]
|
| 137 |
+
|
| 138 |
+
#### Summary
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
## Model Examination [optional]
|
| 143 |
+
|
| 144 |
+
<!-- Relevant interpretability work for the model goes here -->
|
| 145 |
+
|
| 146 |
+
[More Information Needed]
|
| 147 |
+
|
| 148 |
+
## Environmental Impact
|
| 149 |
+
|
| 150 |
+
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
| 151 |
+
|
| 152 |
+
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
| 153 |
+
|
| 154 |
+
- **Hardware Type:** [More Information Needed]
|
| 155 |
+
- **Hours used:** [More Information Needed]
|
| 156 |
+
- **Cloud Provider:** [More Information Needed]
|
| 157 |
+
- **Compute Region:** [More Information Needed]
|
| 158 |
+
- **Carbon Emitted:** [More Information Needed]
|
| 159 |
+
|
| 160 |
+
## Technical Specifications [optional]
|
| 161 |
+
|
| 162 |
+
### Model Architecture and Objective
|
| 163 |
+
|
| 164 |
+
[More Information Needed]
|
| 165 |
+
|
| 166 |
+
### Compute Infrastructure
|
| 167 |
+
|
| 168 |
+
[More Information Needed]
|
| 169 |
+
|
| 170 |
+
#### Hardware
|
| 171 |
+
|
| 172 |
+
[More Information Needed]
|
| 173 |
+
|
| 174 |
+
#### Software
|
| 175 |
+
|
| 176 |
+
[More Information Needed]
|
| 177 |
+
|
| 178 |
+
## Citation [optional]
|
| 179 |
+
|
| 180 |
+
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
| 181 |
+
|
| 182 |
+
**BibTeX:**
|
| 183 |
+
|
| 184 |
+
[More Information Needed]
|
| 185 |
+
|
| 186 |
+
**APA:**
|
| 187 |
+
|
| 188 |
+
[More Information Needed]
|
| 189 |
+
|
| 190 |
+
## Glossary [optional]
|
| 191 |
+
|
| 192 |
+
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
| 193 |
+
|
| 194 |
+
[More Information Needed]
|
| 195 |
+
|
| 196 |
+
## More Information [optional]
|
| 197 |
+
|
| 198 |
+
[More Information Needed]
|
| 199 |
+
|
| 200 |
+
## Model Card Authors [optional]
|
| 201 |
+
|
| 202 |
+
[More Information Needed]
|
| 203 |
+
|
| 204 |
+
## Model Card Contact
|
| 205 |
+
|
| 206 |
+
[More Information Needed]
|
| 207 |
+
### Framework versions
|
| 208 |
+
|
| 209 |
+
- PEFT 0.19.1
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/adapter_config.json
ADDED
|
@@ -0,0 +1,48 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "Qwen/Qwen3-14B-Base",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"layer_replication": null,
|
| 16 |
+
"layers_pattern": null,
|
| 17 |
+
"layers_to_transform": null,
|
| 18 |
+
"loftq_config": {},
|
| 19 |
+
"lora_alpha": 256,
|
| 20 |
+
"lora_bias": false,
|
| 21 |
+
"lora_dropout": 0.00237968804112545,
|
| 22 |
+
"lora_ga_config": null,
|
| 23 |
+
"megatron_config": null,
|
| 24 |
+
"megatron_core": "megatron.core",
|
| 25 |
+
"modules_to_save": null,
|
| 26 |
+
"peft_type": "LORA",
|
| 27 |
+
"peft_version": "0.19.1",
|
| 28 |
+
"qalora_group_size": 16,
|
| 29 |
+
"r": 128,
|
| 30 |
+
"rank_pattern": {},
|
| 31 |
+
"revision": null,
|
| 32 |
+
"target_modules": [
|
| 33 |
+
"gate_proj",
|
| 34 |
+
"q_proj",
|
| 35 |
+
"up_proj",
|
| 36 |
+
"o_proj",
|
| 37 |
+
"v_proj",
|
| 38 |
+
"k_proj",
|
| 39 |
+
"down_proj"
|
| 40 |
+
],
|
| 41 |
+
"target_parameters": null,
|
| 42 |
+
"task_type": "CAUSAL_LM",
|
| 43 |
+
"trainable_token_indices": null,
|
| 44 |
+
"use_bdlora": null,
|
| 45 |
+
"use_dora": false,
|
| 46 |
+
"use_qalora": false,
|
| 47 |
+
"use_rslora": false
|
| 48 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/chat_template.jinja
ADDED
|
@@ -0,0 +1,85 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- if tools %}
|
| 2 |
+
{{- '<|im_start|>system\n' }}
|
| 3 |
+
{%- if messages[0].role == 'system' %}
|
| 4 |
+
{{- messages[0].content + '\n\n' }}
|
| 5 |
+
{%- endif %}
|
| 6 |
+
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
| 7 |
+
{%- for tool in tools %}
|
| 8 |
+
{{- "\n" }}
|
| 9 |
+
{{- tool | tojson }}
|
| 10 |
+
{%- endfor %}
|
| 11 |
+
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
| 12 |
+
{%- else %}
|
| 13 |
+
{%- if messages[0].role == 'system' %}
|
| 14 |
+
{{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
|
| 15 |
+
{%- endif %}
|
| 16 |
+
{%- endif %}
|
| 17 |
+
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
| 18 |
+
{%- for message in messages[::-1] %}
|
| 19 |
+
{%- set index = (messages|length - 1) - loop.index0 %}
|
| 20 |
+
{%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
|
| 21 |
+
{%- set ns.multi_step_tool = false %}
|
| 22 |
+
{%- set ns.last_query_index = index %}
|
| 23 |
+
{%- endif %}
|
| 24 |
+
{%- endfor %}
|
| 25 |
+
{%- for message in messages %}
|
| 26 |
+
{%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
|
| 27 |
+
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
|
| 28 |
+
{%- elif message.role == "assistant" %}
|
| 29 |
+
{%- set content = message.content %}
|
| 30 |
+
{%- set reasoning_content = '' %}
|
| 31 |
+
{%- if message.reasoning_content is defined and message.reasoning_content is not none %}
|
| 32 |
+
{%- set reasoning_content = message.reasoning_content %}
|
| 33 |
+
{%- else %}
|
| 34 |
+
{%- if '</think>' in message.content %}
|
| 35 |
+
{%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
|
| 36 |
+
{%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
| 37 |
+
{%- endif %}
|
| 38 |
+
{%- endif %}
|
| 39 |
+
{%- if loop.index0 > ns.last_query_index %}
|
| 40 |
+
{%- if loop.last or (not loop.last and reasoning_content) %}
|
| 41 |
+
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
|
| 42 |
+
{%- else %}
|
| 43 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 44 |
+
{%- endif %}
|
| 45 |
+
{%- else %}
|
| 46 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 47 |
+
{%- endif %}
|
| 48 |
+
{%- if message.tool_calls %}
|
| 49 |
+
{%- for tool_call in message.tool_calls %}
|
| 50 |
+
{%- if (loop.first and content) or (not loop.first) %}
|
| 51 |
+
{{- '\n' }}
|
| 52 |
+
{%- endif %}
|
| 53 |
+
{%- if tool_call.function %}
|
| 54 |
+
{%- set tool_call = tool_call.function %}
|
| 55 |
+
{%- endif %}
|
| 56 |
+
{{- '<tool_call>\n{"name": "' }}
|
| 57 |
+
{{- tool_call.name }}
|
| 58 |
+
{{- '", "arguments": ' }}
|
| 59 |
+
{%- if tool_call.arguments is string %}
|
| 60 |
+
{{- tool_call.arguments }}
|
| 61 |
+
{%- else %}
|
| 62 |
+
{{- tool_call.arguments | tojson }}
|
| 63 |
+
{%- endif %}
|
| 64 |
+
{{- '}\n</tool_call>' }}
|
| 65 |
+
{%- endfor %}
|
| 66 |
+
{%- endif %}
|
| 67 |
+
{{- '<|im_end|>\n' }}
|
| 68 |
+
{%- elif message.role == "tool" %}
|
| 69 |
+
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
|
| 70 |
+
{{- '<|im_start|>user' }}
|
| 71 |
+
{%- endif %}
|
| 72 |
+
{{- '\n<tool_response>\n' }}
|
| 73 |
+
{{- message.content }}
|
| 74 |
+
{{- '\n</tool_response>' }}
|
| 75 |
+
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
| 76 |
+
{{- '<|im_end|>\n' }}
|
| 77 |
+
{%- endif %}
|
| 78 |
+
{%- endif %}
|
| 79 |
+
{%- endfor %}
|
| 80 |
+
{%- if add_generation_prompt %}
|
| 81 |
+
{{- '<|im_start|>assistant\n' }}
|
| 82 |
+
{%- if enable_thinking is defined and enable_thinking is false %}
|
| 83 |
+
{{- '<think>\n\n</think>\n\n' }}
|
| 84 |
+
{%- endif %}
|
| 85 |
+
{%- endif %}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/tokenizer_config.json
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"backend": "tokenizers",
|
| 4 |
+
"bos_token": null,
|
| 5 |
+
"clean_up_tokenization_spaces": false,
|
| 6 |
+
"eos_token": "<|endoftext|>",
|
| 7 |
+
"errors": "replace",
|
| 8 |
+
"extra_special_tokens": [
|
| 9 |
+
"<|im_start|>",
|
| 10 |
+
"<|im_end|>",
|
| 11 |
+
"<|object_ref_start|>",
|
| 12 |
+
"<|object_ref_end|>",
|
| 13 |
+
"<|box_start|>",
|
| 14 |
+
"<|box_end|>",
|
| 15 |
+
"<|quad_start|>",
|
| 16 |
+
"<|quad_end|>",
|
| 17 |
+
"<|vision_start|>",
|
| 18 |
+
"<|vision_end|>",
|
| 19 |
+
"<|vision_pad|>",
|
| 20 |
+
"<|image_pad|>",
|
| 21 |
+
"<|video_pad|>"
|
| 22 |
+
],
|
| 23 |
+
"is_local": false,
|
| 24 |
+
"model_max_length": 131072,
|
| 25 |
+
"pad_token": "<|endoftext|>",
|
| 26 |
+
"split_special_tokens": false,
|
| 27 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 28 |
+
"unk_token": null
|
| 29 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1167/trainer_state.json
ADDED
|
@@ -0,0 +1,297 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"best_global_step": null,
|
| 3 |
+
"best_metric": null,
|
| 4 |
+
"best_model_checkpoint": null,
|
| 5 |
+
"epoch": 3.0,
|
| 6 |
+
"eval_steps": 500,
|
| 7 |
+
"global_step": 1167,
|
| 8 |
+
"is_hyper_param_search": false,
|
| 9 |
+
"is_local_process_zero": true,
|
| 10 |
+
"is_world_process_zero": true,
|
| 11 |
+
"log_history": [
|
| 12 |
+
{
|
| 13 |
+
"entropy": 1.6110110306739807,
|
| 14 |
+
"epoch": 0.1287001287001287,
|
| 15 |
+
"grad_norm": 0.8611119389533997,
|
| 16 |
+
"learning_rate": 3.4130257962133866e-05,
|
| 17 |
+
"loss": 1.53956298828125,
|
| 18 |
+
"mean_token_accuracy": 0.6652342769503593,
|
| 19 |
+
"num_tokens": 73407.0,
|
| 20 |
+
"step": 50
|
| 21 |
+
},
|
| 22 |
+
{
|
| 23 |
+
"entropy": 0.8662820833921433,
|
| 24 |
+
"epoch": 0.2574002574002574,
|
| 25 |
+
"grad_norm": 0.7405035495758057,
|
| 26 |
+
"learning_rate": 6.895705180104598e-05,
|
| 27 |
+
"loss": 0.7929539489746094,
|
| 28 |
+
"mean_token_accuracy": 0.7786347842216492,
|
| 29 |
+
"num_tokens": 143994.0,
|
| 30 |
+
"step": 100
|
| 31 |
+
},
|
| 32 |
+
{
|
| 33 |
+
"entropy": 0.7912819278240204,
|
| 34 |
+
"epoch": 0.3861003861003861,
|
| 35 |
+
"grad_norm": 0.6105485558509827,
|
| 36 |
+
"learning_rate": 0.00010378384563995809,
|
| 37 |
+
"loss": 0.7236511993408203,
|
| 38 |
+
"mean_token_accuracy": 0.7925467795133591,
|
| 39 |
+
"num_tokens": 216171.0,
|
| 40 |
+
"step": 150
|
| 41 |
+
},
|
| 42 |
+
{
|
| 43 |
+
"entropy": 0.7777709531784057,
|
| 44 |
+
"epoch": 0.5148005148005148,
|
| 45 |
+
"grad_norm": 0.5781793594360352,
|
| 46 |
+
"learning_rate": 0.0001386106394788702,
|
| 47 |
+
"loss": 0.7024919891357422,
|
| 48 |
+
"mean_token_accuracy": 0.7969876372814179,
|
| 49 |
+
"num_tokens": 284702.0,
|
| 50 |
+
"step": 200
|
| 51 |
+
},
|
| 52 |
+
{
|
| 53 |
+
"entropy": 0.7632609683275223,
|
| 54 |
+
"epoch": 0.6435006435006435,
|
| 55 |
+
"grad_norm": 0.6553444862365723,
|
| 56 |
+
"learning_rate": 0.00017343743331778232,
|
| 57 |
+
"loss": 0.7000718688964844,
|
| 58 |
+
"mean_token_accuracy": 0.7994742071628571,
|
| 59 |
+
"num_tokens": 356393.0,
|
| 60 |
+
"step": 250
|
| 61 |
+
},
|
| 62 |
+
{
|
| 63 |
+
"entropy": 0.754165632724762,
|
| 64 |
+
"epoch": 0.7722007722007722,
|
| 65 |
+
"grad_norm": 0.634222149848938,
|
| 66 |
+
"learning_rate": 0.00020826422715669444,
|
| 67 |
+
"loss": 0.6909049987792969,
|
| 68 |
+
"mean_token_accuracy": 0.8016809666156769,
|
| 69 |
+
"num_tokens": 426916.0,
|
| 70 |
+
"step": 300
|
| 71 |
+
},
|
| 72 |
+
{
|
| 73 |
+
"entropy": 0.7530711203813553,
|
| 74 |
+
"epoch": 0.9009009009009009,
|
| 75 |
+
"grad_norm": 0.7020156383514404,
|
| 76 |
+
"learning_rate": 0.00024309102099560653,
|
| 77 |
+
"loss": 0.69554931640625,
|
| 78 |
+
"mean_token_accuracy": 0.8002807641029358,
|
| 79 |
+
"num_tokens": 499603.0,
|
| 80 |
+
"step": 350
|
| 81 |
+
},
|
| 82 |
+
{
|
| 83 |
+
"epoch": 1.0,
|
| 84 |
+
"eval_entropy": 0.5728044164242204,
|
| 85 |
+
"eval_loss": 0.6785851716995239,
|
| 86 |
+
"eval_mean_token_accuracy": 0.798433246993527,
|
| 87 |
+
"eval_num_tokens": 555493.0,
|
| 88 |
+
"eval_runtime": 162.4129,
|
| 89 |
+
"eval_samples_per_second": 9.513,
|
| 90 |
+
"eval_steps_per_second": 1.194,
|
| 91 |
+
"step": 389
|
| 92 |
+
},
|
| 93 |
+
{
|
| 94 |
+
"entropy": 0.7489313946829902,
|
| 95 |
+
"epoch": 1.0283140283140284,
|
| 96 |
+
"grad_norm": 0.7532815933227539,
|
| 97 |
+
"learning_rate": 0.0002709470016827303,
|
| 98 |
+
"loss": 0.6856581878662109,
|
| 99 |
+
"mean_token_accuracy": 0.8023027079273956,
|
| 100 |
+
"num_tokens": 570813.0,
|
| 101 |
+
"step": 400
|
| 102 |
+
},
|
| 103 |
+
{
|
| 104 |
+
"entropy": 0.7238409864902496,
|
| 105 |
+
"epoch": 1.157014157014157,
|
| 106 |
+
"grad_norm": 1.2016215324401855,
|
| 107 |
+
"learning_rate": 0.0002707561443541359,
|
| 108 |
+
"loss": 0.6699818420410156,
|
| 109 |
+
"mean_token_accuracy": 0.8070302194356919,
|
| 110 |
+
"num_tokens": 642956.0,
|
| 111 |
+
"step": 450
|
| 112 |
+
},
|
| 113 |
+
{
|
| 114 |
+
"entropy": 0.7388938587903976,
|
| 115 |
+
"epoch": 1.2857142857142856,
|
| 116 |
+
"grad_norm": 0.7279272079467773,
|
| 117 |
+
"learning_rate": 0.0002702930068622498,
|
| 118 |
+
"loss": 0.6728517150878907,
|
| 119 |
+
"mean_token_accuracy": 0.8049580943584442,
|
| 120 |
+
"num_tokens": 714498.0,
|
| 121 |
+
"step": 500
|
| 122 |
+
},
|
| 123 |
+
{
|
| 124 |
+
"entropy": 0.7284485149383545,
|
| 125 |
+
"epoch": 1.4144144144144144,
|
| 126 |
+
"grad_norm": 0.8563987016677856,
|
| 127 |
+
"learning_rate": 0.0002695585213716931,
|
| 128 |
+
"loss": 0.6657986450195312,
|
| 129 |
+
"mean_token_accuracy": 0.8085588800907135,
|
| 130 |
+
"num_tokens": 785595.0,
|
| 131 |
+
"step": 550
|
| 132 |
+
},
|
| 133 |
+
{
|
| 134 |
+
"entropy": 0.6960959500074386,
|
| 135 |
+
"epoch": 1.5431145431145432,
|
| 136 |
+
"grad_norm": 0.5145474672317505,
|
| 137 |
+
"learning_rate": 0.0002685541661937683,
|
| 138 |
+
"loss": 0.6358638763427734,
|
| 139 |
+
"mean_token_accuracy": 0.8131621342897415,
|
| 140 |
+
"num_tokens": 857551.0,
|
| 141 |
+
"step": 600
|
| 142 |
+
},
|
| 143 |
+
{
|
| 144 |
+
"entropy": 0.7220414417982102,
|
| 145 |
+
"epoch": 1.6718146718146718,
|
| 146 |
+
"grad_norm": 0.9381059408187866,
|
| 147 |
+
"learning_rate": 0.00026728196281103746,
|
| 148 |
+
"loss": 0.6531407928466797,
|
| 149 |
+
"mean_token_accuracy": 0.811619822382927,
|
| 150 |
+
"num_tokens": 926864.0,
|
| 151 |
+
"step": 650
|
| 152 |
+
},
|
| 153 |
+
{
|
| 154 |
+
"entropy": 0.6955459499359131,
|
| 155 |
+
"epoch": 1.8005148005148004,
|
| 156 |
+
"grad_norm": 0.6115108728408813,
|
| 157 |
+
"learning_rate": 0.0002657444718086503,
|
| 158 |
+
"loss": 0.6269588088989257,
|
| 159 |
+
"mean_token_accuracy": 0.8155373805761337,
|
| 160 |
+
"num_tokens": 998824.0,
|
| 161 |
+
"step": 700
|
| 162 |
+
},
|
| 163 |
+
{
|
| 164 |
+
"entropy": 0.6929812705516816,
|
| 165 |
+
"epoch": 1.9292149292149292,
|
| 166 |
+
"grad_norm": 0.749489426612854,
|
| 167 |
+
"learning_rate": 0.0002639447877206115,
|
| 168 |
+
"loss": 0.629054069519043,
|
| 169 |
+
"mean_token_accuracy": 0.8183021235466004,
|
| 170 |
+
"num_tokens": 1069332.0,
|
| 171 |
+
"step": 750
|
| 172 |
+
},
|
| 173 |
+
{
|
| 174 |
+
"epoch": 2.0,
|
| 175 |
+
"eval_entropy": 0.5833694102223387,
|
| 176 |
+
"eval_loss": 0.6328718662261963,
|
| 177 |
+
"eval_mean_token_accuracy": 0.8159044071571114,
|
| 178 |
+
"eval_num_tokens": 1110986.0,
|
| 179 |
+
"eval_runtime": 161.885,
|
| 180 |
+
"eval_samples_per_second": 9.544,
|
| 181 |
+
"eval_steps_per_second": 1.198,
|
| 182 |
+
"step": 778
|
| 183 |
+
},
|
| 184 |
+
{
|
| 185 |
+
"entropy": 0.647241060480927,
|
| 186 |
+
"epoch": 2.056628056628057,
|
| 187 |
+
"grad_norm": 0.7070767879486084,
|
| 188 |
+
"learning_rate": 0.00026188653280135975,
|
| 189 |
+
"loss": 0.5823922348022461,
|
| 190 |
+
"mean_token_accuracy": 0.8265365301960647,
|
| 191 |
+
"num_tokens": 1141195.0,
|
| 192 |
+
"step": 800
|
| 193 |
+
},
|
| 194 |
+
{
|
| 195 |
+
"entropy": 0.5995378407835961,
|
| 196 |
+
"epoch": 2.1853281853281854,
|
| 197 |
+
"grad_norm": 0.8090486526489258,
|
| 198 |
+
"learning_rate": 0.0002595738497351955,
|
| 199 |
+
"loss": 0.5325597763061524,
|
| 200 |
+
"mean_token_accuracy": 0.8369336777925491,
|
| 201 |
+
"num_tokens": 1210708.0,
|
| 202 |
+
"step": 850
|
| 203 |
+
},
|
| 204 |
+
{
|
| 205 |
+
"entropy": 0.6043145382404327,
|
| 206 |
+
"epoch": 2.314028314028314,
|
| 207 |
+
"grad_norm": 0.8279913067817688,
|
| 208 |
+
"learning_rate": 0.00025701139329823054,
|
| 209 |
+
"loss": 0.5414446258544922,
|
| 210 |
+
"mean_token_accuracy": 0.8361396533250809,
|
| 211 |
+
"num_tokens": 1283441.0,
|
| 212 |
+
"step": 900
|
| 213 |
+
},
|
| 214 |
+
{
|
| 215 |
+
"entropy": 0.5953224584460258,
|
| 216 |
+
"epoch": 2.4427284427284426,
|
| 217 |
+
"grad_norm": 0.6075023412704468,
|
| 218 |
+
"learning_rate": 0.00025420432098964183,
|
| 219 |
+
"loss": 0.536654167175293,
|
| 220 |
+
"mean_token_accuracy": 0.8340093129873276,
|
| 221 |
+
"num_tokens": 1356418.0,
|
| 222 |
+
"step": 950
|
| 223 |
+
},
|
| 224 |
+
{
|
| 225 |
+
"entropy": 0.5998479858040809,
|
| 226 |
+
"epoch": 2.571428571428571,
|
| 227 |
+
"grad_norm": 1.0311471223831177,
|
| 228 |
+
"learning_rate": 0.0002511582826510862,
|
| 229 |
+
"loss": 0.5372924423217773,
|
| 230 |
+
"mean_token_accuracy": 0.8366045409440994,
|
| 231 |
+
"num_tokens": 1427797.0,
|
| 232 |
+
"step": 1000
|
| 233 |
+
},
|
| 234 |
+
{
|
| 235 |
+
"entropy": 0.5980158120393753,
|
| 236 |
+
"epoch": 2.7001287001287,
|
| 237 |
+
"grad_norm": 0.5971426367759705,
|
| 238 |
+
"learning_rate": 0.0002478794090951689,
|
| 239 |
+
"loss": 0.5392885208129883,
|
| 240 |
+
"mean_token_accuracy": 0.8347727072238922,
|
| 241 |
+
"num_tokens": 1498082.0,
|
| 242 |
+
"step": 1050
|
| 243 |
+
},
|
| 244 |
+
{
|
| 245 |
+
"entropy": 0.5990680930018425,
|
| 246 |
+
"epoch": 2.828828828828829,
|
| 247 |
+
"grad_norm": 0.5662627220153809,
|
| 248 |
+
"learning_rate": 0.0002443742997658538,
|
| 249 |
+
"loss": 0.5360498428344727,
|
| 250 |
+
"mean_token_accuracy": 0.8371847170591354,
|
| 251 |
+
"num_tokens": 1568798.0,
|
| 252 |
+
"step": 1100
|
| 253 |
+
},
|
| 254 |
+
{
|
| 255 |
+
"entropy": 0.5957039377093315,
|
| 256 |
+
"epoch": 2.9575289575289574,
|
| 257 |
+
"grad_norm": 0.5043798685073853,
|
| 258 |
+
"learning_rate": 0.00024065000945565205,
|
| 259 |
+
"loss": 0.5342231369018555,
|
| 260 |
+
"mean_token_accuracy": 0.8380735236406326,
|
| 261 |
+
"num_tokens": 1643449.0,
|
| 262 |
+
"step": 1150
|
| 263 |
+
},
|
| 264 |
+
{
|
| 265 |
+
"epoch": 3.0,
|
| 266 |
+
"eval_entropy": 0.5346978819861854,
|
| 267 |
+
"eval_loss": 0.6255015134811401,
|
| 268 |
+
"eval_mean_token_accuracy": 0.8202014476368108,
|
| 269 |
+
"eval_num_tokens": 1666479.0,
|
| 270 |
+
"eval_runtime": 161.6098,
|
| 271 |
+
"eval_samples_per_second": 9.56,
|
| 272 |
+
"eval_steps_per_second": 1.2,
|
| 273 |
+
"step": 1167
|
| 274 |
+
}
|
| 275 |
+
],
|
| 276 |
+
"logging_steps": 50,
|
| 277 |
+
"max_steps": 3890,
|
| 278 |
+
"num_input_tokens_seen": 0,
|
| 279 |
+
"num_train_epochs": 10,
|
| 280 |
+
"save_steps": 500,
|
| 281 |
+
"stateful_callbacks": {
|
| 282 |
+
"TrainerControl": {
|
| 283 |
+
"args": {
|
| 284 |
+
"should_epoch_stop": false,
|
| 285 |
+
"should_evaluate": false,
|
| 286 |
+
"should_log": false,
|
| 287 |
+
"should_save": true,
|
| 288 |
+
"should_training_stop": false
|
| 289 |
+
},
|
| 290 |
+
"attributes": {}
|
| 291 |
+
}
|
| 292 |
+
},
|
| 293 |
+
"total_flos": 2.7917443648582656e+17,
|
| 294 |
+
"train_batch_size": 8,
|
| 295 |
+
"trial_name": null,
|
| 296 |
+
"trial_params": null
|
| 297 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/README.md
ADDED
|
@@ -0,0 +1,209 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen3-14B-Base
|
| 3 |
+
library_name: peft
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:Qwen/Qwen3-14B-Base
|
| 7 |
+
- lora
|
| 8 |
+
- sft
|
| 9 |
+
- transformers
|
| 10 |
+
- trl
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Model Card for Model ID
|
| 14 |
+
|
| 15 |
+
<!-- Provide a quick summary of what the model is/does. -->
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
## Model Details
|
| 20 |
+
|
| 21 |
+
### Model Description
|
| 22 |
+
|
| 23 |
+
<!-- Provide a longer summary of what this model is. -->
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
- **Developed by:** [More Information Needed]
|
| 28 |
+
- **Funded by [optional]:** [More Information Needed]
|
| 29 |
+
- **Shared by [optional]:** [More Information Needed]
|
| 30 |
+
- **Model type:** [More Information Needed]
|
| 31 |
+
- **Language(s) (NLP):** [More Information Needed]
|
| 32 |
+
- **License:** [More Information Needed]
|
| 33 |
+
- **Finetuned from model [optional]:** [More Information Needed]
|
| 34 |
+
|
| 35 |
+
### Model Sources [optional]
|
| 36 |
+
|
| 37 |
+
<!-- Provide the basic links for the model. -->
|
| 38 |
+
|
| 39 |
+
- **Repository:** [More Information Needed]
|
| 40 |
+
- **Paper [optional]:** [More Information Needed]
|
| 41 |
+
- **Demo [optional]:** [More Information Needed]
|
| 42 |
+
|
| 43 |
+
## Uses
|
| 44 |
+
|
| 45 |
+
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
| 46 |
+
|
| 47 |
+
### Direct Use
|
| 48 |
+
|
| 49 |
+
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
| 50 |
+
|
| 51 |
+
[More Information Needed]
|
| 52 |
+
|
| 53 |
+
### Downstream Use [optional]
|
| 54 |
+
|
| 55 |
+
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
| 56 |
+
|
| 57 |
+
[More Information Needed]
|
| 58 |
+
|
| 59 |
+
### Out-of-Scope Use
|
| 60 |
+
|
| 61 |
+
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
| 62 |
+
|
| 63 |
+
[More Information Needed]
|
| 64 |
+
|
| 65 |
+
## Bias, Risks, and Limitations
|
| 66 |
+
|
| 67 |
+
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
| 68 |
+
|
| 69 |
+
[More Information Needed]
|
| 70 |
+
|
| 71 |
+
### Recommendations
|
| 72 |
+
|
| 73 |
+
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
| 74 |
+
|
| 75 |
+
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
| 76 |
+
|
| 77 |
+
## How to Get Started with the Model
|
| 78 |
+
|
| 79 |
+
Use the code below to get started with the model.
|
| 80 |
+
|
| 81 |
+
[More Information Needed]
|
| 82 |
+
|
| 83 |
+
## Training Details
|
| 84 |
+
|
| 85 |
+
### Training Data
|
| 86 |
+
|
| 87 |
+
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
| 88 |
+
|
| 89 |
+
[More Information Needed]
|
| 90 |
+
|
| 91 |
+
### Training Procedure
|
| 92 |
+
|
| 93 |
+
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
| 94 |
+
|
| 95 |
+
#### Preprocessing [optional]
|
| 96 |
+
|
| 97 |
+
[More Information Needed]
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
#### Training Hyperparameters
|
| 101 |
+
|
| 102 |
+
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
| 103 |
+
|
| 104 |
+
#### Speeds, Sizes, Times [optional]
|
| 105 |
+
|
| 106 |
+
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
| 107 |
+
|
| 108 |
+
[More Information Needed]
|
| 109 |
+
|
| 110 |
+
## Evaluation
|
| 111 |
+
|
| 112 |
+
<!-- This section describes the evaluation protocols and provides the results. -->
|
| 113 |
+
|
| 114 |
+
### Testing Data, Factors & Metrics
|
| 115 |
+
|
| 116 |
+
#### Testing Data
|
| 117 |
+
|
| 118 |
+
<!-- This should link to a Dataset Card if possible. -->
|
| 119 |
+
|
| 120 |
+
[More Information Needed]
|
| 121 |
+
|
| 122 |
+
#### Factors
|
| 123 |
+
|
| 124 |
+
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
| 125 |
+
|
| 126 |
+
[More Information Needed]
|
| 127 |
+
|
| 128 |
+
#### Metrics
|
| 129 |
+
|
| 130 |
+
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
| 131 |
+
|
| 132 |
+
[More Information Needed]
|
| 133 |
+
|
| 134 |
+
### Results
|
| 135 |
+
|
| 136 |
+
[More Information Needed]
|
| 137 |
+
|
| 138 |
+
#### Summary
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
## Model Examination [optional]
|
| 143 |
+
|
| 144 |
+
<!-- Relevant interpretability work for the model goes here -->
|
| 145 |
+
|
| 146 |
+
[More Information Needed]
|
| 147 |
+
|
| 148 |
+
## Environmental Impact
|
| 149 |
+
|
| 150 |
+
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
| 151 |
+
|
| 152 |
+
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
| 153 |
+
|
| 154 |
+
- **Hardware Type:** [More Information Needed]
|
| 155 |
+
- **Hours used:** [More Information Needed]
|
| 156 |
+
- **Cloud Provider:** [More Information Needed]
|
| 157 |
+
- **Compute Region:** [More Information Needed]
|
| 158 |
+
- **Carbon Emitted:** [More Information Needed]
|
| 159 |
+
|
| 160 |
+
## Technical Specifications [optional]
|
| 161 |
+
|
| 162 |
+
### Model Architecture and Objective
|
| 163 |
+
|
| 164 |
+
[More Information Needed]
|
| 165 |
+
|
| 166 |
+
### Compute Infrastructure
|
| 167 |
+
|
| 168 |
+
[More Information Needed]
|
| 169 |
+
|
| 170 |
+
#### Hardware
|
| 171 |
+
|
| 172 |
+
[More Information Needed]
|
| 173 |
+
|
| 174 |
+
#### Software
|
| 175 |
+
|
| 176 |
+
[More Information Needed]
|
| 177 |
+
|
| 178 |
+
## Citation [optional]
|
| 179 |
+
|
| 180 |
+
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
| 181 |
+
|
| 182 |
+
**BibTeX:**
|
| 183 |
+
|
| 184 |
+
[More Information Needed]
|
| 185 |
+
|
| 186 |
+
**APA:**
|
| 187 |
+
|
| 188 |
+
[More Information Needed]
|
| 189 |
+
|
| 190 |
+
## Glossary [optional]
|
| 191 |
+
|
| 192 |
+
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
| 193 |
+
|
| 194 |
+
[More Information Needed]
|
| 195 |
+
|
| 196 |
+
## More Information [optional]
|
| 197 |
+
|
| 198 |
+
[More Information Needed]
|
| 199 |
+
|
| 200 |
+
## Model Card Authors [optional]
|
| 201 |
+
|
| 202 |
+
[More Information Needed]
|
| 203 |
+
|
| 204 |
+
## Model Card Contact
|
| 205 |
+
|
| 206 |
+
[More Information Needed]
|
| 207 |
+
### Framework versions
|
| 208 |
+
|
| 209 |
+
- PEFT 0.19.1
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/adapter_config.json
ADDED
|
@@ -0,0 +1,48 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "Qwen/Qwen3-14B-Base",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"layer_replication": null,
|
| 16 |
+
"layers_pattern": null,
|
| 17 |
+
"layers_to_transform": null,
|
| 18 |
+
"loftq_config": {},
|
| 19 |
+
"lora_alpha": 256,
|
| 20 |
+
"lora_bias": false,
|
| 21 |
+
"lora_dropout": 0.00237968804112545,
|
| 22 |
+
"lora_ga_config": null,
|
| 23 |
+
"megatron_config": null,
|
| 24 |
+
"megatron_core": "megatron.core",
|
| 25 |
+
"modules_to_save": null,
|
| 26 |
+
"peft_type": "LORA",
|
| 27 |
+
"peft_version": "0.19.1",
|
| 28 |
+
"qalora_group_size": 16,
|
| 29 |
+
"r": 128,
|
| 30 |
+
"rank_pattern": {},
|
| 31 |
+
"revision": null,
|
| 32 |
+
"target_modules": [
|
| 33 |
+
"gate_proj",
|
| 34 |
+
"q_proj",
|
| 35 |
+
"up_proj",
|
| 36 |
+
"o_proj",
|
| 37 |
+
"v_proj",
|
| 38 |
+
"k_proj",
|
| 39 |
+
"down_proj"
|
| 40 |
+
],
|
| 41 |
+
"target_parameters": null,
|
| 42 |
+
"task_type": "CAUSAL_LM",
|
| 43 |
+
"trainable_token_indices": null,
|
| 44 |
+
"use_bdlora": null,
|
| 45 |
+
"use_dora": false,
|
| 46 |
+
"use_qalora": false,
|
| 47 |
+
"use_rslora": false
|
| 48 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/chat_template.jinja
ADDED
|
@@ -0,0 +1,85 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- if tools %}
|
| 2 |
+
{{- '<|im_start|>system\n' }}
|
| 3 |
+
{%- if messages[0].role == 'system' %}
|
| 4 |
+
{{- messages[0].content + '\n\n' }}
|
| 5 |
+
{%- endif %}
|
| 6 |
+
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
| 7 |
+
{%- for tool in tools %}
|
| 8 |
+
{{- "\n" }}
|
| 9 |
+
{{- tool | tojson }}
|
| 10 |
+
{%- endfor %}
|
| 11 |
+
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
| 12 |
+
{%- else %}
|
| 13 |
+
{%- if messages[0].role == 'system' %}
|
| 14 |
+
{{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
|
| 15 |
+
{%- endif %}
|
| 16 |
+
{%- endif %}
|
| 17 |
+
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
| 18 |
+
{%- for message in messages[::-1] %}
|
| 19 |
+
{%- set index = (messages|length - 1) - loop.index0 %}
|
| 20 |
+
{%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
|
| 21 |
+
{%- set ns.multi_step_tool = false %}
|
| 22 |
+
{%- set ns.last_query_index = index %}
|
| 23 |
+
{%- endif %}
|
| 24 |
+
{%- endfor %}
|
| 25 |
+
{%- for message in messages %}
|
| 26 |
+
{%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
|
| 27 |
+
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
|
| 28 |
+
{%- elif message.role == "assistant" %}
|
| 29 |
+
{%- set content = message.content %}
|
| 30 |
+
{%- set reasoning_content = '' %}
|
| 31 |
+
{%- if message.reasoning_content is defined and message.reasoning_content is not none %}
|
| 32 |
+
{%- set reasoning_content = message.reasoning_content %}
|
| 33 |
+
{%- else %}
|
| 34 |
+
{%- if '</think>' in message.content %}
|
| 35 |
+
{%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
|
| 36 |
+
{%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
| 37 |
+
{%- endif %}
|
| 38 |
+
{%- endif %}
|
| 39 |
+
{%- if loop.index0 > ns.last_query_index %}
|
| 40 |
+
{%- if loop.last or (not loop.last and reasoning_content) %}
|
| 41 |
+
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
|
| 42 |
+
{%- else %}
|
| 43 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 44 |
+
{%- endif %}
|
| 45 |
+
{%- else %}
|
| 46 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 47 |
+
{%- endif %}
|
| 48 |
+
{%- if message.tool_calls %}
|
| 49 |
+
{%- for tool_call in message.tool_calls %}
|
| 50 |
+
{%- if (loop.first and content) or (not loop.first) %}
|
| 51 |
+
{{- '\n' }}
|
| 52 |
+
{%- endif %}
|
| 53 |
+
{%- if tool_call.function %}
|
| 54 |
+
{%- set tool_call = tool_call.function %}
|
| 55 |
+
{%- endif %}
|
| 56 |
+
{{- '<tool_call>\n{"name": "' }}
|
| 57 |
+
{{- tool_call.name }}
|
| 58 |
+
{{- '", "arguments": ' }}
|
| 59 |
+
{%- if tool_call.arguments is string %}
|
| 60 |
+
{{- tool_call.arguments }}
|
| 61 |
+
{%- else %}
|
| 62 |
+
{{- tool_call.arguments | tojson }}
|
| 63 |
+
{%- endif %}
|
| 64 |
+
{{- '}\n</tool_call>' }}
|
| 65 |
+
{%- endfor %}
|
| 66 |
+
{%- endif %}
|
| 67 |
+
{{- '<|im_end|>\n' }}
|
| 68 |
+
{%- elif message.role == "tool" %}
|
| 69 |
+
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
|
| 70 |
+
{{- '<|im_start|>user' }}
|
| 71 |
+
{%- endif %}
|
| 72 |
+
{{- '\n<tool_response>\n' }}
|
| 73 |
+
{{- message.content }}
|
| 74 |
+
{{- '\n</tool_response>' }}
|
| 75 |
+
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
| 76 |
+
{{- '<|im_end|>\n' }}
|
| 77 |
+
{%- endif %}
|
| 78 |
+
{%- endif %}
|
| 79 |
+
{%- endfor %}
|
| 80 |
+
{%- if add_generation_prompt %}
|
| 81 |
+
{{- '<|im_start|>assistant\n' }}
|
| 82 |
+
{%- if enable_thinking is defined and enable_thinking is false %}
|
| 83 |
+
{{- '<think>\n\n</think>\n\n' }}
|
| 84 |
+
{%- endif %}
|
| 85 |
+
{%- endif %}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/tokenizer_config.json
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"backend": "tokenizers",
|
| 4 |
+
"bos_token": null,
|
| 5 |
+
"clean_up_tokenization_spaces": false,
|
| 6 |
+
"eos_token": "<|endoftext|>",
|
| 7 |
+
"errors": "replace",
|
| 8 |
+
"extra_special_tokens": [
|
| 9 |
+
"<|im_start|>",
|
| 10 |
+
"<|im_end|>",
|
| 11 |
+
"<|object_ref_start|>",
|
| 12 |
+
"<|object_ref_end|>",
|
| 13 |
+
"<|box_start|>",
|
| 14 |
+
"<|box_end|>",
|
| 15 |
+
"<|quad_start|>",
|
| 16 |
+
"<|quad_end|>",
|
| 17 |
+
"<|vision_start|>",
|
| 18 |
+
"<|vision_end|>",
|
| 19 |
+
"<|vision_pad|>",
|
| 20 |
+
"<|image_pad|>",
|
| 21 |
+
"<|video_pad|>"
|
| 22 |
+
],
|
| 23 |
+
"is_local": false,
|
| 24 |
+
"model_max_length": 131072,
|
| 25 |
+
"pad_token": "<|endoftext|>",
|
| 26 |
+
"split_special_tokens": false,
|
| 27 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 28 |
+
"unk_token": null
|
| 29 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1556/trainer_state.json
ADDED
|
@@ -0,0 +1,388 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"best_global_step": null,
|
| 3 |
+
"best_metric": null,
|
| 4 |
+
"best_model_checkpoint": null,
|
| 5 |
+
"epoch": 4.0,
|
| 6 |
+
"eval_steps": 500,
|
| 7 |
+
"global_step": 1556,
|
| 8 |
+
"is_hyper_param_search": false,
|
| 9 |
+
"is_local_process_zero": true,
|
| 10 |
+
"is_world_process_zero": true,
|
| 11 |
+
"log_history": [
|
| 12 |
+
{
|
| 13 |
+
"entropy": 1.6110110306739807,
|
| 14 |
+
"epoch": 0.1287001287001287,
|
| 15 |
+
"grad_norm": 0.8611119389533997,
|
| 16 |
+
"learning_rate": 3.4130257962133866e-05,
|
| 17 |
+
"loss": 1.53956298828125,
|
| 18 |
+
"mean_token_accuracy": 0.6652342769503593,
|
| 19 |
+
"num_tokens": 73407.0,
|
| 20 |
+
"step": 50
|
| 21 |
+
},
|
| 22 |
+
{
|
| 23 |
+
"entropy": 0.8662820833921433,
|
| 24 |
+
"epoch": 0.2574002574002574,
|
| 25 |
+
"grad_norm": 0.7405035495758057,
|
| 26 |
+
"learning_rate": 6.895705180104598e-05,
|
| 27 |
+
"loss": 0.7929539489746094,
|
| 28 |
+
"mean_token_accuracy": 0.7786347842216492,
|
| 29 |
+
"num_tokens": 143994.0,
|
| 30 |
+
"step": 100
|
| 31 |
+
},
|
| 32 |
+
{
|
| 33 |
+
"entropy": 0.7912819278240204,
|
| 34 |
+
"epoch": 0.3861003861003861,
|
| 35 |
+
"grad_norm": 0.6105485558509827,
|
| 36 |
+
"learning_rate": 0.00010378384563995809,
|
| 37 |
+
"loss": 0.7236511993408203,
|
| 38 |
+
"mean_token_accuracy": 0.7925467795133591,
|
| 39 |
+
"num_tokens": 216171.0,
|
| 40 |
+
"step": 150
|
| 41 |
+
},
|
| 42 |
+
{
|
| 43 |
+
"entropy": 0.7777709531784057,
|
| 44 |
+
"epoch": 0.5148005148005148,
|
| 45 |
+
"grad_norm": 0.5781793594360352,
|
| 46 |
+
"learning_rate": 0.0001386106394788702,
|
| 47 |
+
"loss": 0.7024919891357422,
|
| 48 |
+
"mean_token_accuracy": 0.7969876372814179,
|
| 49 |
+
"num_tokens": 284702.0,
|
| 50 |
+
"step": 200
|
| 51 |
+
},
|
| 52 |
+
{
|
| 53 |
+
"entropy": 0.7632609683275223,
|
| 54 |
+
"epoch": 0.6435006435006435,
|
| 55 |
+
"grad_norm": 0.6553444862365723,
|
| 56 |
+
"learning_rate": 0.00017343743331778232,
|
| 57 |
+
"loss": 0.7000718688964844,
|
| 58 |
+
"mean_token_accuracy": 0.7994742071628571,
|
| 59 |
+
"num_tokens": 356393.0,
|
| 60 |
+
"step": 250
|
| 61 |
+
},
|
| 62 |
+
{
|
| 63 |
+
"entropy": 0.754165632724762,
|
| 64 |
+
"epoch": 0.7722007722007722,
|
| 65 |
+
"grad_norm": 0.634222149848938,
|
| 66 |
+
"learning_rate": 0.00020826422715669444,
|
| 67 |
+
"loss": 0.6909049987792969,
|
| 68 |
+
"mean_token_accuracy": 0.8016809666156769,
|
| 69 |
+
"num_tokens": 426916.0,
|
| 70 |
+
"step": 300
|
| 71 |
+
},
|
| 72 |
+
{
|
| 73 |
+
"entropy": 0.7530711203813553,
|
| 74 |
+
"epoch": 0.9009009009009009,
|
| 75 |
+
"grad_norm": 0.7020156383514404,
|
| 76 |
+
"learning_rate": 0.00024309102099560653,
|
| 77 |
+
"loss": 0.69554931640625,
|
| 78 |
+
"mean_token_accuracy": 0.8002807641029358,
|
| 79 |
+
"num_tokens": 499603.0,
|
| 80 |
+
"step": 350
|
| 81 |
+
},
|
| 82 |
+
{
|
| 83 |
+
"epoch": 1.0,
|
| 84 |
+
"eval_entropy": 0.5728044164242204,
|
| 85 |
+
"eval_loss": 0.6785851716995239,
|
| 86 |
+
"eval_mean_token_accuracy": 0.798433246993527,
|
| 87 |
+
"eval_num_tokens": 555493.0,
|
| 88 |
+
"eval_runtime": 162.4129,
|
| 89 |
+
"eval_samples_per_second": 9.513,
|
| 90 |
+
"eval_steps_per_second": 1.194,
|
| 91 |
+
"step": 389
|
| 92 |
+
},
|
| 93 |
+
{
|
| 94 |
+
"entropy": 0.7489313946829902,
|
| 95 |
+
"epoch": 1.0283140283140284,
|
| 96 |
+
"grad_norm": 0.7532815933227539,
|
| 97 |
+
"learning_rate": 0.0002709470016827303,
|
| 98 |
+
"loss": 0.6856581878662109,
|
| 99 |
+
"mean_token_accuracy": 0.8023027079273956,
|
| 100 |
+
"num_tokens": 570813.0,
|
| 101 |
+
"step": 400
|
| 102 |
+
},
|
| 103 |
+
{
|
| 104 |
+
"entropy": 0.7238409864902496,
|
| 105 |
+
"epoch": 1.157014157014157,
|
| 106 |
+
"grad_norm": 1.2016215324401855,
|
| 107 |
+
"learning_rate": 0.0002707561443541359,
|
| 108 |
+
"loss": 0.6699818420410156,
|
| 109 |
+
"mean_token_accuracy": 0.8070302194356919,
|
| 110 |
+
"num_tokens": 642956.0,
|
| 111 |
+
"step": 450
|
| 112 |
+
},
|
| 113 |
+
{
|
| 114 |
+
"entropy": 0.7388938587903976,
|
| 115 |
+
"epoch": 1.2857142857142856,
|
| 116 |
+
"grad_norm": 0.7279272079467773,
|
| 117 |
+
"learning_rate": 0.0002702930068622498,
|
| 118 |
+
"loss": 0.6728517150878907,
|
| 119 |
+
"mean_token_accuracy": 0.8049580943584442,
|
| 120 |
+
"num_tokens": 714498.0,
|
| 121 |
+
"step": 500
|
| 122 |
+
},
|
| 123 |
+
{
|
| 124 |
+
"entropy": 0.7284485149383545,
|
| 125 |
+
"epoch": 1.4144144144144144,
|
| 126 |
+
"grad_norm": 0.8563987016677856,
|
| 127 |
+
"learning_rate": 0.0002695585213716931,
|
| 128 |
+
"loss": 0.6657986450195312,
|
| 129 |
+
"mean_token_accuracy": 0.8085588800907135,
|
| 130 |
+
"num_tokens": 785595.0,
|
| 131 |
+
"step": 550
|
| 132 |
+
},
|
| 133 |
+
{
|
| 134 |
+
"entropy": 0.6960959500074386,
|
| 135 |
+
"epoch": 1.5431145431145432,
|
| 136 |
+
"grad_norm": 0.5145474672317505,
|
| 137 |
+
"learning_rate": 0.0002685541661937683,
|
| 138 |
+
"loss": 0.6358638763427734,
|
| 139 |
+
"mean_token_accuracy": 0.8131621342897415,
|
| 140 |
+
"num_tokens": 857551.0,
|
| 141 |
+
"step": 600
|
| 142 |
+
},
|
| 143 |
+
{
|
| 144 |
+
"entropy": 0.7220414417982102,
|
| 145 |
+
"epoch": 1.6718146718146718,
|
| 146 |
+
"grad_norm": 0.9381059408187866,
|
| 147 |
+
"learning_rate": 0.00026728196281103746,
|
| 148 |
+
"loss": 0.6531407928466797,
|
| 149 |
+
"mean_token_accuracy": 0.811619822382927,
|
| 150 |
+
"num_tokens": 926864.0,
|
| 151 |
+
"step": 650
|
| 152 |
+
},
|
| 153 |
+
{
|
| 154 |
+
"entropy": 0.6955459499359131,
|
| 155 |
+
"epoch": 1.8005148005148004,
|
| 156 |
+
"grad_norm": 0.6115108728408813,
|
| 157 |
+
"learning_rate": 0.0002657444718086503,
|
| 158 |
+
"loss": 0.6269588088989257,
|
| 159 |
+
"mean_token_accuracy": 0.8155373805761337,
|
| 160 |
+
"num_tokens": 998824.0,
|
| 161 |
+
"step": 700
|
| 162 |
+
},
|
| 163 |
+
{
|
| 164 |
+
"entropy": 0.6929812705516816,
|
| 165 |
+
"epoch": 1.9292149292149292,
|
| 166 |
+
"grad_norm": 0.749489426612854,
|
| 167 |
+
"learning_rate": 0.0002639447877206115,
|
| 168 |
+
"loss": 0.629054069519043,
|
| 169 |
+
"mean_token_accuracy": 0.8183021235466004,
|
| 170 |
+
"num_tokens": 1069332.0,
|
| 171 |
+
"step": 750
|
| 172 |
+
},
|
| 173 |
+
{
|
| 174 |
+
"epoch": 2.0,
|
| 175 |
+
"eval_entropy": 0.5833694102223387,
|
| 176 |
+
"eval_loss": 0.6328718662261963,
|
| 177 |
+
"eval_mean_token_accuracy": 0.8159044071571114,
|
| 178 |
+
"eval_num_tokens": 1110986.0,
|
| 179 |
+
"eval_runtime": 161.885,
|
| 180 |
+
"eval_samples_per_second": 9.544,
|
| 181 |
+
"eval_steps_per_second": 1.198,
|
| 182 |
+
"step": 778
|
| 183 |
+
},
|
| 184 |
+
{
|
| 185 |
+
"entropy": 0.647241060480927,
|
| 186 |
+
"epoch": 2.056628056628057,
|
| 187 |
+
"grad_norm": 0.7070767879486084,
|
| 188 |
+
"learning_rate": 0.00026188653280135975,
|
| 189 |
+
"loss": 0.5823922348022461,
|
| 190 |
+
"mean_token_accuracy": 0.8265365301960647,
|
| 191 |
+
"num_tokens": 1141195.0,
|
| 192 |
+
"step": 800
|
| 193 |
+
},
|
| 194 |
+
{
|
| 195 |
+
"entropy": 0.5995378407835961,
|
| 196 |
+
"epoch": 2.1853281853281854,
|
| 197 |
+
"grad_norm": 0.8090486526489258,
|
| 198 |
+
"learning_rate": 0.0002595738497351955,
|
| 199 |
+
"loss": 0.5325597763061524,
|
| 200 |
+
"mean_token_accuracy": 0.8369336777925491,
|
| 201 |
+
"num_tokens": 1210708.0,
|
| 202 |
+
"step": 850
|
| 203 |
+
},
|
| 204 |
+
{
|
| 205 |
+
"entropy": 0.6043145382404327,
|
| 206 |
+
"epoch": 2.314028314028314,
|
| 207 |
+
"grad_norm": 0.8279913067817688,
|
| 208 |
+
"learning_rate": 0.00025701139329823054,
|
| 209 |
+
"loss": 0.5414446258544922,
|
| 210 |
+
"mean_token_accuracy": 0.8361396533250809,
|
| 211 |
+
"num_tokens": 1283441.0,
|
| 212 |
+
"step": 900
|
| 213 |
+
},
|
| 214 |
+
{
|
| 215 |
+
"entropy": 0.5953224584460258,
|
| 216 |
+
"epoch": 2.4427284427284426,
|
| 217 |
+
"grad_norm": 0.6075023412704468,
|
| 218 |
+
"learning_rate": 0.00025420432098964183,
|
| 219 |
+
"loss": 0.536654167175293,
|
| 220 |
+
"mean_token_accuracy": 0.8340093129873276,
|
| 221 |
+
"num_tokens": 1356418.0,
|
| 222 |
+
"step": 950
|
| 223 |
+
},
|
| 224 |
+
{
|
| 225 |
+
"entropy": 0.5998479858040809,
|
| 226 |
+
"epoch": 2.571428571428571,
|
| 227 |
+
"grad_norm": 1.0311471223831177,
|
| 228 |
+
"learning_rate": 0.0002511582826510862,
|
| 229 |
+
"loss": 0.5372924423217773,
|
| 230 |
+
"mean_token_accuracy": 0.8366045409440994,
|
| 231 |
+
"num_tokens": 1427797.0,
|
| 232 |
+
"step": 1000
|
| 233 |
+
},
|
| 234 |
+
{
|
| 235 |
+
"entropy": 0.5980158120393753,
|
| 236 |
+
"epoch": 2.7001287001287,
|
| 237 |
+
"grad_norm": 0.5971426367759705,
|
| 238 |
+
"learning_rate": 0.0002478794090951689,
|
| 239 |
+
"loss": 0.5392885208129883,
|
| 240 |
+
"mean_token_accuracy": 0.8347727072238922,
|
| 241 |
+
"num_tokens": 1498082.0,
|
| 242 |
+
"step": 1050
|
| 243 |
+
},
|
| 244 |
+
{
|
| 245 |
+
"entropy": 0.5990680930018425,
|
| 246 |
+
"epoch": 2.828828828828829,
|
| 247 |
+
"grad_norm": 0.5662627220153809,
|
| 248 |
+
"learning_rate": 0.0002443742997658538,
|
| 249 |
+
"loss": 0.5360498428344727,
|
| 250 |
+
"mean_token_accuracy": 0.8371847170591354,
|
| 251 |
+
"num_tokens": 1568798.0,
|
| 252 |
+
"step": 1100
|
| 253 |
+
},
|
| 254 |
+
{
|
| 255 |
+
"entropy": 0.5957039377093315,
|
| 256 |
+
"epoch": 2.9575289575289574,
|
| 257 |
+
"grad_norm": 0.5043798685073853,
|
| 258 |
+
"learning_rate": 0.00024065000945565205,
|
| 259 |
+
"loss": 0.5342231369018555,
|
| 260 |
+
"mean_token_accuracy": 0.8380735236406326,
|
| 261 |
+
"num_tokens": 1643449.0,
|
| 262 |
+
"step": 1150
|
| 263 |
+
},
|
| 264 |
+
{
|
| 265 |
+
"epoch": 3.0,
|
| 266 |
+
"eval_entropy": 0.5346978819861854,
|
| 267 |
+
"eval_loss": 0.6255015134811401,
|
| 268 |
+
"eval_mean_token_accuracy": 0.8202014476368108,
|
| 269 |
+
"eval_num_tokens": 1666479.0,
|
| 270 |
+
"eval_runtime": 161.6098,
|
| 271 |
+
"eval_samples_per_second": 9.56,
|
| 272 |
+
"eval_steps_per_second": 1.2,
|
| 273 |
+
"step": 1167
|
| 274 |
+
},
|
| 275 |
+
{
|
| 276 |
+
"entropy": 0.5212209137401196,
|
| 277 |
+
"epoch": 3.0849420849420848,
|
| 278 |
+
"grad_norm": 0.7095440626144409,
|
| 279 |
+
"learning_rate": 0.00023671403410632178,
|
| 280 |
+
"loss": 0.45311901092529294,
|
| 281 |
+
"mean_token_accuracy": 0.856362871449403,
|
| 282 |
+
"num_tokens": 1713536.0,
|
| 283 |
+
"step": 1200
|
| 284 |
+
},
|
| 285 |
+
{
|
| 286 |
+
"entropy": 0.4701593083143234,
|
| 287 |
+
"epoch": 3.213642213642214,
|
| 288 |
+
"grad_norm": 0.6408083438873291,
|
| 289 |
+
"learning_rate": 0.0002325742957216607,
|
| 290 |
+
"loss": 0.39916397094726563,
|
| 291 |
+
"mean_token_accuracy": 0.8698061722517013,
|
| 292 |
+
"num_tokens": 1785609.0,
|
| 293 |
+
"step": 1250
|
| 294 |
+
},
|
| 295 |
+
{
|
| 296 |
+
"entropy": 0.4766591975092888,
|
| 297 |
+
"epoch": 3.3423423423423424,
|
| 298 |
+
"grad_norm": 0.6415093541145325,
|
| 299 |
+
"learning_rate": 0.0002282391264227552,
|
| 300 |
+
"loss": 0.4116698455810547,
|
| 301 |
+
"mean_token_accuracy": 0.8679435575008392,
|
| 302 |
+
"num_tokens": 1858651.0,
|
| 303 |
+
"step": 1300
|
| 304 |
+
},
|
| 305 |
+
{
|
| 306 |
+
"entropy": 0.4937947469949722,
|
| 307 |
+
"epoch": 3.471042471042471,
|
| 308 |
+
"grad_norm": 0.6549825072288513,
|
| 309 |
+
"learning_rate": 0.00022371725167778054,
|
| 310 |
+
"loss": 0.4296376037597656,
|
| 311 |
+
"mean_token_accuracy": 0.8609357953071595,
|
| 312 |
+
"num_tokens": 1928692.0,
|
| 313 |
+
"step": 1350
|
| 314 |
+
},
|
| 315 |
+
{
|
| 316 |
+
"entropy": 0.4920153194665909,
|
| 317 |
+
"epoch": 3.5997425997425996,
|
| 318 |
+
"grad_norm": 0.6452126502990723,
|
| 319 |
+
"learning_rate": 0.00021901777274010406,
|
| 320 |
+
"loss": 0.4307489013671875,
|
| 321 |
+
"mean_token_accuracy": 0.8606827831268311,
|
| 322 |
+
"num_tokens": 1998668.0,
|
| 323 |
+
"step": 1400
|
| 324 |
+
},
|
| 325 |
+
{
|
| 326 |
+
"entropy": 0.490042342543602,
|
| 327 |
+
"epoch": 3.7284427284427286,
|
| 328 |
+
"grad_norm": 0.5727734565734863,
|
| 329 |
+
"learning_rate": 0.0002141501483300395,
|
| 330 |
+
"loss": 0.4295254135131836,
|
| 331 |
+
"mean_token_accuracy": 0.8616545403003693,
|
| 332 |
+
"num_tokens": 2072809.0,
|
| 333 |
+
"step": 1450
|
| 334 |
+
},
|
| 335 |
+
{
|
| 336 |
+
"entropy": 0.49857193052768706,
|
| 337 |
+
"epoch": 3.857142857142857,
|
| 338 |
+
"grad_norm": 0.7732954025268555,
|
| 339 |
+
"learning_rate": 0.00020912417559712133,
|
| 340 |
+
"loss": 0.4289303207397461,
|
| 341 |
+
"mean_token_accuracy": 0.8616443765163422,
|
| 342 |
+
"num_tokens": 2142475.0,
|
| 343 |
+
"step": 1500
|
| 344 |
+
},
|
| 345 |
+
{
|
| 346 |
+
"entropy": 0.4746784272789955,
|
| 347 |
+
"epoch": 3.985842985842986,
|
| 348 |
+
"grad_norm": 0.5791187882423401,
|
| 349 |
+
"learning_rate": 0.00020394997040121726,
|
| 350 |
+
"loss": 0.4180263900756836,
|
| 351 |
+
"mean_token_accuracy": 0.866080379486084,
|
| 352 |
+
"num_tokens": 2214044.0,
|
| 353 |
+
"step": 1550
|
| 354 |
+
},
|
| 355 |
+
{
|
| 356 |
+
"epoch": 4.0,
|
| 357 |
+
"eval_entropy": 0.4793926059585257,
|
| 358 |
+
"eval_loss": 0.627627968788147,
|
| 359 |
+
"eval_mean_token_accuracy": 0.8233870095813397,
|
| 360 |
+
"eval_num_tokens": 2221972.0,
|
| 361 |
+
"eval_runtime": 162.0225,
|
| 362 |
+
"eval_samples_per_second": 9.536,
|
| 363 |
+
"eval_steps_per_second": 1.197,
|
| 364 |
+
"step": 1556
|
| 365 |
+
}
|
| 366 |
+
],
|
| 367 |
+
"logging_steps": 50,
|
| 368 |
+
"max_steps": 3890,
|
| 369 |
+
"num_input_tokens_seen": 0,
|
| 370 |
+
"num_train_epochs": 10,
|
| 371 |
+
"save_steps": 500,
|
| 372 |
+
"stateful_callbacks": {
|
| 373 |
+
"TrainerControl": {
|
| 374 |
+
"args": {
|
| 375 |
+
"should_epoch_stop": false,
|
| 376 |
+
"should_evaluate": false,
|
| 377 |
+
"should_log": false,
|
| 378 |
+
"should_save": true,
|
| 379 |
+
"should_training_stop": false
|
| 380 |
+
},
|
| 381 |
+
"attributes": {}
|
| 382 |
+
}
|
| 383 |
+
},
|
| 384 |
+
"total_flos": 3.723284547240653e+17,
|
| 385 |
+
"train_batch_size": 8,
|
| 386 |
+
"trial_name": null,
|
| 387 |
+
"trial_params": null
|
| 388 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/README.md
ADDED
|
@@ -0,0 +1,209 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen3-14B-Base
|
| 3 |
+
library_name: peft
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:Qwen/Qwen3-14B-Base
|
| 7 |
+
- lora
|
| 8 |
+
- sft
|
| 9 |
+
- transformers
|
| 10 |
+
- trl
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Model Card for Model ID
|
| 14 |
+
|
| 15 |
+
<!-- Provide a quick summary of what the model is/does. -->
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
## Model Details
|
| 20 |
+
|
| 21 |
+
### Model Description
|
| 22 |
+
|
| 23 |
+
<!-- Provide a longer summary of what this model is. -->
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
- **Developed by:** [More Information Needed]
|
| 28 |
+
- **Funded by [optional]:** [More Information Needed]
|
| 29 |
+
- **Shared by [optional]:** [More Information Needed]
|
| 30 |
+
- **Model type:** [More Information Needed]
|
| 31 |
+
- **Language(s) (NLP):** [More Information Needed]
|
| 32 |
+
- **License:** [More Information Needed]
|
| 33 |
+
- **Finetuned from model [optional]:** [More Information Needed]
|
| 34 |
+
|
| 35 |
+
### Model Sources [optional]
|
| 36 |
+
|
| 37 |
+
<!-- Provide the basic links for the model. -->
|
| 38 |
+
|
| 39 |
+
- **Repository:** [More Information Needed]
|
| 40 |
+
- **Paper [optional]:** [More Information Needed]
|
| 41 |
+
- **Demo [optional]:** [More Information Needed]
|
| 42 |
+
|
| 43 |
+
## Uses
|
| 44 |
+
|
| 45 |
+
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
| 46 |
+
|
| 47 |
+
### Direct Use
|
| 48 |
+
|
| 49 |
+
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
| 50 |
+
|
| 51 |
+
[More Information Needed]
|
| 52 |
+
|
| 53 |
+
### Downstream Use [optional]
|
| 54 |
+
|
| 55 |
+
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
| 56 |
+
|
| 57 |
+
[More Information Needed]
|
| 58 |
+
|
| 59 |
+
### Out-of-Scope Use
|
| 60 |
+
|
| 61 |
+
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
| 62 |
+
|
| 63 |
+
[More Information Needed]
|
| 64 |
+
|
| 65 |
+
## Bias, Risks, and Limitations
|
| 66 |
+
|
| 67 |
+
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
| 68 |
+
|
| 69 |
+
[More Information Needed]
|
| 70 |
+
|
| 71 |
+
### Recommendations
|
| 72 |
+
|
| 73 |
+
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
| 74 |
+
|
| 75 |
+
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
| 76 |
+
|
| 77 |
+
## How to Get Started with the Model
|
| 78 |
+
|
| 79 |
+
Use the code below to get started with the model.
|
| 80 |
+
|
| 81 |
+
[More Information Needed]
|
| 82 |
+
|
| 83 |
+
## Training Details
|
| 84 |
+
|
| 85 |
+
### Training Data
|
| 86 |
+
|
| 87 |
+
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
| 88 |
+
|
| 89 |
+
[More Information Needed]
|
| 90 |
+
|
| 91 |
+
### Training Procedure
|
| 92 |
+
|
| 93 |
+
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
| 94 |
+
|
| 95 |
+
#### Preprocessing [optional]
|
| 96 |
+
|
| 97 |
+
[More Information Needed]
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
#### Training Hyperparameters
|
| 101 |
+
|
| 102 |
+
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
| 103 |
+
|
| 104 |
+
#### Speeds, Sizes, Times [optional]
|
| 105 |
+
|
| 106 |
+
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
| 107 |
+
|
| 108 |
+
[More Information Needed]
|
| 109 |
+
|
| 110 |
+
## Evaluation
|
| 111 |
+
|
| 112 |
+
<!-- This section describes the evaluation protocols and provides the results. -->
|
| 113 |
+
|
| 114 |
+
### Testing Data, Factors & Metrics
|
| 115 |
+
|
| 116 |
+
#### Testing Data
|
| 117 |
+
|
| 118 |
+
<!-- This should link to a Dataset Card if possible. -->
|
| 119 |
+
|
| 120 |
+
[More Information Needed]
|
| 121 |
+
|
| 122 |
+
#### Factors
|
| 123 |
+
|
| 124 |
+
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
| 125 |
+
|
| 126 |
+
[More Information Needed]
|
| 127 |
+
|
| 128 |
+
#### Metrics
|
| 129 |
+
|
| 130 |
+
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
| 131 |
+
|
| 132 |
+
[More Information Needed]
|
| 133 |
+
|
| 134 |
+
### Results
|
| 135 |
+
|
| 136 |
+
[More Information Needed]
|
| 137 |
+
|
| 138 |
+
#### Summary
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
## Model Examination [optional]
|
| 143 |
+
|
| 144 |
+
<!-- Relevant interpretability work for the model goes here -->
|
| 145 |
+
|
| 146 |
+
[More Information Needed]
|
| 147 |
+
|
| 148 |
+
## Environmental Impact
|
| 149 |
+
|
| 150 |
+
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
| 151 |
+
|
| 152 |
+
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
| 153 |
+
|
| 154 |
+
- **Hardware Type:** [More Information Needed]
|
| 155 |
+
- **Hours used:** [More Information Needed]
|
| 156 |
+
- **Cloud Provider:** [More Information Needed]
|
| 157 |
+
- **Compute Region:** [More Information Needed]
|
| 158 |
+
- **Carbon Emitted:** [More Information Needed]
|
| 159 |
+
|
| 160 |
+
## Technical Specifications [optional]
|
| 161 |
+
|
| 162 |
+
### Model Architecture and Objective
|
| 163 |
+
|
| 164 |
+
[More Information Needed]
|
| 165 |
+
|
| 166 |
+
### Compute Infrastructure
|
| 167 |
+
|
| 168 |
+
[More Information Needed]
|
| 169 |
+
|
| 170 |
+
#### Hardware
|
| 171 |
+
|
| 172 |
+
[More Information Needed]
|
| 173 |
+
|
| 174 |
+
#### Software
|
| 175 |
+
|
| 176 |
+
[More Information Needed]
|
| 177 |
+
|
| 178 |
+
## Citation [optional]
|
| 179 |
+
|
| 180 |
+
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
| 181 |
+
|
| 182 |
+
**BibTeX:**
|
| 183 |
+
|
| 184 |
+
[More Information Needed]
|
| 185 |
+
|
| 186 |
+
**APA:**
|
| 187 |
+
|
| 188 |
+
[More Information Needed]
|
| 189 |
+
|
| 190 |
+
## Glossary [optional]
|
| 191 |
+
|
| 192 |
+
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
| 193 |
+
|
| 194 |
+
[More Information Needed]
|
| 195 |
+
|
| 196 |
+
## More Information [optional]
|
| 197 |
+
|
| 198 |
+
[More Information Needed]
|
| 199 |
+
|
| 200 |
+
## Model Card Authors [optional]
|
| 201 |
+
|
| 202 |
+
[More Information Needed]
|
| 203 |
+
|
| 204 |
+
## Model Card Contact
|
| 205 |
+
|
| 206 |
+
[More Information Needed]
|
| 207 |
+
### Framework versions
|
| 208 |
+
|
| 209 |
+
- PEFT 0.19.1
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/adapter_config.json
ADDED
|
@@ -0,0 +1,48 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "Qwen/Qwen3-14B-Base",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"layer_replication": null,
|
| 16 |
+
"layers_pattern": null,
|
| 17 |
+
"layers_to_transform": null,
|
| 18 |
+
"loftq_config": {},
|
| 19 |
+
"lora_alpha": 256,
|
| 20 |
+
"lora_bias": false,
|
| 21 |
+
"lora_dropout": 0.00237968804112545,
|
| 22 |
+
"lora_ga_config": null,
|
| 23 |
+
"megatron_config": null,
|
| 24 |
+
"megatron_core": "megatron.core",
|
| 25 |
+
"modules_to_save": null,
|
| 26 |
+
"peft_type": "LORA",
|
| 27 |
+
"peft_version": "0.19.1",
|
| 28 |
+
"qalora_group_size": 16,
|
| 29 |
+
"r": 128,
|
| 30 |
+
"rank_pattern": {},
|
| 31 |
+
"revision": null,
|
| 32 |
+
"target_modules": [
|
| 33 |
+
"gate_proj",
|
| 34 |
+
"q_proj",
|
| 35 |
+
"up_proj",
|
| 36 |
+
"o_proj",
|
| 37 |
+
"v_proj",
|
| 38 |
+
"k_proj",
|
| 39 |
+
"down_proj"
|
| 40 |
+
],
|
| 41 |
+
"target_parameters": null,
|
| 42 |
+
"task_type": "CAUSAL_LM",
|
| 43 |
+
"trainable_token_indices": null,
|
| 44 |
+
"use_bdlora": null,
|
| 45 |
+
"use_dora": false,
|
| 46 |
+
"use_qalora": false,
|
| 47 |
+
"use_rslora": false
|
| 48 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/chat_template.jinja
ADDED
|
@@ -0,0 +1,85 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- if tools %}
|
| 2 |
+
{{- '<|im_start|>system\n' }}
|
| 3 |
+
{%- if messages[0].role == 'system' %}
|
| 4 |
+
{{- messages[0].content + '\n\n' }}
|
| 5 |
+
{%- endif %}
|
| 6 |
+
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
| 7 |
+
{%- for tool in tools %}
|
| 8 |
+
{{- "\n" }}
|
| 9 |
+
{{- tool | tojson }}
|
| 10 |
+
{%- endfor %}
|
| 11 |
+
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
| 12 |
+
{%- else %}
|
| 13 |
+
{%- if messages[0].role == 'system' %}
|
| 14 |
+
{{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
|
| 15 |
+
{%- endif %}
|
| 16 |
+
{%- endif %}
|
| 17 |
+
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
| 18 |
+
{%- for message in messages[::-1] %}
|
| 19 |
+
{%- set index = (messages|length - 1) - loop.index0 %}
|
| 20 |
+
{%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
|
| 21 |
+
{%- set ns.multi_step_tool = false %}
|
| 22 |
+
{%- set ns.last_query_index = index %}
|
| 23 |
+
{%- endif %}
|
| 24 |
+
{%- endfor %}
|
| 25 |
+
{%- for message in messages %}
|
| 26 |
+
{%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
|
| 27 |
+
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
|
| 28 |
+
{%- elif message.role == "assistant" %}
|
| 29 |
+
{%- set content = message.content %}
|
| 30 |
+
{%- set reasoning_content = '' %}
|
| 31 |
+
{%- if message.reasoning_content is defined and message.reasoning_content is not none %}
|
| 32 |
+
{%- set reasoning_content = message.reasoning_content %}
|
| 33 |
+
{%- else %}
|
| 34 |
+
{%- if '</think>' in message.content %}
|
| 35 |
+
{%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
|
| 36 |
+
{%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
| 37 |
+
{%- endif %}
|
| 38 |
+
{%- endif %}
|
| 39 |
+
{%- if loop.index0 > ns.last_query_index %}
|
| 40 |
+
{%- if loop.last or (not loop.last and reasoning_content) %}
|
| 41 |
+
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
|
| 42 |
+
{%- else %}
|
| 43 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 44 |
+
{%- endif %}
|
| 45 |
+
{%- else %}
|
| 46 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 47 |
+
{%- endif %}
|
| 48 |
+
{%- if message.tool_calls %}
|
| 49 |
+
{%- for tool_call in message.tool_calls %}
|
| 50 |
+
{%- if (loop.first and content) or (not loop.first) %}
|
| 51 |
+
{{- '\n' }}
|
| 52 |
+
{%- endif %}
|
| 53 |
+
{%- if tool_call.function %}
|
| 54 |
+
{%- set tool_call = tool_call.function %}
|
| 55 |
+
{%- endif %}
|
| 56 |
+
{{- '<tool_call>\n{"name": "' }}
|
| 57 |
+
{{- tool_call.name }}
|
| 58 |
+
{{- '", "arguments": ' }}
|
| 59 |
+
{%- if tool_call.arguments is string %}
|
| 60 |
+
{{- tool_call.arguments }}
|
| 61 |
+
{%- else %}
|
| 62 |
+
{{- tool_call.arguments | tojson }}
|
| 63 |
+
{%- endif %}
|
| 64 |
+
{{- '}\n</tool_call>' }}
|
| 65 |
+
{%- endfor %}
|
| 66 |
+
{%- endif %}
|
| 67 |
+
{{- '<|im_end|>\n' }}
|
| 68 |
+
{%- elif message.role == "tool" %}
|
| 69 |
+
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
|
| 70 |
+
{{- '<|im_start|>user' }}
|
| 71 |
+
{%- endif %}
|
| 72 |
+
{{- '\n<tool_response>\n' }}
|
| 73 |
+
{{- message.content }}
|
| 74 |
+
{{- '\n</tool_response>' }}
|
| 75 |
+
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
| 76 |
+
{{- '<|im_end|>\n' }}
|
| 77 |
+
{%- endif %}
|
| 78 |
+
{%- endif %}
|
| 79 |
+
{%- endfor %}
|
| 80 |
+
{%- if add_generation_prompt %}
|
| 81 |
+
{{- '<|im_start|>assistant\n' }}
|
| 82 |
+
{%- if enable_thinking is defined and enable_thinking is false %}
|
| 83 |
+
{{- '<think>\n\n</think>\n\n' }}
|
| 84 |
+
{%- endif %}
|
| 85 |
+
{%- endif %}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/tokenizer_config.json
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"backend": "tokenizers",
|
| 4 |
+
"bos_token": null,
|
| 5 |
+
"clean_up_tokenization_spaces": false,
|
| 6 |
+
"eos_token": "<|endoftext|>",
|
| 7 |
+
"errors": "replace",
|
| 8 |
+
"extra_special_tokens": [
|
| 9 |
+
"<|im_start|>",
|
| 10 |
+
"<|im_end|>",
|
| 11 |
+
"<|object_ref_start|>",
|
| 12 |
+
"<|object_ref_end|>",
|
| 13 |
+
"<|box_start|>",
|
| 14 |
+
"<|box_end|>",
|
| 15 |
+
"<|quad_start|>",
|
| 16 |
+
"<|quad_end|>",
|
| 17 |
+
"<|vision_start|>",
|
| 18 |
+
"<|vision_end|>",
|
| 19 |
+
"<|vision_pad|>",
|
| 20 |
+
"<|image_pad|>",
|
| 21 |
+
"<|video_pad|>"
|
| 22 |
+
],
|
| 23 |
+
"is_local": false,
|
| 24 |
+
"model_max_length": 131072,
|
| 25 |
+
"pad_token": "<|endoftext|>",
|
| 26 |
+
"split_special_tokens": false,
|
| 27 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 28 |
+
"unk_token": null
|
| 29 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-1945/trainer_state.json
ADDED
|
@@ -0,0 +1,469 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"best_global_step": null,
|
| 3 |
+
"best_metric": null,
|
| 4 |
+
"best_model_checkpoint": null,
|
| 5 |
+
"epoch": 5.0,
|
| 6 |
+
"eval_steps": 500,
|
| 7 |
+
"global_step": 1945,
|
| 8 |
+
"is_hyper_param_search": false,
|
| 9 |
+
"is_local_process_zero": true,
|
| 10 |
+
"is_world_process_zero": true,
|
| 11 |
+
"log_history": [
|
| 12 |
+
{
|
| 13 |
+
"entropy": 1.6110110306739807,
|
| 14 |
+
"epoch": 0.1287001287001287,
|
| 15 |
+
"grad_norm": 0.8611119389533997,
|
| 16 |
+
"learning_rate": 3.4130257962133866e-05,
|
| 17 |
+
"loss": 1.53956298828125,
|
| 18 |
+
"mean_token_accuracy": 0.6652342769503593,
|
| 19 |
+
"num_tokens": 73407.0,
|
| 20 |
+
"step": 50
|
| 21 |
+
},
|
| 22 |
+
{
|
| 23 |
+
"entropy": 0.8662820833921433,
|
| 24 |
+
"epoch": 0.2574002574002574,
|
| 25 |
+
"grad_norm": 0.7405035495758057,
|
| 26 |
+
"learning_rate": 6.895705180104598e-05,
|
| 27 |
+
"loss": 0.7929539489746094,
|
| 28 |
+
"mean_token_accuracy": 0.7786347842216492,
|
| 29 |
+
"num_tokens": 143994.0,
|
| 30 |
+
"step": 100
|
| 31 |
+
},
|
| 32 |
+
{
|
| 33 |
+
"entropy": 0.7912819278240204,
|
| 34 |
+
"epoch": 0.3861003861003861,
|
| 35 |
+
"grad_norm": 0.6105485558509827,
|
| 36 |
+
"learning_rate": 0.00010378384563995809,
|
| 37 |
+
"loss": 0.7236511993408203,
|
| 38 |
+
"mean_token_accuracy": 0.7925467795133591,
|
| 39 |
+
"num_tokens": 216171.0,
|
| 40 |
+
"step": 150
|
| 41 |
+
},
|
| 42 |
+
{
|
| 43 |
+
"entropy": 0.7777709531784057,
|
| 44 |
+
"epoch": 0.5148005148005148,
|
| 45 |
+
"grad_norm": 0.5781793594360352,
|
| 46 |
+
"learning_rate": 0.0001386106394788702,
|
| 47 |
+
"loss": 0.7024919891357422,
|
| 48 |
+
"mean_token_accuracy": 0.7969876372814179,
|
| 49 |
+
"num_tokens": 284702.0,
|
| 50 |
+
"step": 200
|
| 51 |
+
},
|
| 52 |
+
{
|
| 53 |
+
"entropy": 0.7632609683275223,
|
| 54 |
+
"epoch": 0.6435006435006435,
|
| 55 |
+
"grad_norm": 0.6553444862365723,
|
| 56 |
+
"learning_rate": 0.00017343743331778232,
|
| 57 |
+
"loss": 0.7000718688964844,
|
| 58 |
+
"mean_token_accuracy": 0.7994742071628571,
|
| 59 |
+
"num_tokens": 356393.0,
|
| 60 |
+
"step": 250
|
| 61 |
+
},
|
| 62 |
+
{
|
| 63 |
+
"entropy": 0.754165632724762,
|
| 64 |
+
"epoch": 0.7722007722007722,
|
| 65 |
+
"grad_norm": 0.634222149848938,
|
| 66 |
+
"learning_rate": 0.00020826422715669444,
|
| 67 |
+
"loss": 0.6909049987792969,
|
| 68 |
+
"mean_token_accuracy": 0.8016809666156769,
|
| 69 |
+
"num_tokens": 426916.0,
|
| 70 |
+
"step": 300
|
| 71 |
+
},
|
| 72 |
+
{
|
| 73 |
+
"entropy": 0.7530711203813553,
|
| 74 |
+
"epoch": 0.9009009009009009,
|
| 75 |
+
"grad_norm": 0.7020156383514404,
|
| 76 |
+
"learning_rate": 0.00024309102099560653,
|
| 77 |
+
"loss": 0.69554931640625,
|
| 78 |
+
"mean_token_accuracy": 0.8002807641029358,
|
| 79 |
+
"num_tokens": 499603.0,
|
| 80 |
+
"step": 350
|
| 81 |
+
},
|
| 82 |
+
{
|
| 83 |
+
"epoch": 1.0,
|
| 84 |
+
"eval_entropy": 0.5728044164242204,
|
| 85 |
+
"eval_loss": 0.6785851716995239,
|
| 86 |
+
"eval_mean_token_accuracy": 0.798433246993527,
|
| 87 |
+
"eval_num_tokens": 555493.0,
|
| 88 |
+
"eval_runtime": 162.4129,
|
| 89 |
+
"eval_samples_per_second": 9.513,
|
| 90 |
+
"eval_steps_per_second": 1.194,
|
| 91 |
+
"step": 389
|
| 92 |
+
},
|
| 93 |
+
{
|
| 94 |
+
"entropy": 0.7489313946829902,
|
| 95 |
+
"epoch": 1.0283140283140284,
|
| 96 |
+
"grad_norm": 0.7532815933227539,
|
| 97 |
+
"learning_rate": 0.0002709470016827303,
|
| 98 |
+
"loss": 0.6856581878662109,
|
| 99 |
+
"mean_token_accuracy": 0.8023027079273956,
|
| 100 |
+
"num_tokens": 570813.0,
|
| 101 |
+
"step": 400
|
| 102 |
+
},
|
| 103 |
+
{
|
| 104 |
+
"entropy": 0.7238409864902496,
|
| 105 |
+
"epoch": 1.157014157014157,
|
| 106 |
+
"grad_norm": 1.2016215324401855,
|
| 107 |
+
"learning_rate": 0.0002707561443541359,
|
| 108 |
+
"loss": 0.6699818420410156,
|
| 109 |
+
"mean_token_accuracy": 0.8070302194356919,
|
| 110 |
+
"num_tokens": 642956.0,
|
| 111 |
+
"step": 450
|
| 112 |
+
},
|
| 113 |
+
{
|
| 114 |
+
"entropy": 0.7388938587903976,
|
| 115 |
+
"epoch": 1.2857142857142856,
|
| 116 |
+
"grad_norm": 0.7279272079467773,
|
| 117 |
+
"learning_rate": 0.0002702930068622498,
|
| 118 |
+
"loss": 0.6728517150878907,
|
| 119 |
+
"mean_token_accuracy": 0.8049580943584442,
|
| 120 |
+
"num_tokens": 714498.0,
|
| 121 |
+
"step": 500
|
| 122 |
+
},
|
| 123 |
+
{
|
| 124 |
+
"entropy": 0.7284485149383545,
|
| 125 |
+
"epoch": 1.4144144144144144,
|
| 126 |
+
"grad_norm": 0.8563987016677856,
|
| 127 |
+
"learning_rate": 0.0002695585213716931,
|
| 128 |
+
"loss": 0.6657986450195312,
|
| 129 |
+
"mean_token_accuracy": 0.8085588800907135,
|
| 130 |
+
"num_tokens": 785595.0,
|
| 131 |
+
"step": 550
|
| 132 |
+
},
|
| 133 |
+
{
|
| 134 |
+
"entropy": 0.6960959500074386,
|
| 135 |
+
"epoch": 1.5431145431145432,
|
| 136 |
+
"grad_norm": 0.5145474672317505,
|
| 137 |
+
"learning_rate": 0.0002685541661937683,
|
| 138 |
+
"loss": 0.6358638763427734,
|
| 139 |
+
"mean_token_accuracy": 0.8131621342897415,
|
| 140 |
+
"num_tokens": 857551.0,
|
| 141 |
+
"step": 600
|
| 142 |
+
},
|
| 143 |
+
{
|
| 144 |
+
"entropy": 0.7220414417982102,
|
| 145 |
+
"epoch": 1.6718146718146718,
|
| 146 |
+
"grad_norm": 0.9381059408187866,
|
| 147 |
+
"learning_rate": 0.00026728196281103746,
|
| 148 |
+
"loss": 0.6531407928466797,
|
| 149 |
+
"mean_token_accuracy": 0.811619822382927,
|
| 150 |
+
"num_tokens": 926864.0,
|
| 151 |
+
"step": 650
|
| 152 |
+
},
|
| 153 |
+
{
|
| 154 |
+
"entropy": 0.6955459499359131,
|
| 155 |
+
"epoch": 1.8005148005148004,
|
| 156 |
+
"grad_norm": 0.6115108728408813,
|
| 157 |
+
"learning_rate": 0.0002657444718086503,
|
| 158 |
+
"loss": 0.6269588088989257,
|
| 159 |
+
"mean_token_accuracy": 0.8155373805761337,
|
| 160 |
+
"num_tokens": 998824.0,
|
| 161 |
+
"step": 700
|
| 162 |
+
},
|
| 163 |
+
{
|
| 164 |
+
"entropy": 0.6929812705516816,
|
| 165 |
+
"epoch": 1.9292149292149292,
|
| 166 |
+
"grad_norm": 0.749489426612854,
|
| 167 |
+
"learning_rate": 0.0002639447877206115,
|
| 168 |
+
"loss": 0.629054069519043,
|
| 169 |
+
"mean_token_accuracy": 0.8183021235466004,
|
| 170 |
+
"num_tokens": 1069332.0,
|
| 171 |
+
"step": 750
|
| 172 |
+
},
|
| 173 |
+
{
|
| 174 |
+
"epoch": 2.0,
|
| 175 |
+
"eval_entropy": 0.5833694102223387,
|
| 176 |
+
"eval_loss": 0.6328718662261963,
|
| 177 |
+
"eval_mean_token_accuracy": 0.8159044071571114,
|
| 178 |
+
"eval_num_tokens": 1110986.0,
|
| 179 |
+
"eval_runtime": 161.885,
|
| 180 |
+
"eval_samples_per_second": 9.544,
|
| 181 |
+
"eval_steps_per_second": 1.198,
|
| 182 |
+
"step": 778
|
| 183 |
+
},
|
| 184 |
+
{
|
| 185 |
+
"entropy": 0.647241060480927,
|
| 186 |
+
"epoch": 2.056628056628057,
|
| 187 |
+
"grad_norm": 0.7070767879486084,
|
| 188 |
+
"learning_rate": 0.00026188653280135975,
|
| 189 |
+
"loss": 0.5823922348022461,
|
| 190 |
+
"mean_token_accuracy": 0.8265365301960647,
|
| 191 |
+
"num_tokens": 1141195.0,
|
| 192 |
+
"step": 800
|
| 193 |
+
},
|
| 194 |
+
{
|
| 195 |
+
"entropy": 0.5995378407835961,
|
| 196 |
+
"epoch": 2.1853281853281854,
|
| 197 |
+
"grad_norm": 0.8090486526489258,
|
| 198 |
+
"learning_rate": 0.0002595738497351955,
|
| 199 |
+
"loss": 0.5325597763061524,
|
| 200 |
+
"mean_token_accuracy": 0.8369336777925491,
|
| 201 |
+
"num_tokens": 1210708.0,
|
| 202 |
+
"step": 850
|
| 203 |
+
},
|
| 204 |
+
{
|
| 205 |
+
"entropy": 0.6043145382404327,
|
| 206 |
+
"epoch": 2.314028314028314,
|
| 207 |
+
"grad_norm": 0.8279913067817688,
|
| 208 |
+
"learning_rate": 0.00025701139329823054,
|
| 209 |
+
"loss": 0.5414446258544922,
|
| 210 |
+
"mean_token_accuracy": 0.8361396533250809,
|
| 211 |
+
"num_tokens": 1283441.0,
|
| 212 |
+
"step": 900
|
| 213 |
+
},
|
| 214 |
+
{
|
| 215 |
+
"entropy": 0.5953224584460258,
|
| 216 |
+
"epoch": 2.4427284427284426,
|
| 217 |
+
"grad_norm": 0.6075023412704468,
|
| 218 |
+
"learning_rate": 0.00025420432098964183,
|
| 219 |
+
"loss": 0.536654167175293,
|
| 220 |
+
"mean_token_accuracy": 0.8340093129873276,
|
| 221 |
+
"num_tokens": 1356418.0,
|
| 222 |
+
"step": 950
|
| 223 |
+
},
|
| 224 |
+
{
|
| 225 |
+
"entropy": 0.5998479858040809,
|
| 226 |
+
"epoch": 2.571428571428571,
|
| 227 |
+
"grad_norm": 1.0311471223831177,
|
| 228 |
+
"learning_rate": 0.0002511582826510862,
|
| 229 |
+
"loss": 0.5372924423217773,
|
| 230 |
+
"mean_token_accuracy": 0.8366045409440994,
|
| 231 |
+
"num_tokens": 1427797.0,
|
| 232 |
+
"step": 1000
|
| 233 |
+
},
|
| 234 |
+
{
|
| 235 |
+
"entropy": 0.5980158120393753,
|
| 236 |
+
"epoch": 2.7001287001287,
|
| 237 |
+
"grad_norm": 0.5971426367759705,
|
| 238 |
+
"learning_rate": 0.0002478794090951689,
|
| 239 |
+
"loss": 0.5392885208129883,
|
| 240 |
+
"mean_token_accuracy": 0.8347727072238922,
|
| 241 |
+
"num_tokens": 1498082.0,
|
| 242 |
+
"step": 1050
|
| 243 |
+
},
|
| 244 |
+
{
|
| 245 |
+
"entropy": 0.5990680930018425,
|
| 246 |
+
"epoch": 2.828828828828829,
|
| 247 |
+
"grad_norm": 0.5662627220153809,
|
| 248 |
+
"learning_rate": 0.0002443742997658538,
|
| 249 |
+
"loss": 0.5360498428344727,
|
| 250 |
+
"mean_token_accuracy": 0.8371847170591354,
|
| 251 |
+
"num_tokens": 1568798.0,
|
| 252 |
+
"step": 1100
|
| 253 |
+
},
|
| 254 |
+
{
|
| 255 |
+
"entropy": 0.5957039377093315,
|
| 256 |
+
"epoch": 2.9575289575289574,
|
| 257 |
+
"grad_norm": 0.5043798685073853,
|
| 258 |
+
"learning_rate": 0.00024065000945565205,
|
| 259 |
+
"loss": 0.5342231369018555,
|
| 260 |
+
"mean_token_accuracy": 0.8380735236406326,
|
| 261 |
+
"num_tokens": 1643449.0,
|
| 262 |
+
"step": 1150
|
| 263 |
+
},
|
| 264 |
+
{
|
| 265 |
+
"epoch": 3.0,
|
| 266 |
+
"eval_entropy": 0.5346978819861854,
|
| 267 |
+
"eval_loss": 0.6255015134811401,
|
| 268 |
+
"eval_mean_token_accuracy": 0.8202014476368108,
|
| 269 |
+
"eval_num_tokens": 1666479.0,
|
| 270 |
+
"eval_runtime": 161.6098,
|
| 271 |
+
"eval_samples_per_second": 9.56,
|
| 272 |
+
"eval_steps_per_second": 1.2,
|
| 273 |
+
"step": 1167
|
| 274 |
+
},
|
| 275 |
+
{
|
| 276 |
+
"entropy": 0.5212209137401196,
|
| 277 |
+
"epoch": 3.0849420849420848,
|
| 278 |
+
"grad_norm": 0.7095440626144409,
|
| 279 |
+
"learning_rate": 0.00023671403410632178,
|
| 280 |
+
"loss": 0.45311901092529294,
|
| 281 |
+
"mean_token_accuracy": 0.856362871449403,
|
| 282 |
+
"num_tokens": 1713536.0,
|
| 283 |
+
"step": 1200
|
| 284 |
+
},
|
| 285 |
+
{
|
| 286 |
+
"entropy": 0.4701593083143234,
|
| 287 |
+
"epoch": 3.213642213642214,
|
| 288 |
+
"grad_norm": 0.6408083438873291,
|
| 289 |
+
"learning_rate": 0.0002325742957216607,
|
| 290 |
+
"loss": 0.39916397094726563,
|
| 291 |
+
"mean_token_accuracy": 0.8698061722517013,
|
| 292 |
+
"num_tokens": 1785609.0,
|
| 293 |
+
"step": 1250
|
| 294 |
+
},
|
| 295 |
+
{
|
| 296 |
+
"entropy": 0.4766591975092888,
|
| 297 |
+
"epoch": 3.3423423423423424,
|
| 298 |
+
"grad_norm": 0.6415093541145325,
|
| 299 |
+
"learning_rate": 0.0002282391264227552,
|
| 300 |
+
"loss": 0.4116698455810547,
|
| 301 |
+
"mean_token_accuracy": 0.8679435575008392,
|
| 302 |
+
"num_tokens": 1858651.0,
|
| 303 |
+
"step": 1300
|
| 304 |
+
},
|
| 305 |
+
{
|
| 306 |
+
"entropy": 0.4937947469949722,
|
| 307 |
+
"epoch": 3.471042471042471,
|
| 308 |
+
"grad_norm": 0.6549825072288513,
|
| 309 |
+
"learning_rate": 0.00022371725167778054,
|
| 310 |
+
"loss": 0.4296376037597656,
|
| 311 |
+
"mean_token_accuracy": 0.8609357953071595,
|
| 312 |
+
"num_tokens": 1928692.0,
|
| 313 |
+
"step": 1350
|
| 314 |
+
},
|
| 315 |
+
{
|
| 316 |
+
"entropy": 0.4920153194665909,
|
| 317 |
+
"epoch": 3.5997425997425996,
|
| 318 |
+
"grad_norm": 0.6452126502990723,
|
| 319 |
+
"learning_rate": 0.00021901777274010406,
|
| 320 |
+
"loss": 0.4307489013671875,
|
| 321 |
+
"mean_token_accuracy": 0.8606827831268311,
|
| 322 |
+
"num_tokens": 1998668.0,
|
| 323 |
+
"step": 1400
|
| 324 |
+
},
|
| 325 |
+
{
|
| 326 |
+
"entropy": 0.490042342543602,
|
| 327 |
+
"epoch": 3.7284427284427286,
|
| 328 |
+
"grad_norm": 0.5727734565734863,
|
| 329 |
+
"learning_rate": 0.0002141501483300395,
|
| 330 |
+
"loss": 0.4295254135131836,
|
| 331 |
+
"mean_token_accuracy": 0.8616545403003693,
|
| 332 |
+
"num_tokens": 2072809.0,
|
| 333 |
+
"step": 1450
|
| 334 |
+
},
|
| 335 |
+
{
|
| 336 |
+
"entropy": 0.49857193052768706,
|
| 337 |
+
"epoch": 3.857142857142857,
|
| 338 |
+
"grad_norm": 0.7732954025268555,
|
| 339 |
+
"learning_rate": 0.00020912417559712133,
|
| 340 |
+
"loss": 0.4289303207397461,
|
| 341 |
+
"mean_token_accuracy": 0.8616443765163422,
|
| 342 |
+
"num_tokens": 2142475.0,
|
| 343 |
+
"step": 1500
|
| 344 |
+
},
|
| 345 |
+
{
|
| 346 |
+
"entropy": 0.4746784272789955,
|
| 347 |
+
"epoch": 3.985842985842986,
|
| 348 |
+
"grad_norm": 0.5791187882423401,
|
| 349 |
+
"learning_rate": 0.00020394997040121726,
|
| 350 |
+
"loss": 0.4180263900756836,
|
| 351 |
+
"mean_token_accuracy": 0.866080379486084,
|
| 352 |
+
"num_tokens": 2214044.0,
|
| 353 |
+
"step": 1550
|
| 354 |
+
},
|
| 355 |
+
{
|
| 356 |
+
"epoch": 4.0,
|
| 357 |
+
"eval_entropy": 0.4793926059585257,
|
| 358 |
+
"eval_loss": 0.627627968788147,
|
| 359 |
+
"eval_mean_token_accuracy": 0.8233870095813397,
|
| 360 |
+
"eval_num_tokens": 2221972.0,
|
| 361 |
+
"eval_runtime": 162.0225,
|
| 362 |
+
"eval_samples_per_second": 9.536,
|
| 363 |
+
"eval_steps_per_second": 1.197,
|
| 364 |
+
"step": 1556
|
| 365 |
+
},
|
| 366 |
+
{
|
| 367 |
+
"entropy": 0.38195489000792454,
|
| 368 |
+
"epoch": 4.113256113256114,
|
| 369 |
+
"grad_norm": 0.6002617478370667,
|
| 370 |
+
"learning_rate": 0.0001986379469521669,
|
| 371 |
+
"loss": 0.30819049835205076,
|
| 372 |
+
"mean_token_accuracy": 0.8977848634575353,
|
| 373 |
+
"num_tokens": 2282164.0,
|
| 374 |
+
"step": 1600
|
| 375 |
+
},
|
| 376 |
+
{
|
| 377 |
+
"entropy": 0.3655787402391434,
|
| 378 |
+
"epoch": 4.241956241956242,
|
| 379 |
+
"grad_norm": 0.7100041508674622,
|
| 380 |
+
"learning_rate": 0.00019319879684892634,
|
| 381 |
+
"loss": 0.29959835052490236,
|
| 382 |
+
"mean_token_accuracy": 0.8991208010911942,
|
| 383 |
+
"num_tokens": 2353213.0,
|
| 384 |
+
"step": 1650
|
| 385 |
+
},
|
| 386 |
+
{
|
| 387 |
+
"entropy": 0.3821141055226326,
|
| 388 |
+
"epoch": 4.370656370656371,
|
| 389 |
+
"grad_norm": 0.5848307013511658,
|
| 390 |
+
"learning_rate": 0.00018764346756040715,
|
| 391 |
+
"loss": 0.313802490234375,
|
| 392 |
+
"mean_token_accuracy": 0.895167955160141,
|
| 393 |
+
"num_tokens": 2425068.0,
|
| 394 |
+
"step": 1700
|
| 395 |
+
},
|
| 396 |
+
{
|
| 397 |
+
"entropy": 0.37083797007799146,
|
| 398 |
+
"epoch": 4.499356499356499,
|
| 399 |
+
"grad_norm": 0.6447024941444397,
|
| 400 |
+
"learning_rate": 0.00018198314039132143,
|
| 401 |
+
"loss": 0.30583988189697264,
|
| 402 |
+
"mean_token_accuracy": 0.8961733293533325,
|
| 403 |
+
"num_tokens": 2498321.0,
|
| 404 |
+
"step": 1750
|
| 405 |
+
},
|
| 406 |
+
{
|
| 407 |
+
"entropy": 0.3791545969247818,
|
| 408 |
+
"epoch": 4.628056628056628,
|
| 409 |
+
"grad_norm": 0.6575382351875305,
|
| 410 |
+
"learning_rate": 0.00017622920797738184,
|
| 411 |
+
"loss": 0.3088031005859375,
|
| 412 |
+
"mean_token_accuracy": 0.8960050916671753,
|
| 413 |
+
"num_tokens": 2570321.0,
|
| 414 |
+
"step": 1800
|
| 415 |
+
},
|
| 416 |
+
{
|
| 417 |
+
"entropy": 0.3946831756830215,
|
| 418 |
+
"epoch": 4.756756756756757,
|
| 419 |
+
"grad_norm": 0.5351552963256836,
|
| 420 |
+
"learning_rate": 0.00017039325135515207,
|
| 421 |
+
"loss": 0.3229162979125977,
|
| 422 |
+
"mean_token_accuracy": 0.8920552498102188,
|
| 423 |
+
"num_tokens": 2642851.0,
|
| 424 |
+
"step": 1850
|
| 425 |
+
},
|
| 426 |
+
{
|
| 427 |
+
"entropy": 0.37298239797353744,
|
| 428 |
+
"epoch": 4.885456885456885,
|
| 429 |
+
"grad_norm": 0.7624587416648865,
|
| 430 |
+
"learning_rate": 0.00016448701665269964,
|
| 431 |
+
"loss": 0.3067934799194336,
|
| 432 |
+
"mean_token_accuracy": 0.8951873427629471,
|
| 433 |
+
"num_tokens": 2715629.0,
|
| 434 |
+
"step": 1900
|
| 435 |
+
},
|
| 436 |
+
{
|
| 437 |
+
"epoch": 5.0,
|
| 438 |
+
"eval_entropy": 0.3738150204887095,
|
| 439 |
+
"eval_loss": 0.7131896615028381,
|
| 440 |
+
"eval_mean_token_accuracy": 0.8207752468045225,
|
| 441 |
+
"eval_num_tokens": 2777465.0,
|
| 442 |
+
"eval_runtime": 162.1244,
|
| 443 |
+
"eval_samples_per_second": 9.53,
|
| 444 |
+
"eval_steps_per_second": 1.197,
|
| 445 |
+
"step": 1945
|
| 446 |
+
}
|
| 447 |
+
],
|
| 448 |
+
"logging_steps": 50,
|
| 449 |
+
"max_steps": 3890,
|
| 450 |
+
"num_input_tokens_seen": 0,
|
| 451 |
+
"num_train_epochs": 10,
|
| 452 |
+
"save_steps": 500,
|
| 453 |
+
"stateful_callbacks": {
|
| 454 |
+
"TrainerControl": {
|
| 455 |
+
"args": {
|
| 456 |
+
"should_epoch_stop": false,
|
| 457 |
+
"should_evaluate": false,
|
| 458 |
+
"should_log": false,
|
| 459 |
+
"should_save": true,
|
| 460 |
+
"should_training_stop": false
|
| 461 |
+
},
|
| 462 |
+
"attributes": {}
|
| 463 |
+
}
|
| 464 |
+
},
|
| 465 |
+
"total_flos": 4.653233039031091e+17,
|
| 466 |
+
"train_batch_size": 8,
|
| 467 |
+
"trial_name": null,
|
| 468 |
+
"trial_params": null
|
| 469 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/README.md
ADDED
|
@@ -0,0 +1,209 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen3-14B-Base
|
| 3 |
+
library_name: peft
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:Qwen/Qwen3-14B-Base
|
| 7 |
+
- lora
|
| 8 |
+
- sft
|
| 9 |
+
- transformers
|
| 10 |
+
- trl
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Model Card for Model ID
|
| 14 |
+
|
| 15 |
+
<!-- Provide a quick summary of what the model is/does. -->
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
## Model Details
|
| 20 |
+
|
| 21 |
+
### Model Description
|
| 22 |
+
|
| 23 |
+
<!-- Provide a longer summary of what this model is. -->
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
- **Developed by:** [More Information Needed]
|
| 28 |
+
- **Funded by [optional]:** [More Information Needed]
|
| 29 |
+
- **Shared by [optional]:** [More Information Needed]
|
| 30 |
+
- **Model type:** [More Information Needed]
|
| 31 |
+
- **Language(s) (NLP):** [More Information Needed]
|
| 32 |
+
- **License:** [More Information Needed]
|
| 33 |
+
- **Finetuned from model [optional]:** [More Information Needed]
|
| 34 |
+
|
| 35 |
+
### Model Sources [optional]
|
| 36 |
+
|
| 37 |
+
<!-- Provide the basic links for the model. -->
|
| 38 |
+
|
| 39 |
+
- **Repository:** [More Information Needed]
|
| 40 |
+
- **Paper [optional]:** [More Information Needed]
|
| 41 |
+
- **Demo [optional]:** [More Information Needed]
|
| 42 |
+
|
| 43 |
+
## Uses
|
| 44 |
+
|
| 45 |
+
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
| 46 |
+
|
| 47 |
+
### Direct Use
|
| 48 |
+
|
| 49 |
+
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
| 50 |
+
|
| 51 |
+
[More Information Needed]
|
| 52 |
+
|
| 53 |
+
### Downstream Use [optional]
|
| 54 |
+
|
| 55 |
+
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
| 56 |
+
|
| 57 |
+
[More Information Needed]
|
| 58 |
+
|
| 59 |
+
### Out-of-Scope Use
|
| 60 |
+
|
| 61 |
+
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
| 62 |
+
|
| 63 |
+
[More Information Needed]
|
| 64 |
+
|
| 65 |
+
## Bias, Risks, and Limitations
|
| 66 |
+
|
| 67 |
+
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
| 68 |
+
|
| 69 |
+
[More Information Needed]
|
| 70 |
+
|
| 71 |
+
### Recommendations
|
| 72 |
+
|
| 73 |
+
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
| 74 |
+
|
| 75 |
+
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
| 76 |
+
|
| 77 |
+
## How to Get Started with the Model
|
| 78 |
+
|
| 79 |
+
Use the code below to get started with the model.
|
| 80 |
+
|
| 81 |
+
[More Information Needed]
|
| 82 |
+
|
| 83 |
+
## Training Details
|
| 84 |
+
|
| 85 |
+
### Training Data
|
| 86 |
+
|
| 87 |
+
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
| 88 |
+
|
| 89 |
+
[More Information Needed]
|
| 90 |
+
|
| 91 |
+
### Training Procedure
|
| 92 |
+
|
| 93 |
+
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
| 94 |
+
|
| 95 |
+
#### Preprocessing [optional]
|
| 96 |
+
|
| 97 |
+
[More Information Needed]
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
#### Training Hyperparameters
|
| 101 |
+
|
| 102 |
+
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
| 103 |
+
|
| 104 |
+
#### Speeds, Sizes, Times [optional]
|
| 105 |
+
|
| 106 |
+
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
| 107 |
+
|
| 108 |
+
[More Information Needed]
|
| 109 |
+
|
| 110 |
+
## Evaluation
|
| 111 |
+
|
| 112 |
+
<!-- This section describes the evaluation protocols and provides the results. -->
|
| 113 |
+
|
| 114 |
+
### Testing Data, Factors & Metrics
|
| 115 |
+
|
| 116 |
+
#### Testing Data
|
| 117 |
+
|
| 118 |
+
<!-- This should link to a Dataset Card if possible. -->
|
| 119 |
+
|
| 120 |
+
[More Information Needed]
|
| 121 |
+
|
| 122 |
+
#### Factors
|
| 123 |
+
|
| 124 |
+
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
| 125 |
+
|
| 126 |
+
[More Information Needed]
|
| 127 |
+
|
| 128 |
+
#### Metrics
|
| 129 |
+
|
| 130 |
+
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
| 131 |
+
|
| 132 |
+
[More Information Needed]
|
| 133 |
+
|
| 134 |
+
### Results
|
| 135 |
+
|
| 136 |
+
[More Information Needed]
|
| 137 |
+
|
| 138 |
+
#### Summary
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
## Model Examination [optional]
|
| 143 |
+
|
| 144 |
+
<!-- Relevant interpretability work for the model goes here -->
|
| 145 |
+
|
| 146 |
+
[More Information Needed]
|
| 147 |
+
|
| 148 |
+
## Environmental Impact
|
| 149 |
+
|
| 150 |
+
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
| 151 |
+
|
| 152 |
+
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
| 153 |
+
|
| 154 |
+
- **Hardware Type:** [More Information Needed]
|
| 155 |
+
- **Hours used:** [More Information Needed]
|
| 156 |
+
- **Cloud Provider:** [More Information Needed]
|
| 157 |
+
- **Compute Region:** [More Information Needed]
|
| 158 |
+
- **Carbon Emitted:** [More Information Needed]
|
| 159 |
+
|
| 160 |
+
## Technical Specifications [optional]
|
| 161 |
+
|
| 162 |
+
### Model Architecture and Objective
|
| 163 |
+
|
| 164 |
+
[More Information Needed]
|
| 165 |
+
|
| 166 |
+
### Compute Infrastructure
|
| 167 |
+
|
| 168 |
+
[More Information Needed]
|
| 169 |
+
|
| 170 |
+
#### Hardware
|
| 171 |
+
|
| 172 |
+
[More Information Needed]
|
| 173 |
+
|
| 174 |
+
#### Software
|
| 175 |
+
|
| 176 |
+
[More Information Needed]
|
| 177 |
+
|
| 178 |
+
## Citation [optional]
|
| 179 |
+
|
| 180 |
+
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
| 181 |
+
|
| 182 |
+
**BibTeX:**
|
| 183 |
+
|
| 184 |
+
[More Information Needed]
|
| 185 |
+
|
| 186 |
+
**APA:**
|
| 187 |
+
|
| 188 |
+
[More Information Needed]
|
| 189 |
+
|
| 190 |
+
## Glossary [optional]
|
| 191 |
+
|
| 192 |
+
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
| 193 |
+
|
| 194 |
+
[More Information Needed]
|
| 195 |
+
|
| 196 |
+
## More Information [optional]
|
| 197 |
+
|
| 198 |
+
[More Information Needed]
|
| 199 |
+
|
| 200 |
+
## Model Card Authors [optional]
|
| 201 |
+
|
| 202 |
+
[More Information Needed]
|
| 203 |
+
|
| 204 |
+
## Model Card Contact
|
| 205 |
+
|
| 206 |
+
[More Information Needed]
|
| 207 |
+
### Framework versions
|
| 208 |
+
|
| 209 |
+
- PEFT 0.19.1
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/adapter_config.json
ADDED
|
@@ -0,0 +1,48 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "Qwen/Qwen3-14B-Base",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"layer_replication": null,
|
| 16 |
+
"layers_pattern": null,
|
| 17 |
+
"layers_to_transform": null,
|
| 18 |
+
"loftq_config": {},
|
| 19 |
+
"lora_alpha": 256,
|
| 20 |
+
"lora_bias": false,
|
| 21 |
+
"lora_dropout": 0.00237968804112545,
|
| 22 |
+
"lora_ga_config": null,
|
| 23 |
+
"megatron_config": null,
|
| 24 |
+
"megatron_core": "megatron.core",
|
| 25 |
+
"modules_to_save": null,
|
| 26 |
+
"peft_type": "LORA",
|
| 27 |
+
"peft_version": "0.19.1",
|
| 28 |
+
"qalora_group_size": 16,
|
| 29 |
+
"r": 128,
|
| 30 |
+
"rank_pattern": {},
|
| 31 |
+
"revision": null,
|
| 32 |
+
"target_modules": [
|
| 33 |
+
"gate_proj",
|
| 34 |
+
"q_proj",
|
| 35 |
+
"up_proj",
|
| 36 |
+
"o_proj",
|
| 37 |
+
"v_proj",
|
| 38 |
+
"k_proj",
|
| 39 |
+
"down_proj"
|
| 40 |
+
],
|
| 41 |
+
"target_parameters": null,
|
| 42 |
+
"task_type": "CAUSAL_LM",
|
| 43 |
+
"trainable_token_indices": null,
|
| 44 |
+
"use_bdlora": null,
|
| 45 |
+
"use_dora": false,
|
| 46 |
+
"use_qalora": false,
|
| 47 |
+
"use_rslora": false
|
| 48 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/chat_template.jinja
ADDED
|
@@ -0,0 +1,85 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- if tools %}
|
| 2 |
+
{{- '<|im_start|>system\n' }}
|
| 3 |
+
{%- if messages[0].role == 'system' %}
|
| 4 |
+
{{- messages[0].content + '\n\n' }}
|
| 5 |
+
{%- endif %}
|
| 6 |
+
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
| 7 |
+
{%- for tool in tools %}
|
| 8 |
+
{{- "\n" }}
|
| 9 |
+
{{- tool | tojson }}
|
| 10 |
+
{%- endfor %}
|
| 11 |
+
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
| 12 |
+
{%- else %}
|
| 13 |
+
{%- if messages[0].role == 'system' %}
|
| 14 |
+
{{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
|
| 15 |
+
{%- endif %}
|
| 16 |
+
{%- endif %}
|
| 17 |
+
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
| 18 |
+
{%- for message in messages[::-1] %}
|
| 19 |
+
{%- set index = (messages|length - 1) - loop.index0 %}
|
| 20 |
+
{%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
|
| 21 |
+
{%- set ns.multi_step_tool = false %}
|
| 22 |
+
{%- set ns.last_query_index = index %}
|
| 23 |
+
{%- endif %}
|
| 24 |
+
{%- endfor %}
|
| 25 |
+
{%- for message in messages %}
|
| 26 |
+
{%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
|
| 27 |
+
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
|
| 28 |
+
{%- elif message.role == "assistant" %}
|
| 29 |
+
{%- set content = message.content %}
|
| 30 |
+
{%- set reasoning_content = '' %}
|
| 31 |
+
{%- if message.reasoning_content is defined and message.reasoning_content is not none %}
|
| 32 |
+
{%- set reasoning_content = message.reasoning_content %}
|
| 33 |
+
{%- else %}
|
| 34 |
+
{%- if '</think>' in message.content %}
|
| 35 |
+
{%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
|
| 36 |
+
{%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
| 37 |
+
{%- endif %}
|
| 38 |
+
{%- endif %}
|
| 39 |
+
{%- if loop.index0 > ns.last_query_index %}
|
| 40 |
+
{%- if loop.last or (not loop.last and reasoning_content) %}
|
| 41 |
+
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
|
| 42 |
+
{%- else %}
|
| 43 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 44 |
+
{%- endif %}
|
| 45 |
+
{%- else %}
|
| 46 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 47 |
+
{%- endif %}
|
| 48 |
+
{%- if message.tool_calls %}
|
| 49 |
+
{%- for tool_call in message.tool_calls %}
|
| 50 |
+
{%- if (loop.first and content) or (not loop.first) %}
|
| 51 |
+
{{- '\n' }}
|
| 52 |
+
{%- endif %}
|
| 53 |
+
{%- if tool_call.function %}
|
| 54 |
+
{%- set tool_call = tool_call.function %}
|
| 55 |
+
{%- endif %}
|
| 56 |
+
{{- '<tool_call>\n{"name": "' }}
|
| 57 |
+
{{- tool_call.name }}
|
| 58 |
+
{{- '", "arguments": ' }}
|
| 59 |
+
{%- if tool_call.arguments is string %}
|
| 60 |
+
{{- tool_call.arguments }}
|
| 61 |
+
{%- else %}
|
| 62 |
+
{{- tool_call.arguments | tojson }}
|
| 63 |
+
{%- endif %}
|
| 64 |
+
{{- '}\n</tool_call>' }}
|
| 65 |
+
{%- endfor %}
|
| 66 |
+
{%- endif %}
|
| 67 |
+
{{- '<|im_end|>\n' }}
|
| 68 |
+
{%- elif message.role == "tool" %}
|
| 69 |
+
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
|
| 70 |
+
{{- '<|im_start|>user' }}
|
| 71 |
+
{%- endif %}
|
| 72 |
+
{{- '\n<tool_response>\n' }}
|
| 73 |
+
{{- message.content }}
|
| 74 |
+
{{- '\n</tool_response>' }}
|
| 75 |
+
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
| 76 |
+
{{- '<|im_end|>\n' }}
|
| 77 |
+
{%- endif %}
|
| 78 |
+
{%- endif %}
|
| 79 |
+
{%- endfor %}
|
| 80 |
+
{%- if add_generation_prompt %}
|
| 81 |
+
{{- '<|im_start|>assistant\n' }}
|
| 82 |
+
{%- if enable_thinking is defined and enable_thinking is false %}
|
| 83 |
+
{{- '<think>\n\n</think>\n\n' }}
|
| 84 |
+
{%- endif %}
|
| 85 |
+
{%- endif %}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/tokenizer_config.json
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"backend": "tokenizers",
|
| 4 |
+
"bos_token": null,
|
| 5 |
+
"clean_up_tokenization_spaces": false,
|
| 6 |
+
"eos_token": "<|endoftext|>",
|
| 7 |
+
"errors": "replace",
|
| 8 |
+
"extra_special_tokens": [
|
| 9 |
+
"<|im_start|>",
|
| 10 |
+
"<|im_end|>",
|
| 11 |
+
"<|object_ref_start|>",
|
| 12 |
+
"<|object_ref_end|>",
|
| 13 |
+
"<|box_start|>",
|
| 14 |
+
"<|box_end|>",
|
| 15 |
+
"<|quad_start|>",
|
| 16 |
+
"<|quad_end|>",
|
| 17 |
+
"<|vision_start|>",
|
| 18 |
+
"<|vision_end|>",
|
| 19 |
+
"<|vision_pad|>",
|
| 20 |
+
"<|image_pad|>",
|
| 21 |
+
"<|video_pad|>"
|
| 22 |
+
],
|
| 23 |
+
"is_local": false,
|
| 24 |
+
"model_max_length": 131072,
|
| 25 |
+
"pad_token": "<|endoftext|>",
|
| 26 |
+
"split_special_tokens": false,
|
| 27 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 28 |
+
"unk_token": null
|
| 29 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2334/trainer_state.json
ADDED
|
@@ -0,0 +1,560 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"best_global_step": null,
|
| 3 |
+
"best_metric": null,
|
| 4 |
+
"best_model_checkpoint": null,
|
| 5 |
+
"epoch": 6.0,
|
| 6 |
+
"eval_steps": 500,
|
| 7 |
+
"global_step": 2334,
|
| 8 |
+
"is_hyper_param_search": false,
|
| 9 |
+
"is_local_process_zero": true,
|
| 10 |
+
"is_world_process_zero": true,
|
| 11 |
+
"log_history": [
|
| 12 |
+
{
|
| 13 |
+
"entropy": 1.6110110306739807,
|
| 14 |
+
"epoch": 0.1287001287001287,
|
| 15 |
+
"grad_norm": 0.8611119389533997,
|
| 16 |
+
"learning_rate": 3.4130257962133866e-05,
|
| 17 |
+
"loss": 1.53956298828125,
|
| 18 |
+
"mean_token_accuracy": 0.6652342769503593,
|
| 19 |
+
"num_tokens": 73407.0,
|
| 20 |
+
"step": 50
|
| 21 |
+
},
|
| 22 |
+
{
|
| 23 |
+
"entropy": 0.8662820833921433,
|
| 24 |
+
"epoch": 0.2574002574002574,
|
| 25 |
+
"grad_norm": 0.7405035495758057,
|
| 26 |
+
"learning_rate": 6.895705180104598e-05,
|
| 27 |
+
"loss": 0.7929539489746094,
|
| 28 |
+
"mean_token_accuracy": 0.7786347842216492,
|
| 29 |
+
"num_tokens": 143994.0,
|
| 30 |
+
"step": 100
|
| 31 |
+
},
|
| 32 |
+
{
|
| 33 |
+
"entropy": 0.7912819278240204,
|
| 34 |
+
"epoch": 0.3861003861003861,
|
| 35 |
+
"grad_norm": 0.6105485558509827,
|
| 36 |
+
"learning_rate": 0.00010378384563995809,
|
| 37 |
+
"loss": 0.7236511993408203,
|
| 38 |
+
"mean_token_accuracy": 0.7925467795133591,
|
| 39 |
+
"num_tokens": 216171.0,
|
| 40 |
+
"step": 150
|
| 41 |
+
},
|
| 42 |
+
{
|
| 43 |
+
"entropy": 0.7777709531784057,
|
| 44 |
+
"epoch": 0.5148005148005148,
|
| 45 |
+
"grad_norm": 0.5781793594360352,
|
| 46 |
+
"learning_rate": 0.0001386106394788702,
|
| 47 |
+
"loss": 0.7024919891357422,
|
| 48 |
+
"mean_token_accuracy": 0.7969876372814179,
|
| 49 |
+
"num_tokens": 284702.0,
|
| 50 |
+
"step": 200
|
| 51 |
+
},
|
| 52 |
+
{
|
| 53 |
+
"entropy": 0.7632609683275223,
|
| 54 |
+
"epoch": 0.6435006435006435,
|
| 55 |
+
"grad_norm": 0.6553444862365723,
|
| 56 |
+
"learning_rate": 0.00017343743331778232,
|
| 57 |
+
"loss": 0.7000718688964844,
|
| 58 |
+
"mean_token_accuracy": 0.7994742071628571,
|
| 59 |
+
"num_tokens": 356393.0,
|
| 60 |
+
"step": 250
|
| 61 |
+
},
|
| 62 |
+
{
|
| 63 |
+
"entropy": 0.754165632724762,
|
| 64 |
+
"epoch": 0.7722007722007722,
|
| 65 |
+
"grad_norm": 0.634222149848938,
|
| 66 |
+
"learning_rate": 0.00020826422715669444,
|
| 67 |
+
"loss": 0.6909049987792969,
|
| 68 |
+
"mean_token_accuracy": 0.8016809666156769,
|
| 69 |
+
"num_tokens": 426916.0,
|
| 70 |
+
"step": 300
|
| 71 |
+
},
|
| 72 |
+
{
|
| 73 |
+
"entropy": 0.7530711203813553,
|
| 74 |
+
"epoch": 0.9009009009009009,
|
| 75 |
+
"grad_norm": 0.7020156383514404,
|
| 76 |
+
"learning_rate": 0.00024309102099560653,
|
| 77 |
+
"loss": 0.69554931640625,
|
| 78 |
+
"mean_token_accuracy": 0.8002807641029358,
|
| 79 |
+
"num_tokens": 499603.0,
|
| 80 |
+
"step": 350
|
| 81 |
+
},
|
| 82 |
+
{
|
| 83 |
+
"epoch": 1.0,
|
| 84 |
+
"eval_entropy": 0.5728044164242204,
|
| 85 |
+
"eval_loss": 0.6785851716995239,
|
| 86 |
+
"eval_mean_token_accuracy": 0.798433246993527,
|
| 87 |
+
"eval_num_tokens": 555493.0,
|
| 88 |
+
"eval_runtime": 162.4129,
|
| 89 |
+
"eval_samples_per_second": 9.513,
|
| 90 |
+
"eval_steps_per_second": 1.194,
|
| 91 |
+
"step": 389
|
| 92 |
+
},
|
| 93 |
+
{
|
| 94 |
+
"entropy": 0.7489313946829902,
|
| 95 |
+
"epoch": 1.0283140283140284,
|
| 96 |
+
"grad_norm": 0.7532815933227539,
|
| 97 |
+
"learning_rate": 0.0002709470016827303,
|
| 98 |
+
"loss": 0.6856581878662109,
|
| 99 |
+
"mean_token_accuracy": 0.8023027079273956,
|
| 100 |
+
"num_tokens": 570813.0,
|
| 101 |
+
"step": 400
|
| 102 |
+
},
|
| 103 |
+
{
|
| 104 |
+
"entropy": 0.7238409864902496,
|
| 105 |
+
"epoch": 1.157014157014157,
|
| 106 |
+
"grad_norm": 1.2016215324401855,
|
| 107 |
+
"learning_rate": 0.0002707561443541359,
|
| 108 |
+
"loss": 0.6699818420410156,
|
| 109 |
+
"mean_token_accuracy": 0.8070302194356919,
|
| 110 |
+
"num_tokens": 642956.0,
|
| 111 |
+
"step": 450
|
| 112 |
+
},
|
| 113 |
+
{
|
| 114 |
+
"entropy": 0.7388938587903976,
|
| 115 |
+
"epoch": 1.2857142857142856,
|
| 116 |
+
"grad_norm": 0.7279272079467773,
|
| 117 |
+
"learning_rate": 0.0002702930068622498,
|
| 118 |
+
"loss": 0.6728517150878907,
|
| 119 |
+
"mean_token_accuracy": 0.8049580943584442,
|
| 120 |
+
"num_tokens": 714498.0,
|
| 121 |
+
"step": 500
|
| 122 |
+
},
|
| 123 |
+
{
|
| 124 |
+
"entropy": 0.7284485149383545,
|
| 125 |
+
"epoch": 1.4144144144144144,
|
| 126 |
+
"grad_norm": 0.8563987016677856,
|
| 127 |
+
"learning_rate": 0.0002695585213716931,
|
| 128 |
+
"loss": 0.6657986450195312,
|
| 129 |
+
"mean_token_accuracy": 0.8085588800907135,
|
| 130 |
+
"num_tokens": 785595.0,
|
| 131 |
+
"step": 550
|
| 132 |
+
},
|
| 133 |
+
{
|
| 134 |
+
"entropy": 0.6960959500074386,
|
| 135 |
+
"epoch": 1.5431145431145432,
|
| 136 |
+
"grad_norm": 0.5145474672317505,
|
| 137 |
+
"learning_rate": 0.0002685541661937683,
|
| 138 |
+
"loss": 0.6358638763427734,
|
| 139 |
+
"mean_token_accuracy": 0.8131621342897415,
|
| 140 |
+
"num_tokens": 857551.0,
|
| 141 |
+
"step": 600
|
| 142 |
+
},
|
| 143 |
+
{
|
| 144 |
+
"entropy": 0.7220414417982102,
|
| 145 |
+
"epoch": 1.6718146718146718,
|
| 146 |
+
"grad_norm": 0.9381059408187866,
|
| 147 |
+
"learning_rate": 0.00026728196281103746,
|
| 148 |
+
"loss": 0.6531407928466797,
|
| 149 |
+
"mean_token_accuracy": 0.811619822382927,
|
| 150 |
+
"num_tokens": 926864.0,
|
| 151 |
+
"step": 650
|
| 152 |
+
},
|
| 153 |
+
{
|
| 154 |
+
"entropy": 0.6955459499359131,
|
| 155 |
+
"epoch": 1.8005148005148004,
|
| 156 |
+
"grad_norm": 0.6115108728408813,
|
| 157 |
+
"learning_rate": 0.0002657444718086503,
|
| 158 |
+
"loss": 0.6269588088989257,
|
| 159 |
+
"mean_token_accuracy": 0.8155373805761337,
|
| 160 |
+
"num_tokens": 998824.0,
|
| 161 |
+
"step": 700
|
| 162 |
+
},
|
| 163 |
+
{
|
| 164 |
+
"entropy": 0.6929812705516816,
|
| 165 |
+
"epoch": 1.9292149292149292,
|
| 166 |
+
"grad_norm": 0.749489426612854,
|
| 167 |
+
"learning_rate": 0.0002639447877206115,
|
| 168 |
+
"loss": 0.629054069519043,
|
| 169 |
+
"mean_token_accuracy": 0.8183021235466004,
|
| 170 |
+
"num_tokens": 1069332.0,
|
| 171 |
+
"step": 750
|
| 172 |
+
},
|
| 173 |
+
{
|
| 174 |
+
"epoch": 2.0,
|
| 175 |
+
"eval_entropy": 0.5833694102223387,
|
| 176 |
+
"eval_loss": 0.6328718662261963,
|
| 177 |
+
"eval_mean_token_accuracy": 0.8159044071571114,
|
| 178 |
+
"eval_num_tokens": 1110986.0,
|
| 179 |
+
"eval_runtime": 161.885,
|
| 180 |
+
"eval_samples_per_second": 9.544,
|
| 181 |
+
"eval_steps_per_second": 1.198,
|
| 182 |
+
"step": 778
|
| 183 |
+
},
|
| 184 |
+
{
|
| 185 |
+
"entropy": 0.647241060480927,
|
| 186 |
+
"epoch": 2.056628056628057,
|
| 187 |
+
"grad_norm": 0.7070767879486084,
|
| 188 |
+
"learning_rate": 0.00026188653280135975,
|
| 189 |
+
"loss": 0.5823922348022461,
|
| 190 |
+
"mean_token_accuracy": 0.8265365301960647,
|
| 191 |
+
"num_tokens": 1141195.0,
|
| 192 |
+
"step": 800
|
| 193 |
+
},
|
| 194 |
+
{
|
| 195 |
+
"entropy": 0.5995378407835961,
|
| 196 |
+
"epoch": 2.1853281853281854,
|
| 197 |
+
"grad_norm": 0.8090486526489258,
|
| 198 |
+
"learning_rate": 0.0002595738497351955,
|
| 199 |
+
"loss": 0.5325597763061524,
|
| 200 |
+
"mean_token_accuracy": 0.8369336777925491,
|
| 201 |
+
"num_tokens": 1210708.0,
|
| 202 |
+
"step": 850
|
| 203 |
+
},
|
| 204 |
+
{
|
| 205 |
+
"entropy": 0.6043145382404327,
|
| 206 |
+
"epoch": 2.314028314028314,
|
| 207 |
+
"grad_norm": 0.8279913067817688,
|
| 208 |
+
"learning_rate": 0.00025701139329823054,
|
| 209 |
+
"loss": 0.5414446258544922,
|
| 210 |
+
"mean_token_accuracy": 0.8361396533250809,
|
| 211 |
+
"num_tokens": 1283441.0,
|
| 212 |
+
"step": 900
|
| 213 |
+
},
|
| 214 |
+
{
|
| 215 |
+
"entropy": 0.5953224584460258,
|
| 216 |
+
"epoch": 2.4427284427284426,
|
| 217 |
+
"grad_norm": 0.6075023412704468,
|
| 218 |
+
"learning_rate": 0.00025420432098964183,
|
| 219 |
+
"loss": 0.536654167175293,
|
| 220 |
+
"mean_token_accuracy": 0.8340093129873276,
|
| 221 |
+
"num_tokens": 1356418.0,
|
| 222 |
+
"step": 950
|
| 223 |
+
},
|
| 224 |
+
{
|
| 225 |
+
"entropy": 0.5998479858040809,
|
| 226 |
+
"epoch": 2.571428571428571,
|
| 227 |
+
"grad_norm": 1.0311471223831177,
|
| 228 |
+
"learning_rate": 0.0002511582826510862,
|
| 229 |
+
"loss": 0.5372924423217773,
|
| 230 |
+
"mean_token_accuracy": 0.8366045409440994,
|
| 231 |
+
"num_tokens": 1427797.0,
|
| 232 |
+
"step": 1000
|
| 233 |
+
},
|
| 234 |
+
{
|
| 235 |
+
"entropy": 0.5980158120393753,
|
| 236 |
+
"epoch": 2.7001287001287,
|
| 237 |
+
"grad_norm": 0.5971426367759705,
|
| 238 |
+
"learning_rate": 0.0002478794090951689,
|
| 239 |
+
"loss": 0.5392885208129883,
|
| 240 |
+
"mean_token_accuracy": 0.8347727072238922,
|
| 241 |
+
"num_tokens": 1498082.0,
|
| 242 |
+
"step": 1050
|
| 243 |
+
},
|
| 244 |
+
{
|
| 245 |
+
"entropy": 0.5990680930018425,
|
| 246 |
+
"epoch": 2.828828828828829,
|
| 247 |
+
"grad_norm": 0.5662627220153809,
|
| 248 |
+
"learning_rate": 0.0002443742997658538,
|
| 249 |
+
"loss": 0.5360498428344727,
|
| 250 |
+
"mean_token_accuracy": 0.8371847170591354,
|
| 251 |
+
"num_tokens": 1568798.0,
|
| 252 |
+
"step": 1100
|
| 253 |
+
},
|
| 254 |
+
{
|
| 255 |
+
"entropy": 0.5957039377093315,
|
| 256 |
+
"epoch": 2.9575289575289574,
|
| 257 |
+
"grad_norm": 0.5043798685073853,
|
| 258 |
+
"learning_rate": 0.00024065000945565205,
|
| 259 |
+
"loss": 0.5342231369018555,
|
| 260 |
+
"mean_token_accuracy": 0.8380735236406326,
|
| 261 |
+
"num_tokens": 1643449.0,
|
| 262 |
+
"step": 1150
|
| 263 |
+
},
|
| 264 |
+
{
|
| 265 |
+
"epoch": 3.0,
|
| 266 |
+
"eval_entropy": 0.5346978819861854,
|
| 267 |
+
"eval_loss": 0.6255015134811401,
|
| 268 |
+
"eval_mean_token_accuracy": 0.8202014476368108,
|
| 269 |
+
"eval_num_tokens": 1666479.0,
|
| 270 |
+
"eval_runtime": 161.6098,
|
| 271 |
+
"eval_samples_per_second": 9.56,
|
| 272 |
+
"eval_steps_per_second": 1.2,
|
| 273 |
+
"step": 1167
|
| 274 |
+
},
|
| 275 |
+
{
|
| 276 |
+
"entropy": 0.5212209137401196,
|
| 277 |
+
"epoch": 3.0849420849420848,
|
| 278 |
+
"grad_norm": 0.7095440626144409,
|
| 279 |
+
"learning_rate": 0.00023671403410632178,
|
| 280 |
+
"loss": 0.45311901092529294,
|
| 281 |
+
"mean_token_accuracy": 0.856362871449403,
|
| 282 |
+
"num_tokens": 1713536.0,
|
| 283 |
+
"step": 1200
|
| 284 |
+
},
|
| 285 |
+
{
|
| 286 |
+
"entropy": 0.4701593083143234,
|
| 287 |
+
"epoch": 3.213642213642214,
|
| 288 |
+
"grad_norm": 0.6408083438873291,
|
| 289 |
+
"learning_rate": 0.0002325742957216607,
|
| 290 |
+
"loss": 0.39916397094726563,
|
| 291 |
+
"mean_token_accuracy": 0.8698061722517013,
|
| 292 |
+
"num_tokens": 1785609.0,
|
| 293 |
+
"step": 1250
|
| 294 |
+
},
|
| 295 |
+
{
|
| 296 |
+
"entropy": 0.4766591975092888,
|
| 297 |
+
"epoch": 3.3423423423423424,
|
| 298 |
+
"grad_norm": 0.6415093541145325,
|
| 299 |
+
"learning_rate": 0.0002282391264227552,
|
| 300 |
+
"loss": 0.4116698455810547,
|
| 301 |
+
"mean_token_accuracy": 0.8679435575008392,
|
| 302 |
+
"num_tokens": 1858651.0,
|
| 303 |
+
"step": 1300
|
| 304 |
+
},
|
| 305 |
+
{
|
| 306 |
+
"entropy": 0.4937947469949722,
|
| 307 |
+
"epoch": 3.471042471042471,
|
| 308 |
+
"grad_norm": 0.6549825072288513,
|
| 309 |
+
"learning_rate": 0.00022371725167778054,
|
| 310 |
+
"loss": 0.4296376037597656,
|
| 311 |
+
"mean_token_accuracy": 0.8609357953071595,
|
| 312 |
+
"num_tokens": 1928692.0,
|
| 313 |
+
"step": 1350
|
| 314 |
+
},
|
| 315 |
+
{
|
| 316 |
+
"entropy": 0.4920153194665909,
|
| 317 |
+
"epoch": 3.5997425997425996,
|
| 318 |
+
"grad_norm": 0.6452126502990723,
|
| 319 |
+
"learning_rate": 0.00021901777274010406,
|
| 320 |
+
"loss": 0.4307489013671875,
|
| 321 |
+
"mean_token_accuracy": 0.8606827831268311,
|
| 322 |
+
"num_tokens": 1998668.0,
|
| 323 |
+
"step": 1400
|
| 324 |
+
},
|
| 325 |
+
{
|
| 326 |
+
"entropy": 0.490042342543602,
|
| 327 |
+
"epoch": 3.7284427284427286,
|
| 328 |
+
"grad_norm": 0.5727734565734863,
|
| 329 |
+
"learning_rate": 0.0002141501483300395,
|
| 330 |
+
"loss": 0.4295254135131836,
|
| 331 |
+
"mean_token_accuracy": 0.8616545403003693,
|
| 332 |
+
"num_tokens": 2072809.0,
|
| 333 |
+
"step": 1450
|
| 334 |
+
},
|
| 335 |
+
{
|
| 336 |
+
"entropy": 0.49857193052768706,
|
| 337 |
+
"epoch": 3.857142857142857,
|
| 338 |
+
"grad_norm": 0.7732954025268555,
|
| 339 |
+
"learning_rate": 0.00020912417559712133,
|
| 340 |
+
"loss": 0.4289303207397461,
|
| 341 |
+
"mean_token_accuracy": 0.8616443765163422,
|
| 342 |
+
"num_tokens": 2142475.0,
|
| 343 |
+
"step": 1500
|
| 344 |
+
},
|
| 345 |
+
{
|
| 346 |
+
"entropy": 0.4746784272789955,
|
| 347 |
+
"epoch": 3.985842985842986,
|
| 348 |
+
"grad_norm": 0.5791187882423401,
|
| 349 |
+
"learning_rate": 0.00020394997040121726,
|
| 350 |
+
"loss": 0.4180263900756836,
|
| 351 |
+
"mean_token_accuracy": 0.866080379486084,
|
| 352 |
+
"num_tokens": 2214044.0,
|
| 353 |
+
"step": 1550
|
| 354 |
+
},
|
| 355 |
+
{
|
| 356 |
+
"epoch": 4.0,
|
| 357 |
+
"eval_entropy": 0.4793926059585257,
|
| 358 |
+
"eval_loss": 0.627627968788147,
|
| 359 |
+
"eval_mean_token_accuracy": 0.8233870095813397,
|
| 360 |
+
"eval_num_tokens": 2221972.0,
|
| 361 |
+
"eval_runtime": 162.0225,
|
| 362 |
+
"eval_samples_per_second": 9.536,
|
| 363 |
+
"eval_steps_per_second": 1.197,
|
| 364 |
+
"step": 1556
|
| 365 |
+
},
|
| 366 |
+
{
|
| 367 |
+
"entropy": 0.38195489000792454,
|
| 368 |
+
"epoch": 4.113256113256114,
|
| 369 |
+
"grad_norm": 0.6002617478370667,
|
| 370 |
+
"learning_rate": 0.0001986379469521669,
|
| 371 |
+
"loss": 0.30819049835205076,
|
| 372 |
+
"mean_token_accuracy": 0.8977848634575353,
|
| 373 |
+
"num_tokens": 2282164.0,
|
| 374 |
+
"step": 1600
|
| 375 |
+
},
|
| 376 |
+
{
|
| 377 |
+
"entropy": 0.3655787402391434,
|
| 378 |
+
"epoch": 4.241956241956242,
|
| 379 |
+
"grad_norm": 0.7100041508674622,
|
| 380 |
+
"learning_rate": 0.00019319879684892634,
|
| 381 |
+
"loss": 0.29959835052490236,
|
| 382 |
+
"mean_token_accuracy": 0.8991208010911942,
|
| 383 |
+
"num_tokens": 2353213.0,
|
| 384 |
+
"step": 1650
|
| 385 |
+
},
|
| 386 |
+
{
|
| 387 |
+
"entropy": 0.3821141055226326,
|
| 388 |
+
"epoch": 4.370656370656371,
|
| 389 |
+
"grad_norm": 0.5848307013511658,
|
| 390 |
+
"learning_rate": 0.00018764346756040715,
|
| 391 |
+
"loss": 0.313802490234375,
|
| 392 |
+
"mean_token_accuracy": 0.895167955160141,
|
| 393 |
+
"num_tokens": 2425068.0,
|
| 394 |
+
"step": 1700
|
| 395 |
+
},
|
| 396 |
+
{
|
| 397 |
+
"entropy": 0.37083797007799146,
|
| 398 |
+
"epoch": 4.499356499356499,
|
| 399 |
+
"grad_norm": 0.6447024941444397,
|
| 400 |
+
"learning_rate": 0.00018198314039132143,
|
| 401 |
+
"loss": 0.30583988189697264,
|
| 402 |
+
"mean_token_accuracy": 0.8961733293533325,
|
| 403 |
+
"num_tokens": 2498321.0,
|
| 404 |
+
"step": 1750
|
| 405 |
+
},
|
| 406 |
+
{
|
| 407 |
+
"entropy": 0.3791545969247818,
|
| 408 |
+
"epoch": 4.628056628056628,
|
| 409 |
+
"grad_norm": 0.6575382351875305,
|
| 410 |
+
"learning_rate": 0.00017622920797738184,
|
| 411 |
+
"loss": 0.3088031005859375,
|
| 412 |
+
"mean_token_accuracy": 0.8960050916671753,
|
| 413 |
+
"num_tokens": 2570321.0,
|
| 414 |
+
"step": 1800
|
| 415 |
+
},
|
| 416 |
+
{
|
| 417 |
+
"entropy": 0.3946831756830215,
|
| 418 |
+
"epoch": 4.756756756756757,
|
| 419 |
+
"grad_norm": 0.5351552963256836,
|
| 420 |
+
"learning_rate": 0.00017039325135515207,
|
| 421 |
+
"loss": 0.3229162979125977,
|
| 422 |
+
"mean_token_accuracy": 0.8920552498102188,
|
| 423 |
+
"num_tokens": 2642851.0,
|
| 424 |
+
"step": 1850
|
| 425 |
+
},
|
| 426 |
+
{
|
| 427 |
+
"entropy": 0.37298239797353744,
|
| 428 |
+
"epoch": 4.885456885456885,
|
| 429 |
+
"grad_norm": 0.7624587416648865,
|
| 430 |
+
"learning_rate": 0.00016448701665269964,
|
| 431 |
+
"loss": 0.3067934799194336,
|
| 432 |
+
"mean_token_accuracy": 0.8951873427629471,
|
| 433 |
+
"num_tokens": 2715629.0,
|
| 434 |
+
"step": 1900
|
| 435 |
+
},
|
| 436 |
+
{
|
| 437 |
+
"epoch": 5.0,
|
| 438 |
+
"eval_entropy": 0.3738150204887095,
|
| 439 |
+
"eval_loss": 0.7131896615028381,
|
| 440 |
+
"eval_mean_token_accuracy": 0.8207752468045225,
|
| 441 |
+
"eval_num_tokens": 2777465.0,
|
| 442 |
+
"eval_runtime": 162.1244,
|
| 443 |
+
"eval_samples_per_second": 9.53,
|
| 444 |
+
"eval_steps_per_second": 1.197,
|
| 445 |
+
"step": 1945
|
| 446 |
+
},
|
| 447 |
+
{
|
| 448 |
+
"entropy": 0.377820266617669,
|
| 449 |
+
"epoch": 5.012870012870013,
|
| 450 |
+
"grad_norm": 0.41990190744400024,
|
| 451 |
+
"learning_rate": 0.00015852239144796624,
|
| 452 |
+
"loss": 0.3058685111999512,
|
| 453 |
+
"mean_token_accuracy": 0.8964343480389527,
|
| 454 |
+
"num_tokens": 2784896.0,
|
| 455 |
+
"step": 1950
|
| 456 |
+
},
|
| 457 |
+
{
|
| 458 |
+
"entropy": 0.2714502356946468,
|
| 459 |
+
"epoch": 5.141570141570142,
|
| 460 |
+
"grad_norm": 0.422568678855896,
|
| 461 |
+
"learning_rate": 0.00015251138084243995,
|
| 462 |
+
"loss": 0.2093442153930664,
|
| 463 |
+
"mean_token_accuracy": 0.9311346983909607,
|
| 464 |
+
"num_tokens": 2854374.0,
|
| 465 |
+
"step": 2000
|
| 466 |
+
},
|
| 467 |
+
{
|
| 468 |
+
"entropy": 0.268475965410471,
|
| 469 |
+
"epoch": 5.27027027027027,
|
| 470 |
+
"grad_norm": 0.6637414693832397,
|
| 471 |
+
"learning_rate": 0.0001464660832982852,
|
| 472 |
+
"loss": 0.20736080169677734,
|
| 473 |
+
"mean_token_accuracy": 0.9289199805259705,
|
| 474 |
+
"num_tokens": 2927362.0,
|
| 475 |
+
"step": 2050
|
| 476 |
+
},
|
| 477 |
+
{
|
| 478 |
+
"entropy": 0.2644876340031624,
|
| 479 |
+
"epoch": 5.398970398970399,
|
| 480 |
+
"grad_norm": 0.47317707538604736,
|
| 481 |
+
"learning_rate": 0.00014039866628756467,
|
| 482 |
+
"loss": 0.20464908599853515,
|
| 483 |
+
"mean_token_accuracy": 0.9300856202840805,
|
| 484 |
+
"num_tokens": 3000143.0,
|
| 485 |
+
"step": 2100
|
| 486 |
+
},
|
| 487 |
+
{
|
| 488 |
+
"entropy": 0.2675253136456013,
|
| 489 |
+
"epoch": 5.527670527670527,
|
| 490 |
+
"grad_norm": 0.5253982543945312,
|
| 491 |
+
"learning_rate": 0.00013432134180256338,
|
| 492 |
+
"loss": 0.21154335021972656,
|
| 493 |
+
"mean_token_accuracy": 0.9283734840154648,
|
| 494 |
+
"num_tokens": 3072561.0,
|
| 495 |
+
"step": 2150
|
| 496 |
+
},
|
| 497 |
+
{
|
| 498 |
+
"entropy": 0.27213907435536383,
|
| 499 |
+
"epoch": 5.656370656370656,
|
| 500 |
+
"grad_norm": 0.46738553047180176,
|
| 501 |
+
"learning_rate": 0.00012824634177650664,
|
| 502 |
+
"loss": 0.21339216232299804,
|
| 503 |
+
"mean_token_accuracy": 0.9272083270549775,
|
| 504 |
+
"num_tokens": 3144831.0,
|
| 505 |
+
"step": 2200
|
| 506 |
+
},
|
| 507 |
+
{
|
| 508 |
+
"entropy": 0.2785488124191761,
|
| 509 |
+
"epoch": 5.785070785070785,
|
| 510 |
+
"grad_norm": 0.4469502866268158,
|
| 511 |
+
"learning_rate": 0.00012218589346414205,
|
| 512 |
+
"loss": 0.21601097106933595,
|
| 513 |
+
"mean_token_accuracy": 0.9255663657188415,
|
| 514 |
+
"num_tokens": 3215960.0,
|
| 515 |
+
"step": 2250
|
| 516 |
+
},
|
| 517 |
+
{
|
| 518 |
+
"entropy": 0.2699935150146484,
|
| 519 |
+
"epoch": 5.913770913770914,
|
| 520 |
+
"grad_norm": 0.7359778881072998,
|
| 521 |
+
"learning_rate": 0.00011615219483173828,
|
| 522 |
+
"loss": 0.20725584030151367,
|
| 523 |
+
"mean_token_accuracy": 0.9286630594730377,
|
| 524 |
+
"num_tokens": 3287499.0,
|
| 525 |
+
"step": 2300
|
| 526 |
+
},
|
| 527 |
+
{
|
| 528 |
+
"epoch": 6.0,
|
| 529 |
+
"eval_entropy": 0.26417383682174783,
|
| 530 |
+
"eval_loss": 0.8880229592323303,
|
| 531 |
+
"eval_mean_token_accuracy": 0.8159987201395723,
|
| 532 |
+
"eval_num_tokens": 3332958.0,
|
| 533 |
+
"eval_runtime": 162.0991,
|
| 534 |
+
"eval_samples_per_second": 9.531,
|
| 535 |
+
"eval_steps_per_second": 1.197,
|
| 536 |
+
"step": 2334
|
| 537 |
+
}
|
| 538 |
+
],
|
| 539 |
+
"logging_steps": 50,
|
| 540 |
+
"max_steps": 3890,
|
| 541 |
+
"num_input_tokens_seen": 0,
|
| 542 |
+
"num_train_epochs": 10,
|
| 543 |
+
"save_steps": 500,
|
| 544 |
+
"stateful_callbacks": {
|
| 545 |
+
"TrainerControl": {
|
| 546 |
+
"args": {
|
| 547 |
+
"should_epoch_stop": false,
|
| 548 |
+
"should_evaluate": false,
|
| 549 |
+
"should_log": false,
|
| 550 |
+
"should_save": true,
|
| 551 |
+
"should_training_stop": false
|
| 552 |
+
},
|
| 553 |
+
"attributes": {}
|
| 554 |
+
}
|
| 555 |
+
},
|
| 556 |
+
"total_flos": 5.587061113467187e+17,
|
| 557 |
+
"train_batch_size": 8,
|
| 558 |
+
"trial_name": null,
|
| 559 |
+
"trial_params": null
|
| 560 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/README.md
ADDED
|
@@ -0,0 +1,209 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen3-14B-Base
|
| 3 |
+
library_name: peft
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:Qwen/Qwen3-14B-Base
|
| 7 |
+
- lora
|
| 8 |
+
- sft
|
| 9 |
+
- transformers
|
| 10 |
+
- trl
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Model Card for Model ID
|
| 14 |
+
|
| 15 |
+
<!-- Provide a quick summary of what the model is/does. -->
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
## Model Details
|
| 20 |
+
|
| 21 |
+
### Model Description
|
| 22 |
+
|
| 23 |
+
<!-- Provide a longer summary of what this model is. -->
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
- **Developed by:** [More Information Needed]
|
| 28 |
+
- **Funded by [optional]:** [More Information Needed]
|
| 29 |
+
- **Shared by [optional]:** [More Information Needed]
|
| 30 |
+
- **Model type:** [More Information Needed]
|
| 31 |
+
- **Language(s) (NLP):** [More Information Needed]
|
| 32 |
+
- **License:** [More Information Needed]
|
| 33 |
+
- **Finetuned from model [optional]:** [More Information Needed]
|
| 34 |
+
|
| 35 |
+
### Model Sources [optional]
|
| 36 |
+
|
| 37 |
+
<!-- Provide the basic links for the model. -->
|
| 38 |
+
|
| 39 |
+
- **Repository:** [More Information Needed]
|
| 40 |
+
- **Paper [optional]:** [More Information Needed]
|
| 41 |
+
- **Demo [optional]:** [More Information Needed]
|
| 42 |
+
|
| 43 |
+
## Uses
|
| 44 |
+
|
| 45 |
+
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
| 46 |
+
|
| 47 |
+
### Direct Use
|
| 48 |
+
|
| 49 |
+
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
| 50 |
+
|
| 51 |
+
[More Information Needed]
|
| 52 |
+
|
| 53 |
+
### Downstream Use [optional]
|
| 54 |
+
|
| 55 |
+
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
| 56 |
+
|
| 57 |
+
[More Information Needed]
|
| 58 |
+
|
| 59 |
+
### Out-of-Scope Use
|
| 60 |
+
|
| 61 |
+
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
| 62 |
+
|
| 63 |
+
[More Information Needed]
|
| 64 |
+
|
| 65 |
+
## Bias, Risks, and Limitations
|
| 66 |
+
|
| 67 |
+
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
| 68 |
+
|
| 69 |
+
[More Information Needed]
|
| 70 |
+
|
| 71 |
+
### Recommendations
|
| 72 |
+
|
| 73 |
+
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
| 74 |
+
|
| 75 |
+
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
| 76 |
+
|
| 77 |
+
## How to Get Started with the Model
|
| 78 |
+
|
| 79 |
+
Use the code below to get started with the model.
|
| 80 |
+
|
| 81 |
+
[More Information Needed]
|
| 82 |
+
|
| 83 |
+
## Training Details
|
| 84 |
+
|
| 85 |
+
### Training Data
|
| 86 |
+
|
| 87 |
+
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
| 88 |
+
|
| 89 |
+
[More Information Needed]
|
| 90 |
+
|
| 91 |
+
### Training Procedure
|
| 92 |
+
|
| 93 |
+
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
| 94 |
+
|
| 95 |
+
#### Preprocessing [optional]
|
| 96 |
+
|
| 97 |
+
[More Information Needed]
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
#### Training Hyperparameters
|
| 101 |
+
|
| 102 |
+
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
| 103 |
+
|
| 104 |
+
#### Speeds, Sizes, Times [optional]
|
| 105 |
+
|
| 106 |
+
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
| 107 |
+
|
| 108 |
+
[More Information Needed]
|
| 109 |
+
|
| 110 |
+
## Evaluation
|
| 111 |
+
|
| 112 |
+
<!-- This section describes the evaluation protocols and provides the results. -->
|
| 113 |
+
|
| 114 |
+
### Testing Data, Factors & Metrics
|
| 115 |
+
|
| 116 |
+
#### Testing Data
|
| 117 |
+
|
| 118 |
+
<!-- This should link to a Dataset Card if possible. -->
|
| 119 |
+
|
| 120 |
+
[More Information Needed]
|
| 121 |
+
|
| 122 |
+
#### Factors
|
| 123 |
+
|
| 124 |
+
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
| 125 |
+
|
| 126 |
+
[More Information Needed]
|
| 127 |
+
|
| 128 |
+
#### Metrics
|
| 129 |
+
|
| 130 |
+
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
| 131 |
+
|
| 132 |
+
[More Information Needed]
|
| 133 |
+
|
| 134 |
+
### Results
|
| 135 |
+
|
| 136 |
+
[More Information Needed]
|
| 137 |
+
|
| 138 |
+
#### Summary
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
## Model Examination [optional]
|
| 143 |
+
|
| 144 |
+
<!-- Relevant interpretability work for the model goes here -->
|
| 145 |
+
|
| 146 |
+
[More Information Needed]
|
| 147 |
+
|
| 148 |
+
## Environmental Impact
|
| 149 |
+
|
| 150 |
+
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
| 151 |
+
|
| 152 |
+
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
| 153 |
+
|
| 154 |
+
- **Hardware Type:** [More Information Needed]
|
| 155 |
+
- **Hours used:** [More Information Needed]
|
| 156 |
+
- **Cloud Provider:** [More Information Needed]
|
| 157 |
+
- **Compute Region:** [More Information Needed]
|
| 158 |
+
- **Carbon Emitted:** [More Information Needed]
|
| 159 |
+
|
| 160 |
+
## Technical Specifications [optional]
|
| 161 |
+
|
| 162 |
+
### Model Architecture and Objective
|
| 163 |
+
|
| 164 |
+
[More Information Needed]
|
| 165 |
+
|
| 166 |
+
### Compute Infrastructure
|
| 167 |
+
|
| 168 |
+
[More Information Needed]
|
| 169 |
+
|
| 170 |
+
#### Hardware
|
| 171 |
+
|
| 172 |
+
[More Information Needed]
|
| 173 |
+
|
| 174 |
+
#### Software
|
| 175 |
+
|
| 176 |
+
[More Information Needed]
|
| 177 |
+
|
| 178 |
+
## Citation [optional]
|
| 179 |
+
|
| 180 |
+
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
| 181 |
+
|
| 182 |
+
**BibTeX:**
|
| 183 |
+
|
| 184 |
+
[More Information Needed]
|
| 185 |
+
|
| 186 |
+
**APA:**
|
| 187 |
+
|
| 188 |
+
[More Information Needed]
|
| 189 |
+
|
| 190 |
+
## Glossary [optional]
|
| 191 |
+
|
| 192 |
+
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
| 193 |
+
|
| 194 |
+
[More Information Needed]
|
| 195 |
+
|
| 196 |
+
## More Information [optional]
|
| 197 |
+
|
| 198 |
+
[More Information Needed]
|
| 199 |
+
|
| 200 |
+
## Model Card Authors [optional]
|
| 201 |
+
|
| 202 |
+
[More Information Needed]
|
| 203 |
+
|
| 204 |
+
## Model Card Contact
|
| 205 |
+
|
| 206 |
+
[More Information Needed]
|
| 207 |
+
### Framework versions
|
| 208 |
+
|
| 209 |
+
- PEFT 0.19.1
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/adapter_config.json
ADDED
|
@@ -0,0 +1,48 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "Qwen/Qwen3-14B-Base",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"layer_replication": null,
|
| 16 |
+
"layers_pattern": null,
|
| 17 |
+
"layers_to_transform": null,
|
| 18 |
+
"loftq_config": {},
|
| 19 |
+
"lora_alpha": 256,
|
| 20 |
+
"lora_bias": false,
|
| 21 |
+
"lora_dropout": 0.00237968804112545,
|
| 22 |
+
"lora_ga_config": null,
|
| 23 |
+
"megatron_config": null,
|
| 24 |
+
"megatron_core": "megatron.core",
|
| 25 |
+
"modules_to_save": null,
|
| 26 |
+
"peft_type": "LORA",
|
| 27 |
+
"peft_version": "0.19.1",
|
| 28 |
+
"qalora_group_size": 16,
|
| 29 |
+
"r": 128,
|
| 30 |
+
"rank_pattern": {},
|
| 31 |
+
"revision": null,
|
| 32 |
+
"target_modules": [
|
| 33 |
+
"gate_proj",
|
| 34 |
+
"q_proj",
|
| 35 |
+
"up_proj",
|
| 36 |
+
"o_proj",
|
| 37 |
+
"v_proj",
|
| 38 |
+
"k_proj",
|
| 39 |
+
"down_proj"
|
| 40 |
+
],
|
| 41 |
+
"target_parameters": null,
|
| 42 |
+
"task_type": "CAUSAL_LM",
|
| 43 |
+
"trainable_token_indices": null,
|
| 44 |
+
"use_bdlora": null,
|
| 45 |
+
"use_dora": false,
|
| 46 |
+
"use_qalora": false,
|
| 47 |
+
"use_rslora": false
|
| 48 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/chat_template.jinja
ADDED
|
@@ -0,0 +1,85 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- if tools %}
|
| 2 |
+
{{- '<|im_start|>system\n' }}
|
| 3 |
+
{%- if messages[0].role == 'system' %}
|
| 4 |
+
{{- messages[0].content + '\n\n' }}
|
| 5 |
+
{%- endif %}
|
| 6 |
+
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
| 7 |
+
{%- for tool in tools %}
|
| 8 |
+
{{- "\n" }}
|
| 9 |
+
{{- tool | tojson }}
|
| 10 |
+
{%- endfor %}
|
| 11 |
+
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
| 12 |
+
{%- else %}
|
| 13 |
+
{%- if messages[0].role == 'system' %}
|
| 14 |
+
{{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
|
| 15 |
+
{%- endif %}
|
| 16 |
+
{%- endif %}
|
| 17 |
+
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
| 18 |
+
{%- for message in messages[::-1] %}
|
| 19 |
+
{%- set index = (messages|length - 1) - loop.index0 %}
|
| 20 |
+
{%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
|
| 21 |
+
{%- set ns.multi_step_tool = false %}
|
| 22 |
+
{%- set ns.last_query_index = index %}
|
| 23 |
+
{%- endif %}
|
| 24 |
+
{%- endfor %}
|
| 25 |
+
{%- for message in messages %}
|
| 26 |
+
{%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
|
| 27 |
+
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
|
| 28 |
+
{%- elif message.role == "assistant" %}
|
| 29 |
+
{%- set content = message.content %}
|
| 30 |
+
{%- set reasoning_content = '' %}
|
| 31 |
+
{%- if message.reasoning_content is defined and message.reasoning_content is not none %}
|
| 32 |
+
{%- set reasoning_content = message.reasoning_content %}
|
| 33 |
+
{%- else %}
|
| 34 |
+
{%- if '</think>' in message.content %}
|
| 35 |
+
{%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
|
| 36 |
+
{%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
| 37 |
+
{%- endif %}
|
| 38 |
+
{%- endif %}
|
| 39 |
+
{%- if loop.index0 > ns.last_query_index %}
|
| 40 |
+
{%- if loop.last or (not loop.last and reasoning_content) %}
|
| 41 |
+
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
|
| 42 |
+
{%- else %}
|
| 43 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 44 |
+
{%- endif %}
|
| 45 |
+
{%- else %}
|
| 46 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 47 |
+
{%- endif %}
|
| 48 |
+
{%- if message.tool_calls %}
|
| 49 |
+
{%- for tool_call in message.tool_calls %}
|
| 50 |
+
{%- if (loop.first and content) or (not loop.first) %}
|
| 51 |
+
{{- '\n' }}
|
| 52 |
+
{%- endif %}
|
| 53 |
+
{%- if tool_call.function %}
|
| 54 |
+
{%- set tool_call = tool_call.function %}
|
| 55 |
+
{%- endif %}
|
| 56 |
+
{{- '<tool_call>\n{"name": "' }}
|
| 57 |
+
{{- tool_call.name }}
|
| 58 |
+
{{- '", "arguments": ' }}
|
| 59 |
+
{%- if tool_call.arguments is string %}
|
| 60 |
+
{{- tool_call.arguments }}
|
| 61 |
+
{%- else %}
|
| 62 |
+
{{- tool_call.arguments | tojson }}
|
| 63 |
+
{%- endif %}
|
| 64 |
+
{{- '}\n</tool_call>' }}
|
| 65 |
+
{%- endfor %}
|
| 66 |
+
{%- endif %}
|
| 67 |
+
{{- '<|im_end|>\n' }}
|
| 68 |
+
{%- elif message.role == "tool" %}
|
| 69 |
+
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
|
| 70 |
+
{{- '<|im_start|>user' }}
|
| 71 |
+
{%- endif %}
|
| 72 |
+
{{- '\n<tool_response>\n' }}
|
| 73 |
+
{{- message.content }}
|
| 74 |
+
{{- '\n</tool_response>' }}
|
| 75 |
+
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
| 76 |
+
{{- '<|im_end|>\n' }}
|
| 77 |
+
{%- endif %}
|
| 78 |
+
{%- endif %}
|
| 79 |
+
{%- endfor %}
|
| 80 |
+
{%- if add_generation_prompt %}
|
| 81 |
+
{{- '<|im_start|>assistant\n' }}
|
| 82 |
+
{%- if enable_thinking is defined and enable_thinking is false %}
|
| 83 |
+
{{- '<think>\n\n</think>\n\n' }}
|
| 84 |
+
{%- endif %}
|
| 85 |
+
{%- endif %}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/tokenizer_config.json
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"backend": "tokenizers",
|
| 4 |
+
"bos_token": null,
|
| 5 |
+
"clean_up_tokenization_spaces": false,
|
| 6 |
+
"eos_token": "<|endoftext|>",
|
| 7 |
+
"errors": "replace",
|
| 8 |
+
"extra_special_tokens": [
|
| 9 |
+
"<|im_start|>",
|
| 10 |
+
"<|im_end|>",
|
| 11 |
+
"<|object_ref_start|>",
|
| 12 |
+
"<|object_ref_end|>",
|
| 13 |
+
"<|box_start|>",
|
| 14 |
+
"<|box_end|>",
|
| 15 |
+
"<|quad_start|>",
|
| 16 |
+
"<|quad_end|>",
|
| 17 |
+
"<|vision_start|>",
|
| 18 |
+
"<|vision_end|>",
|
| 19 |
+
"<|vision_pad|>",
|
| 20 |
+
"<|image_pad|>",
|
| 21 |
+
"<|video_pad|>"
|
| 22 |
+
],
|
| 23 |
+
"is_local": false,
|
| 24 |
+
"model_max_length": 131072,
|
| 25 |
+
"pad_token": "<|endoftext|>",
|
| 26 |
+
"split_special_tokens": false,
|
| 27 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 28 |
+
"unk_token": null
|
| 29 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-2723/trainer_state.json
ADDED
|
@@ -0,0 +1,651 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"best_global_step": null,
|
| 3 |
+
"best_metric": null,
|
| 4 |
+
"best_model_checkpoint": null,
|
| 5 |
+
"epoch": 7.0,
|
| 6 |
+
"eval_steps": 500,
|
| 7 |
+
"global_step": 2723,
|
| 8 |
+
"is_hyper_param_search": false,
|
| 9 |
+
"is_local_process_zero": true,
|
| 10 |
+
"is_world_process_zero": true,
|
| 11 |
+
"log_history": [
|
| 12 |
+
{
|
| 13 |
+
"entropy": 1.6110110306739807,
|
| 14 |
+
"epoch": 0.1287001287001287,
|
| 15 |
+
"grad_norm": 0.8611119389533997,
|
| 16 |
+
"learning_rate": 3.4130257962133866e-05,
|
| 17 |
+
"loss": 1.53956298828125,
|
| 18 |
+
"mean_token_accuracy": 0.6652342769503593,
|
| 19 |
+
"num_tokens": 73407.0,
|
| 20 |
+
"step": 50
|
| 21 |
+
},
|
| 22 |
+
{
|
| 23 |
+
"entropy": 0.8662820833921433,
|
| 24 |
+
"epoch": 0.2574002574002574,
|
| 25 |
+
"grad_norm": 0.7405035495758057,
|
| 26 |
+
"learning_rate": 6.895705180104598e-05,
|
| 27 |
+
"loss": 0.7929539489746094,
|
| 28 |
+
"mean_token_accuracy": 0.7786347842216492,
|
| 29 |
+
"num_tokens": 143994.0,
|
| 30 |
+
"step": 100
|
| 31 |
+
},
|
| 32 |
+
{
|
| 33 |
+
"entropy": 0.7912819278240204,
|
| 34 |
+
"epoch": 0.3861003861003861,
|
| 35 |
+
"grad_norm": 0.6105485558509827,
|
| 36 |
+
"learning_rate": 0.00010378384563995809,
|
| 37 |
+
"loss": 0.7236511993408203,
|
| 38 |
+
"mean_token_accuracy": 0.7925467795133591,
|
| 39 |
+
"num_tokens": 216171.0,
|
| 40 |
+
"step": 150
|
| 41 |
+
},
|
| 42 |
+
{
|
| 43 |
+
"entropy": 0.7777709531784057,
|
| 44 |
+
"epoch": 0.5148005148005148,
|
| 45 |
+
"grad_norm": 0.5781793594360352,
|
| 46 |
+
"learning_rate": 0.0001386106394788702,
|
| 47 |
+
"loss": 0.7024919891357422,
|
| 48 |
+
"mean_token_accuracy": 0.7969876372814179,
|
| 49 |
+
"num_tokens": 284702.0,
|
| 50 |
+
"step": 200
|
| 51 |
+
},
|
| 52 |
+
{
|
| 53 |
+
"entropy": 0.7632609683275223,
|
| 54 |
+
"epoch": 0.6435006435006435,
|
| 55 |
+
"grad_norm": 0.6553444862365723,
|
| 56 |
+
"learning_rate": 0.00017343743331778232,
|
| 57 |
+
"loss": 0.7000718688964844,
|
| 58 |
+
"mean_token_accuracy": 0.7994742071628571,
|
| 59 |
+
"num_tokens": 356393.0,
|
| 60 |
+
"step": 250
|
| 61 |
+
},
|
| 62 |
+
{
|
| 63 |
+
"entropy": 0.754165632724762,
|
| 64 |
+
"epoch": 0.7722007722007722,
|
| 65 |
+
"grad_norm": 0.634222149848938,
|
| 66 |
+
"learning_rate": 0.00020826422715669444,
|
| 67 |
+
"loss": 0.6909049987792969,
|
| 68 |
+
"mean_token_accuracy": 0.8016809666156769,
|
| 69 |
+
"num_tokens": 426916.0,
|
| 70 |
+
"step": 300
|
| 71 |
+
},
|
| 72 |
+
{
|
| 73 |
+
"entropy": 0.7530711203813553,
|
| 74 |
+
"epoch": 0.9009009009009009,
|
| 75 |
+
"grad_norm": 0.7020156383514404,
|
| 76 |
+
"learning_rate": 0.00024309102099560653,
|
| 77 |
+
"loss": 0.69554931640625,
|
| 78 |
+
"mean_token_accuracy": 0.8002807641029358,
|
| 79 |
+
"num_tokens": 499603.0,
|
| 80 |
+
"step": 350
|
| 81 |
+
},
|
| 82 |
+
{
|
| 83 |
+
"epoch": 1.0,
|
| 84 |
+
"eval_entropy": 0.5728044164242204,
|
| 85 |
+
"eval_loss": 0.6785851716995239,
|
| 86 |
+
"eval_mean_token_accuracy": 0.798433246993527,
|
| 87 |
+
"eval_num_tokens": 555493.0,
|
| 88 |
+
"eval_runtime": 162.4129,
|
| 89 |
+
"eval_samples_per_second": 9.513,
|
| 90 |
+
"eval_steps_per_second": 1.194,
|
| 91 |
+
"step": 389
|
| 92 |
+
},
|
| 93 |
+
{
|
| 94 |
+
"entropy": 0.7489313946829902,
|
| 95 |
+
"epoch": 1.0283140283140284,
|
| 96 |
+
"grad_norm": 0.7532815933227539,
|
| 97 |
+
"learning_rate": 0.0002709470016827303,
|
| 98 |
+
"loss": 0.6856581878662109,
|
| 99 |
+
"mean_token_accuracy": 0.8023027079273956,
|
| 100 |
+
"num_tokens": 570813.0,
|
| 101 |
+
"step": 400
|
| 102 |
+
},
|
| 103 |
+
{
|
| 104 |
+
"entropy": 0.7238409864902496,
|
| 105 |
+
"epoch": 1.157014157014157,
|
| 106 |
+
"grad_norm": 1.2016215324401855,
|
| 107 |
+
"learning_rate": 0.0002707561443541359,
|
| 108 |
+
"loss": 0.6699818420410156,
|
| 109 |
+
"mean_token_accuracy": 0.8070302194356919,
|
| 110 |
+
"num_tokens": 642956.0,
|
| 111 |
+
"step": 450
|
| 112 |
+
},
|
| 113 |
+
{
|
| 114 |
+
"entropy": 0.7388938587903976,
|
| 115 |
+
"epoch": 1.2857142857142856,
|
| 116 |
+
"grad_norm": 0.7279272079467773,
|
| 117 |
+
"learning_rate": 0.0002702930068622498,
|
| 118 |
+
"loss": 0.6728517150878907,
|
| 119 |
+
"mean_token_accuracy": 0.8049580943584442,
|
| 120 |
+
"num_tokens": 714498.0,
|
| 121 |
+
"step": 500
|
| 122 |
+
},
|
| 123 |
+
{
|
| 124 |
+
"entropy": 0.7284485149383545,
|
| 125 |
+
"epoch": 1.4144144144144144,
|
| 126 |
+
"grad_norm": 0.8563987016677856,
|
| 127 |
+
"learning_rate": 0.0002695585213716931,
|
| 128 |
+
"loss": 0.6657986450195312,
|
| 129 |
+
"mean_token_accuracy": 0.8085588800907135,
|
| 130 |
+
"num_tokens": 785595.0,
|
| 131 |
+
"step": 550
|
| 132 |
+
},
|
| 133 |
+
{
|
| 134 |
+
"entropy": 0.6960959500074386,
|
| 135 |
+
"epoch": 1.5431145431145432,
|
| 136 |
+
"grad_norm": 0.5145474672317505,
|
| 137 |
+
"learning_rate": 0.0002685541661937683,
|
| 138 |
+
"loss": 0.6358638763427734,
|
| 139 |
+
"mean_token_accuracy": 0.8131621342897415,
|
| 140 |
+
"num_tokens": 857551.0,
|
| 141 |
+
"step": 600
|
| 142 |
+
},
|
| 143 |
+
{
|
| 144 |
+
"entropy": 0.7220414417982102,
|
| 145 |
+
"epoch": 1.6718146718146718,
|
| 146 |
+
"grad_norm": 0.9381059408187866,
|
| 147 |
+
"learning_rate": 0.00026728196281103746,
|
| 148 |
+
"loss": 0.6531407928466797,
|
| 149 |
+
"mean_token_accuracy": 0.811619822382927,
|
| 150 |
+
"num_tokens": 926864.0,
|
| 151 |
+
"step": 650
|
| 152 |
+
},
|
| 153 |
+
{
|
| 154 |
+
"entropy": 0.6955459499359131,
|
| 155 |
+
"epoch": 1.8005148005148004,
|
| 156 |
+
"grad_norm": 0.6115108728408813,
|
| 157 |
+
"learning_rate": 0.0002657444718086503,
|
| 158 |
+
"loss": 0.6269588088989257,
|
| 159 |
+
"mean_token_accuracy": 0.8155373805761337,
|
| 160 |
+
"num_tokens": 998824.0,
|
| 161 |
+
"step": 700
|
| 162 |
+
},
|
| 163 |
+
{
|
| 164 |
+
"entropy": 0.6929812705516816,
|
| 165 |
+
"epoch": 1.9292149292149292,
|
| 166 |
+
"grad_norm": 0.749489426612854,
|
| 167 |
+
"learning_rate": 0.0002639447877206115,
|
| 168 |
+
"loss": 0.629054069519043,
|
| 169 |
+
"mean_token_accuracy": 0.8183021235466004,
|
| 170 |
+
"num_tokens": 1069332.0,
|
| 171 |
+
"step": 750
|
| 172 |
+
},
|
| 173 |
+
{
|
| 174 |
+
"epoch": 2.0,
|
| 175 |
+
"eval_entropy": 0.5833694102223387,
|
| 176 |
+
"eval_loss": 0.6328718662261963,
|
| 177 |
+
"eval_mean_token_accuracy": 0.8159044071571114,
|
| 178 |
+
"eval_num_tokens": 1110986.0,
|
| 179 |
+
"eval_runtime": 161.885,
|
| 180 |
+
"eval_samples_per_second": 9.544,
|
| 181 |
+
"eval_steps_per_second": 1.198,
|
| 182 |
+
"step": 778
|
| 183 |
+
},
|
| 184 |
+
{
|
| 185 |
+
"entropy": 0.647241060480927,
|
| 186 |
+
"epoch": 2.056628056628057,
|
| 187 |
+
"grad_norm": 0.7070767879486084,
|
| 188 |
+
"learning_rate": 0.00026188653280135975,
|
| 189 |
+
"loss": 0.5823922348022461,
|
| 190 |
+
"mean_token_accuracy": 0.8265365301960647,
|
| 191 |
+
"num_tokens": 1141195.0,
|
| 192 |
+
"step": 800
|
| 193 |
+
},
|
| 194 |
+
{
|
| 195 |
+
"entropy": 0.5995378407835961,
|
| 196 |
+
"epoch": 2.1853281853281854,
|
| 197 |
+
"grad_norm": 0.8090486526489258,
|
| 198 |
+
"learning_rate": 0.0002595738497351955,
|
| 199 |
+
"loss": 0.5325597763061524,
|
| 200 |
+
"mean_token_accuracy": 0.8369336777925491,
|
| 201 |
+
"num_tokens": 1210708.0,
|
| 202 |
+
"step": 850
|
| 203 |
+
},
|
| 204 |
+
{
|
| 205 |
+
"entropy": 0.6043145382404327,
|
| 206 |
+
"epoch": 2.314028314028314,
|
| 207 |
+
"grad_norm": 0.8279913067817688,
|
| 208 |
+
"learning_rate": 0.00025701139329823054,
|
| 209 |
+
"loss": 0.5414446258544922,
|
| 210 |
+
"mean_token_accuracy": 0.8361396533250809,
|
| 211 |
+
"num_tokens": 1283441.0,
|
| 212 |
+
"step": 900
|
| 213 |
+
},
|
| 214 |
+
{
|
| 215 |
+
"entropy": 0.5953224584460258,
|
| 216 |
+
"epoch": 2.4427284427284426,
|
| 217 |
+
"grad_norm": 0.6075023412704468,
|
| 218 |
+
"learning_rate": 0.00025420432098964183,
|
| 219 |
+
"loss": 0.536654167175293,
|
| 220 |
+
"mean_token_accuracy": 0.8340093129873276,
|
| 221 |
+
"num_tokens": 1356418.0,
|
| 222 |
+
"step": 950
|
| 223 |
+
},
|
| 224 |
+
{
|
| 225 |
+
"entropy": 0.5998479858040809,
|
| 226 |
+
"epoch": 2.571428571428571,
|
| 227 |
+
"grad_norm": 1.0311471223831177,
|
| 228 |
+
"learning_rate": 0.0002511582826510862,
|
| 229 |
+
"loss": 0.5372924423217773,
|
| 230 |
+
"mean_token_accuracy": 0.8366045409440994,
|
| 231 |
+
"num_tokens": 1427797.0,
|
| 232 |
+
"step": 1000
|
| 233 |
+
},
|
| 234 |
+
{
|
| 235 |
+
"entropy": 0.5980158120393753,
|
| 236 |
+
"epoch": 2.7001287001287,
|
| 237 |
+
"grad_norm": 0.5971426367759705,
|
| 238 |
+
"learning_rate": 0.0002478794090951689,
|
| 239 |
+
"loss": 0.5392885208129883,
|
| 240 |
+
"mean_token_accuracy": 0.8347727072238922,
|
| 241 |
+
"num_tokens": 1498082.0,
|
| 242 |
+
"step": 1050
|
| 243 |
+
},
|
| 244 |
+
{
|
| 245 |
+
"entropy": 0.5990680930018425,
|
| 246 |
+
"epoch": 2.828828828828829,
|
| 247 |
+
"grad_norm": 0.5662627220153809,
|
| 248 |
+
"learning_rate": 0.0002443742997658538,
|
| 249 |
+
"loss": 0.5360498428344727,
|
| 250 |
+
"mean_token_accuracy": 0.8371847170591354,
|
| 251 |
+
"num_tokens": 1568798.0,
|
| 252 |
+
"step": 1100
|
| 253 |
+
},
|
| 254 |
+
{
|
| 255 |
+
"entropy": 0.5957039377093315,
|
| 256 |
+
"epoch": 2.9575289575289574,
|
| 257 |
+
"grad_norm": 0.5043798685073853,
|
| 258 |
+
"learning_rate": 0.00024065000945565205,
|
| 259 |
+
"loss": 0.5342231369018555,
|
| 260 |
+
"mean_token_accuracy": 0.8380735236406326,
|
| 261 |
+
"num_tokens": 1643449.0,
|
| 262 |
+
"step": 1150
|
| 263 |
+
},
|
| 264 |
+
{
|
| 265 |
+
"epoch": 3.0,
|
| 266 |
+
"eval_entropy": 0.5346978819861854,
|
| 267 |
+
"eval_loss": 0.6255015134811401,
|
| 268 |
+
"eval_mean_token_accuracy": 0.8202014476368108,
|
| 269 |
+
"eval_num_tokens": 1666479.0,
|
| 270 |
+
"eval_runtime": 161.6098,
|
| 271 |
+
"eval_samples_per_second": 9.56,
|
| 272 |
+
"eval_steps_per_second": 1.2,
|
| 273 |
+
"step": 1167
|
| 274 |
+
},
|
| 275 |
+
{
|
| 276 |
+
"entropy": 0.5212209137401196,
|
| 277 |
+
"epoch": 3.0849420849420848,
|
| 278 |
+
"grad_norm": 0.7095440626144409,
|
| 279 |
+
"learning_rate": 0.00023671403410632178,
|
| 280 |
+
"loss": 0.45311901092529294,
|
| 281 |
+
"mean_token_accuracy": 0.856362871449403,
|
| 282 |
+
"num_tokens": 1713536.0,
|
| 283 |
+
"step": 1200
|
| 284 |
+
},
|
| 285 |
+
{
|
| 286 |
+
"entropy": 0.4701593083143234,
|
| 287 |
+
"epoch": 3.213642213642214,
|
| 288 |
+
"grad_norm": 0.6408083438873291,
|
| 289 |
+
"learning_rate": 0.0002325742957216607,
|
| 290 |
+
"loss": 0.39916397094726563,
|
| 291 |
+
"mean_token_accuracy": 0.8698061722517013,
|
| 292 |
+
"num_tokens": 1785609.0,
|
| 293 |
+
"step": 1250
|
| 294 |
+
},
|
| 295 |
+
{
|
| 296 |
+
"entropy": 0.4766591975092888,
|
| 297 |
+
"epoch": 3.3423423423423424,
|
| 298 |
+
"grad_norm": 0.6415093541145325,
|
| 299 |
+
"learning_rate": 0.0002282391264227552,
|
| 300 |
+
"loss": 0.4116698455810547,
|
| 301 |
+
"mean_token_accuracy": 0.8679435575008392,
|
| 302 |
+
"num_tokens": 1858651.0,
|
| 303 |
+
"step": 1300
|
| 304 |
+
},
|
| 305 |
+
{
|
| 306 |
+
"entropy": 0.4937947469949722,
|
| 307 |
+
"epoch": 3.471042471042471,
|
| 308 |
+
"grad_norm": 0.6549825072288513,
|
| 309 |
+
"learning_rate": 0.00022371725167778054,
|
| 310 |
+
"loss": 0.4296376037597656,
|
| 311 |
+
"mean_token_accuracy": 0.8609357953071595,
|
| 312 |
+
"num_tokens": 1928692.0,
|
| 313 |
+
"step": 1350
|
| 314 |
+
},
|
| 315 |
+
{
|
| 316 |
+
"entropy": 0.4920153194665909,
|
| 317 |
+
"epoch": 3.5997425997425996,
|
| 318 |
+
"grad_norm": 0.6452126502990723,
|
| 319 |
+
"learning_rate": 0.00021901777274010406,
|
| 320 |
+
"loss": 0.4307489013671875,
|
| 321 |
+
"mean_token_accuracy": 0.8606827831268311,
|
| 322 |
+
"num_tokens": 1998668.0,
|
| 323 |
+
"step": 1400
|
| 324 |
+
},
|
| 325 |
+
{
|
| 326 |
+
"entropy": 0.490042342543602,
|
| 327 |
+
"epoch": 3.7284427284427286,
|
| 328 |
+
"grad_norm": 0.5727734565734863,
|
| 329 |
+
"learning_rate": 0.0002141501483300395,
|
| 330 |
+
"loss": 0.4295254135131836,
|
| 331 |
+
"mean_token_accuracy": 0.8616545403003693,
|
| 332 |
+
"num_tokens": 2072809.0,
|
| 333 |
+
"step": 1450
|
| 334 |
+
},
|
| 335 |
+
{
|
| 336 |
+
"entropy": 0.49857193052768706,
|
| 337 |
+
"epoch": 3.857142857142857,
|
| 338 |
+
"grad_norm": 0.7732954025268555,
|
| 339 |
+
"learning_rate": 0.00020912417559712133,
|
| 340 |
+
"loss": 0.4289303207397461,
|
| 341 |
+
"mean_token_accuracy": 0.8616443765163422,
|
| 342 |
+
"num_tokens": 2142475.0,
|
| 343 |
+
"step": 1500
|
| 344 |
+
},
|
| 345 |
+
{
|
| 346 |
+
"entropy": 0.4746784272789955,
|
| 347 |
+
"epoch": 3.985842985842986,
|
| 348 |
+
"grad_norm": 0.5791187882423401,
|
| 349 |
+
"learning_rate": 0.00020394997040121726,
|
| 350 |
+
"loss": 0.4180263900756836,
|
| 351 |
+
"mean_token_accuracy": 0.866080379486084,
|
| 352 |
+
"num_tokens": 2214044.0,
|
| 353 |
+
"step": 1550
|
| 354 |
+
},
|
| 355 |
+
{
|
| 356 |
+
"epoch": 4.0,
|
| 357 |
+
"eval_entropy": 0.4793926059585257,
|
| 358 |
+
"eval_loss": 0.627627968788147,
|
| 359 |
+
"eval_mean_token_accuracy": 0.8233870095813397,
|
| 360 |
+
"eval_num_tokens": 2221972.0,
|
| 361 |
+
"eval_runtime": 162.0225,
|
| 362 |
+
"eval_samples_per_second": 9.536,
|
| 363 |
+
"eval_steps_per_second": 1.197,
|
| 364 |
+
"step": 1556
|
| 365 |
+
},
|
| 366 |
+
{
|
| 367 |
+
"entropy": 0.38195489000792454,
|
| 368 |
+
"epoch": 4.113256113256114,
|
| 369 |
+
"grad_norm": 0.6002617478370667,
|
| 370 |
+
"learning_rate": 0.0001986379469521669,
|
| 371 |
+
"loss": 0.30819049835205076,
|
| 372 |
+
"mean_token_accuracy": 0.8977848634575353,
|
| 373 |
+
"num_tokens": 2282164.0,
|
| 374 |
+
"step": 1600
|
| 375 |
+
},
|
| 376 |
+
{
|
| 377 |
+
"entropy": 0.3655787402391434,
|
| 378 |
+
"epoch": 4.241956241956242,
|
| 379 |
+
"grad_norm": 0.7100041508674622,
|
| 380 |
+
"learning_rate": 0.00019319879684892634,
|
| 381 |
+
"loss": 0.29959835052490236,
|
| 382 |
+
"mean_token_accuracy": 0.8991208010911942,
|
| 383 |
+
"num_tokens": 2353213.0,
|
| 384 |
+
"step": 1650
|
| 385 |
+
},
|
| 386 |
+
{
|
| 387 |
+
"entropy": 0.3821141055226326,
|
| 388 |
+
"epoch": 4.370656370656371,
|
| 389 |
+
"grad_norm": 0.5848307013511658,
|
| 390 |
+
"learning_rate": 0.00018764346756040715,
|
| 391 |
+
"loss": 0.313802490234375,
|
| 392 |
+
"mean_token_accuracy": 0.895167955160141,
|
| 393 |
+
"num_tokens": 2425068.0,
|
| 394 |
+
"step": 1700
|
| 395 |
+
},
|
| 396 |
+
{
|
| 397 |
+
"entropy": 0.37083797007799146,
|
| 398 |
+
"epoch": 4.499356499356499,
|
| 399 |
+
"grad_norm": 0.6447024941444397,
|
| 400 |
+
"learning_rate": 0.00018198314039132143,
|
| 401 |
+
"loss": 0.30583988189697264,
|
| 402 |
+
"mean_token_accuracy": 0.8961733293533325,
|
| 403 |
+
"num_tokens": 2498321.0,
|
| 404 |
+
"step": 1750
|
| 405 |
+
},
|
| 406 |
+
{
|
| 407 |
+
"entropy": 0.3791545969247818,
|
| 408 |
+
"epoch": 4.628056628056628,
|
| 409 |
+
"grad_norm": 0.6575382351875305,
|
| 410 |
+
"learning_rate": 0.00017622920797738184,
|
| 411 |
+
"loss": 0.3088031005859375,
|
| 412 |
+
"mean_token_accuracy": 0.8960050916671753,
|
| 413 |
+
"num_tokens": 2570321.0,
|
| 414 |
+
"step": 1800
|
| 415 |
+
},
|
| 416 |
+
{
|
| 417 |
+
"entropy": 0.3946831756830215,
|
| 418 |
+
"epoch": 4.756756756756757,
|
| 419 |
+
"grad_norm": 0.5351552963256836,
|
| 420 |
+
"learning_rate": 0.00017039325135515207,
|
| 421 |
+
"loss": 0.3229162979125977,
|
| 422 |
+
"mean_token_accuracy": 0.8920552498102188,
|
| 423 |
+
"num_tokens": 2642851.0,
|
| 424 |
+
"step": 1850
|
| 425 |
+
},
|
| 426 |
+
{
|
| 427 |
+
"entropy": 0.37298239797353744,
|
| 428 |
+
"epoch": 4.885456885456885,
|
| 429 |
+
"grad_norm": 0.7624587416648865,
|
| 430 |
+
"learning_rate": 0.00016448701665269964,
|
| 431 |
+
"loss": 0.3067934799194336,
|
| 432 |
+
"mean_token_accuracy": 0.8951873427629471,
|
| 433 |
+
"num_tokens": 2715629.0,
|
| 434 |
+
"step": 1900
|
| 435 |
+
},
|
| 436 |
+
{
|
| 437 |
+
"epoch": 5.0,
|
| 438 |
+
"eval_entropy": 0.3738150204887095,
|
| 439 |
+
"eval_loss": 0.7131896615028381,
|
| 440 |
+
"eval_mean_token_accuracy": 0.8207752468045225,
|
| 441 |
+
"eval_num_tokens": 2777465.0,
|
| 442 |
+
"eval_runtime": 162.1244,
|
| 443 |
+
"eval_samples_per_second": 9.53,
|
| 444 |
+
"eval_steps_per_second": 1.197,
|
| 445 |
+
"step": 1945
|
| 446 |
+
},
|
| 447 |
+
{
|
| 448 |
+
"entropy": 0.377820266617669,
|
| 449 |
+
"epoch": 5.012870012870013,
|
| 450 |
+
"grad_norm": 0.41990190744400024,
|
| 451 |
+
"learning_rate": 0.00015852239144796624,
|
| 452 |
+
"loss": 0.3058685111999512,
|
| 453 |
+
"mean_token_accuracy": 0.8964343480389527,
|
| 454 |
+
"num_tokens": 2784896.0,
|
| 455 |
+
"step": 1950
|
| 456 |
+
},
|
| 457 |
+
{
|
| 458 |
+
"entropy": 0.2714502356946468,
|
| 459 |
+
"epoch": 5.141570141570142,
|
| 460 |
+
"grad_norm": 0.422568678855896,
|
| 461 |
+
"learning_rate": 0.00015251138084243995,
|
| 462 |
+
"loss": 0.2093442153930664,
|
| 463 |
+
"mean_token_accuracy": 0.9311346983909607,
|
| 464 |
+
"num_tokens": 2854374.0,
|
| 465 |
+
"step": 2000
|
| 466 |
+
},
|
| 467 |
+
{
|
| 468 |
+
"entropy": 0.268475965410471,
|
| 469 |
+
"epoch": 5.27027027027027,
|
| 470 |
+
"grad_norm": 0.6637414693832397,
|
| 471 |
+
"learning_rate": 0.0001464660832982852,
|
| 472 |
+
"loss": 0.20736080169677734,
|
| 473 |
+
"mean_token_accuracy": 0.9289199805259705,
|
| 474 |
+
"num_tokens": 2927362.0,
|
| 475 |
+
"step": 2050
|
| 476 |
+
},
|
| 477 |
+
{
|
| 478 |
+
"entropy": 0.2644876340031624,
|
| 479 |
+
"epoch": 5.398970398970399,
|
| 480 |
+
"grad_norm": 0.47317707538604736,
|
| 481 |
+
"learning_rate": 0.00014039866628756467,
|
| 482 |
+
"loss": 0.20464908599853515,
|
| 483 |
+
"mean_token_accuracy": 0.9300856202840805,
|
| 484 |
+
"num_tokens": 3000143.0,
|
| 485 |
+
"step": 2100
|
| 486 |
+
},
|
| 487 |
+
{
|
| 488 |
+
"entropy": 0.2675253136456013,
|
| 489 |
+
"epoch": 5.527670527670527,
|
| 490 |
+
"grad_norm": 0.5253982543945312,
|
| 491 |
+
"learning_rate": 0.00013432134180256338,
|
| 492 |
+
"loss": 0.21154335021972656,
|
| 493 |
+
"mean_token_accuracy": 0.9283734840154648,
|
| 494 |
+
"num_tokens": 3072561.0,
|
| 495 |
+
"step": 2150
|
| 496 |
+
},
|
| 497 |
+
{
|
| 498 |
+
"entropy": 0.27213907435536383,
|
| 499 |
+
"epoch": 5.656370656370656,
|
| 500 |
+
"grad_norm": 0.46738553047180176,
|
| 501 |
+
"learning_rate": 0.00012824634177650664,
|
| 502 |
+
"loss": 0.21339216232299804,
|
| 503 |
+
"mean_token_accuracy": 0.9272083270549775,
|
| 504 |
+
"num_tokens": 3144831.0,
|
| 505 |
+
"step": 2200
|
| 506 |
+
},
|
| 507 |
+
{
|
| 508 |
+
"entropy": 0.2785488124191761,
|
| 509 |
+
"epoch": 5.785070785070785,
|
| 510 |
+
"grad_norm": 0.4469502866268158,
|
| 511 |
+
"learning_rate": 0.00012218589346414205,
|
| 512 |
+
"loss": 0.21601097106933595,
|
| 513 |
+
"mean_token_accuracy": 0.9255663657188415,
|
| 514 |
+
"num_tokens": 3215960.0,
|
| 515 |
+
"step": 2250
|
| 516 |
+
},
|
| 517 |
+
{
|
| 518 |
+
"entropy": 0.2699935150146484,
|
| 519 |
+
"epoch": 5.913770913770914,
|
| 520 |
+
"grad_norm": 0.7359778881072998,
|
| 521 |
+
"learning_rate": 0.00011615219483173828,
|
| 522 |
+
"loss": 0.20725584030151367,
|
| 523 |
+
"mean_token_accuracy": 0.9286630594730377,
|
| 524 |
+
"num_tokens": 3287499.0,
|
| 525 |
+
"step": 2300
|
| 526 |
+
},
|
| 527 |
+
{
|
| 528 |
+
"epoch": 6.0,
|
| 529 |
+
"eval_entropy": 0.26417383682174783,
|
| 530 |
+
"eval_loss": 0.8880229592323303,
|
| 531 |
+
"eval_mean_token_accuracy": 0.8159987201395723,
|
| 532 |
+
"eval_num_tokens": 3332958.0,
|
| 533 |
+
"eval_runtime": 162.0991,
|
| 534 |
+
"eval_samples_per_second": 9.531,
|
| 535 |
+
"eval_steps_per_second": 1.197,
|
| 536 |
+
"step": 2334
|
| 537 |
+
},
|
| 538 |
+
{
|
| 539 |
+
"entropy": 0.24814540704693458,
|
| 540 |
+
"epoch": 6.041184041184041,
|
| 541 |
+
"grad_norm": 0.4953760802745819,
|
| 542 |
+
"learning_rate": 0.00011015739000603316,
|
| 543 |
+
"loss": 0.18749794006347656,
|
| 544 |
+
"mean_token_accuracy": 0.9370789509831052,
|
| 545 |
+
"num_tokens": 3356879.0,
|
| 546 |
+
"step": 2350
|
| 547 |
+
},
|
| 548 |
+
{
|
| 549 |
+
"entropy": 0.19976271741092205,
|
| 550 |
+
"epoch": 6.1698841698841695,
|
| 551 |
+
"grad_norm": 0.4834803342819214,
|
| 552 |
+
"learning_rate": 0.00010421354483154553,
|
| 553 |
+
"loss": 0.14283526420593262,
|
| 554 |
+
"mean_token_accuracy": 0.9521516615152359,
|
| 555 |
+
"num_tokens": 3427587.0,
|
| 556 |
+
"step": 2400
|
| 557 |
+
},
|
| 558 |
+
{
|
| 559 |
+
"entropy": 0.2060488449037075,
|
| 560 |
+
"epoch": 6.298584298584299,
|
| 561 |
+
"grad_norm": 0.4888673722743988,
|
| 562 |
+
"learning_rate": 9.8332622585447e-05,
|
| 563 |
+
"loss": 0.14414511680603026,
|
| 564 |
+
"mean_token_accuracy": 0.9510996866226197,
|
| 565 |
+
"num_tokens": 3498688.0,
|
| 566 |
+
"step": 2450
|
| 567 |
+
},
|
| 568 |
+
{
|
| 569 |
+
"entropy": 0.2059111550450325,
|
| 570 |
+
"epoch": 6.427284427284428,
|
| 571 |
+
"grad_norm": 0.4064404368400574,
|
| 572 |
+
"learning_rate": 9.252645989887253e-05,
|
| 573 |
+
"loss": 0.14820143699645996,
|
| 574 |
+
"mean_token_accuracy": 0.9507584601640702,
|
| 575 |
+
"num_tokens": 3566137.0,
|
| 576 |
+
"step": 2500
|
| 577 |
+
},
|
| 578 |
+
{
|
| 579 |
+
"entropy": 0.19700154662132263,
|
| 580 |
+
"epoch": 6.555984555984556,
|
| 581 |
+
"grad_norm": 0.467965304851532,
|
| 582 |
+
"learning_rate": 8.680674293313417e-05,
|
| 583 |
+
"loss": 0.14303470611572267,
|
| 584 |
+
"mean_token_accuracy": 0.9515972435474396,
|
| 585 |
+
"num_tokens": 3639573.0,
|
| 586 |
+
"step": 2550
|
| 587 |
+
},
|
| 588 |
+
{
|
| 589 |
+
"entropy": 0.20180423602461814,
|
| 590 |
+
"epoch": 6.684684684684685,
|
| 591 |
+
"grad_norm": 0.36836138367652893,
|
| 592 |
+
"learning_rate": 8.118498385878736e-05,
|
| 593 |
+
"loss": 0.14280882835388184,
|
| 594 |
+
"mean_token_accuracy": 0.9515993863344192,
|
| 595 |
+
"num_tokens": 3710433.0,
|
| 596 |
+
"step": 2600
|
| 597 |
+
},
|
| 598 |
+
{
|
| 599 |
+
"entropy": 0.20024395987391472,
|
| 600 |
+
"epoch": 6.813384813384813,
|
| 601 |
+
"grad_norm": 0.38375866413116455,
|
| 602 |
+
"learning_rate": 7.567249768489171e-05,
|
| 603 |
+
"loss": 0.1427844524383545,
|
| 604 |
+
"mean_token_accuracy": 0.9524166631698608,
|
| 605 |
+
"num_tokens": 3781550.0,
|
| 606 |
+
"step": 2650
|
| 607 |
+
},
|
| 608 |
+
{
|
| 609 |
+
"entropy": 0.19561587080359458,
|
| 610 |
+
"epoch": 6.942084942084942,
|
| 611 |
+
"grad_norm": 0.41185441613197327,
|
| 612 |
+
"learning_rate": 7.028037948510187e-05,
|
| 613 |
+
"loss": 0.13993803024291993,
|
| 614 |
+
"mean_token_accuracy": 0.9522478264570237,
|
| 615 |
+
"num_tokens": 3854952.0,
|
| 616 |
+
"step": 2700
|
| 617 |
+
},
|
| 618 |
+
{
|
| 619 |
+
"epoch": 7.0,
|
| 620 |
+
"eval_entropy": 0.19501976062034823,
|
| 621 |
+
"eval_loss": 1.0653952360153198,
|
| 622 |
+
"eval_mean_token_accuracy": 0.8205490803595671,
|
| 623 |
+
"eval_num_tokens": 3888451.0,
|
| 624 |
+
"eval_runtime": 161.8533,
|
| 625 |
+
"eval_samples_per_second": 9.546,
|
| 626 |
+
"eval_steps_per_second": 1.199,
|
| 627 |
+
"step": 2723
|
| 628 |
+
}
|
| 629 |
+
],
|
| 630 |
+
"logging_steps": 50,
|
| 631 |
+
"max_steps": 3890,
|
| 632 |
+
"num_input_tokens_seen": 0,
|
| 633 |
+
"num_train_epochs": 10,
|
| 634 |
+
"save_steps": 500,
|
| 635 |
+
"stateful_callbacks": {
|
| 636 |
+
"TrainerControl": {
|
| 637 |
+
"args": {
|
| 638 |
+
"should_epoch_stop": false,
|
| 639 |
+
"should_evaluate": false,
|
| 640 |
+
"should_log": false,
|
| 641 |
+
"should_save": true,
|
| 642 |
+
"should_training_stop": false
|
| 643 |
+
},
|
| 644 |
+
"attributes": {}
|
| 645 |
+
}
|
| 646 |
+
},
|
| 647 |
+
"total_flos": 6.516671077296845e+17,
|
| 648 |
+
"train_batch_size": 8,
|
| 649 |
+
"trial_name": null,
|
| 650 |
+
"trial_params": null
|
| 651 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/README.md
ADDED
|
@@ -0,0 +1,209 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen3-14B-Base
|
| 3 |
+
library_name: peft
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:Qwen/Qwen3-14B-Base
|
| 7 |
+
- lora
|
| 8 |
+
- sft
|
| 9 |
+
- transformers
|
| 10 |
+
- trl
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Model Card for Model ID
|
| 14 |
+
|
| 15 |
+
<!-- Provide a quick summary of what the model is/does. -->
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
## Model Details
|
| 20 |
+
|
| 21 |
+
### Model Description
|
| 22 |
+
|
| 23 |
+
<!-- Provide a longer summary of what this model is. -->
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
- **Developed by:** [More Information Needed]
|
| 28 |
+
- **Funded by [optional]:** [More Information Needed]
|
| 29 |
+
- **Shared by [optional]:** [More Information Needed]
|
| 30 |
+
- **Model type:** [More Information Needed]
|
| 31 |
+
- **Language(s) (NLP):** [More Information Needed]
|
| 32 |
+
- **License:** [More Information Needed]
|
| 33 |
+
- **Finetuned from model [optional]:** [More Information Needed]
|
| 34 |
+
|
| 35 |
+
### Model Sources [optional]
|
| 36 |
+
|
| 37 |
+
<!-- Provide the basic links for the model. -->
|
| 38 |
+
|
| 39 |
+
- **Repository:** [More Information Needed]
|
| 40 |
+
- **Paper [optional]:** [More Information Needed]
|
| 41 |
+
- **Demo [optional]:** [More Information Needed]
|
| 42 |
+
|
| 43 |
+
## Uses
|
| 44 |
+
|
| 45 |
+
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
| 46 |
+
|
| 47 |
+
### Direct Use
|
| 48 |
+
|
| 49 |
+
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
| 50 |
+
|
| 51 |
+
[More Information Needed]
|
| 52 |
+
|
| 53 |
+
### Downstream Use [optional]
|
| 54 |
+
|
| 55 |
+
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
| 56 |
+
|
| 57 |
+
[More Information Needed]
|
| 58 |
+
|
| 59 |
+
### Out-of-Scope Use
|
| 60 |
+
|
| 61 |
+
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
| 62 |
+
|
| 63 |
+
[More Information Needed]
|
| 64 |
+
|
| 65 |
+
## Bias, Risks, and Limitations
|
| 66 |
+
|
| 67 |
+
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
| 68 |
+
|
| 69 |
+
[More Information Needed]
|
| 70 |
+
|
| 71 |
+
### Recommendations
|
| 72 |
+
|
| 73 |
+
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
| 74 |
+
|
| 75 |
+
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
| 76 |
+
|
| 77 |
+
## How to Get Started with the Model
|
| 78 |
+
|
| 79 |
+
Use the code below to get started with the model.
|
| 80 |
+
|
| 81 |
+
[More Information Needed]
|
| 82 |
+
|
| 83 |
+
## Training Details
|
| 84 |
+
|
| 85 |
+
### Training Data
|
| 86 |
+
|
| 87 |
+
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
| 88 |
+
|
| 89 |
+
[More Information Needed]
|
| 90 |
+
|
| 91 |
+
### Training Procedure
|
| 92 |
+
|
| 93 |
+
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
| 94 |
+
|
| 95 |
+
#### Preprocessing [optional]
|
| 96 |
+
|
| 97 |
+
[More Information Needed]
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
#### Training Hyperparameters
|
| 101 |
+
|
| 102 |
+
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
| 103 |
+
|
| 104 |
+
#### Speeds, Sizes, Times [optional]
|
| 105 |
+
|
| 106 |
+
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
| 107 |
+
|
| 108 |
+
[More Information Needed]
|
| 109 |
+
|
| 110 |
+
## Evaluation
|
| 111 |
+
|
| 112 |
+
<!-- This section describes the evaluation protocols and provides the results. -->
|
| 113 |
+
|
| 114 |
+
### Testing Data, Factors & Metrics
|
| 115 |
+
|
| 116 |
+
#### Testing Data
|
| 117 |
+
|
| 118 |
+
<!-- This should link to a Dataset Card if possible. -->
|
| 119 |
+
|
| 120 |
+
[More Information Needed]
|
| 121 |
+
|
| 122 |
+
#### Factors
|
| 123 |
+
|
| 124 |
+
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
| 125 |
+
|
| 126 |
+
[More Information Needed]
|
| 127 |
+
|
| 128 |
+
#### Metrics
|
| 129 |
+
|
| 130 |
+
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
| 131 |
+
|
| 132 |
+
[More Information Needed]
|
| 133 |
+
|
| 134 |
+
### Results
|
| 135 |
+
|
| 136 |
+
[More Information Needed]
|
| 137 |
+
|
| 138 |
+
#### Summary
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
## Model Examination [optional]
|
| 143 |
+
|
| 144 |
+
<!-- Relevant interpretability work for the model goes here -->
|
| 145 |
+
|
| 146 |
+
[More Information Needed]
|
| 147 |
+
|
| 148 |
+
## Environmental Impact
|
| 149 |
+
|
| 150 |
+
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
| 151 |
+
|
| 152 |
+
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
| 153 |
+
|
| 154 |
+
- **Hardware Type:** [More Information Needed]
|
| 155 |
+
- **Hours used:** [More Information Needed]
|
| 156 |
+
- **Cloud Provider:** [More Information Needed]
|
| 157 |
+
- **Compute Region:** [More Information Needed]
|
| 158 |
+
- **Carbon Emitted:** [More Information Needed]
|
| 159 |
+
|
| 160 |
+
## Technical Specifications [optional]
|
| 161 |
+
|
| 162 |
+
### Model Architecture and Objective
|
| 163 |
+
|
| 164 |
+
[More Information Needed]
|
| 165 |
+
|
| 166 |
+
### Compute Infrastructure
|
| 167 |
+
|
| 168 |
+
[More Information Needed]
|
| 169 |
+
|
| 170 |
+
#### Hardware
|
| 171 |
+
|
| 172 |
+
[More Information Needed]
|
| 173 |
+
|
| 174 |
+
#### Software
|
| 175 |
+
|
| 176 |
+
[More Information Needed]
|
| 177 |
+
|
| 178 |
+
## Citation [optional]
|
| 179 |
+
|
| 180 |
+
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
| 181 |
+
|
| 182 |
+
**BibTeX:**
|
| 183 |
+
|
| 184 |
+
[More Information Needed]
|
| 185 |
+
|
| 186 |
+
**APA:**
|
| 187 |
+
|
| 188 |
+
[More Information Needed]
|
| 189 |
+
|
| 190 |
+
## Glossary [optional]
|
| 191 |
+
|
| 192 |
+
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
| 193 |
+
|
| 194 |
+
[More Information Needed]
|
| 195 |
+
|
| 196 |
+
## More Information [optional]
|
| 197 |
+
|
| 198 |
+
[More Information Needed]
|
| 199 |
+
|
| 200 |
+
## Model Card Authors [optional]
|
| 201 |
+
|
| 202 |
+
[More Information Needed]
|
| 203 |
+
|
| 204 |
+
## Model Card Contact
|
| 205 |
+
|
| 206 |
+
[More Information Needed]
|
| 207 |
+
### Framework versions
|
| 208 |
+
|
| 209 |
+
- PEFT 0.19.1
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/adapter_config.json
ADDED
|
@@ -0,0 +1,48 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "Qwen/Qwen3-14B-Base",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"layer_replication": null,
|
| 16 |
+
"layers_pattern": null,
|
| 17 |
+
"layers_to_transform": null,
|
| 18 |
+
"loftq_config": {},
|
| 19 |
+
"lora_alpha": 256,
|
| 20 |
+
"lora_bias": false,
|
| 21 |
+
"lora_dropout": 0.00237968804112545,
|
| 22 |
+
"lora_ga_config": null,
|
| 23 |
+
"megatron_config": null,
|
| 24 |
+
"megatron_core": "megatron.core",
|
| 25 |
+
"modules_to_save": null,
|
| 26 |
+
"peft_type": "LORA",
|
| 27 |
+
"peft_version": "0.19.1",
|
| 28 |
+
"qalora_group_size": 16,
|
| 29 |
+
"r": 128,
|
| 30 |
+
"rank_pattern": {},
|
| 31 |
+
"revision": null,
|
| 32 |
+
"target_modules": [
|
| 33 |
+
"gate_proj",
|
| 34 |
+
"q_proj",
|
| 35 |
+
"up_proj",
|
| 36 |
+
"o_proj",
|
| 37 |
+
"v_proj",
|
| 38 |
+
"k_proj",
|
| 39 |
+
"down_proj"
|
| 40 |
+
],
|
| 41 |
+
"target_parameters": null,
|
| 42 |
+
"task_type": "CAUSAL_LM",
|
| 43 |
+
"trainable_token_indices": null,
|
| 44 |
+
"use_bdlora": null,
|
| 45 |
+
"use_dora": false,
|
| 46 |
+
"use_qalora": false,
|
| 47 |
+
"use_rslora": false
|
| 48 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/chat_template.jinja
ADDED
|
@@ -0,0 +1,85 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- if tools %}
|
| 2 |
+
{{- '<|im_start|>system\n' }}
|
| 3 |
+
{%- if messages[0].role == 'system' %}
|
| 4 |
+
{{- messages[0].content + '\n\n' }}
|
| 5 |
+
{%- endif %}
|
| 6 |
+
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
| 7 |
+
{%- for tool in tools %}
|
| 8 |
+
{{- "\n" }}
|
| 9 |
+
{{- tool | tojson }}
|
| 10 |
+
{%- endfor %}
|
| 11 |
+
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
| 12 |
+
{%- else %}
|
| 13 |
+
{%- if messages[0].role == 'system' %}
|
| 14 |
+
{{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
|
| 15 |
+
{%- endif %}
|
| 16 |
+
{%- endif %}
|
| 17 |
+
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
| 18 |
+
{%- for message in messages[::-1] %}
|
| 19 |
+
{%- set index = (messages|length - 1) - loop.index0 %}
|
| 20 |
+
{%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
|
| 21 |
+
{%- set ns.multi_step_tool = false %}
|
| 22 |
+
{%- set ns.last_query_index = index %}
|
| 23 |
+
{%- endif %}
|
| 24 |
+
{%- endfor %}
|
| 25 |
+
{%- for message in messages %}
|
| 26 |
+
{%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
|
| 27 |
+
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
|
| 28 |
+
{%- elif message.role == "assistant" %}
|
| 29 |
+
{%- set content = message.content %}
|
| 30 |
+
{%- set reasoning_content = '' %}
|
| 31 |
+
{%- if message.reasoning_content is defined and message.reasoning_content is not none %}
|
| 32 |
+
{%- set reasoning_content = message.reasoning_content %}
|
| 33 |
+
{%- else %}
|
| 34 |
+
{%- if '</think>' in message.content %}
|
| 35 |
+
{%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
|
| 36 |
+
{%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
| 37 |
+
{%- endif %}
|
| 38 |
+
{%- endif %}
|
| 39 |
+
{%- if loop.index0 > ns.last_query_index %}
|
| 40 |
+
{%- if loop.last or (not loop.last and reasoning_content) %}
|
| 41 |
+
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
|
| 42 |
+
{%- else %}
|
| 43 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 44 |
+
{%- endif %}
|
| 45 |
+
{%- else %}
|
| 46 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 47 |
+
{%- endif %}
|
| 48 |
+
{%- if message.tool_calls %}
|
| 49 |
+
{%- for tool_call in message.tool_calls %}
|
| 50 |
+
{%- if (loop.first and content) or (not loop.first) %}
|
| 51 |
+
{{- '\n' }}
|
| 52 |
+
{%- endif %}
|
| 53 |
+
{%- if tool_call.function %}
|
| 54 |
+
{%- set tool_call = tool_call.function %}
|
| 55 |
+
{%- endif %}
|
| 56 |
+
{{- '<tool_call>\n{"name": "' }}
|
| 57 |
+
{{- tool_call.name }}
|
| 58 |
+
{{- '", "arguments": ' }}
|
| 59 |
+
{%- if tool_call.arguments is string %}
|
| 60 |
+
{{- tool_call.arguments }}
|
| 61 |
+
{%- else %}
|
| 62 |
+
{{- tool_call.arguments | tojson }}
|
| 63 |
+
{%- endif %}
|
| 64 |
+
{{- '}\n</tool_call>' }}
|
| 65 |
+
{%- endfor %}
|
| 66 |
+
{%- endif %}
|
| 67 |
+
{{- '<|im_end|>\n' }}
|
| 68 |
+
{%- elif message.role == "tool" %}
|
| 69 |
+
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
|
| 70 |
+
{{- '<|im_start|>user' }}
|
| 71 |
+
{%- endif %}
|
| 72 |
+
{{- '\n<tool_response>\n' }}
|
| 73 |
+
{{- message.content }}
|
| 74 |
+
{{- '\n</tool_response>' }}
|
| 75 |
+
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
| 76 |
+
{{- '<|im_end|>\n' }}
|
| 77 |
+
{%- endif %}
|
| 78 |
+
{%- endif %}
|
| 79 |
+
{%- endfor %}
|
| 80 |
+
{%- if add_generation_prompt %}
|
| 81 |
+
{{- '<|im_start|>assistant\n' }}
|
| 82 |
+
{%- if enable_thinking is defined and enable_thinking is false %}
|
| 83 |
+
{{- '<think>\n\n</think>\n\n' }}
|
| 84 |
+
{%- endif %}
|
| 85 |
+
{%- endif %}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/tokenizer_config.json
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"backend": "tokenizers",
|
| 4 |
+
"bos_token": null,
|
| 5 |
+
"clean_up_tokenization_spaces": false,
|
| 6 |
+
"eos_token": "<|endoftext|>",
|
| 7 |
+
"errors": "replace",
|
| 8 |
+
"extra_special_tokens": [
|
| 9 |
+
"<|im_start|>",
|
| 10 |
+
"<|im_end|>",
|
| 11 |
+
"<|object_ref_start|>",
|
| 12 |
+
"<|object_ref_end|>",
|
| 13 |
+
"<|box_start|>",
|
| 14 |
+
"<|box_end|>",
|
| 15 |
+
"<|quad_start|>",
|
| 16 |
+
"<|quad_end|>",
|
| 17 |
+
"<|vision_start|>",
|
| 18 |
+
"<|vision_end|>",
|
| 19 |
+
"<|vision_pad|>",
|
| 20 |
+
"<|image_pad|>",
|
| 21 |
+
"<|video_pad|>"
|
| 22 |
+
],
|
| 23 |
+
"is_local": false,
|
| 24 |
+
"model_max_length": 131072,
|
| 25 |
+
"pad_token": "<|endoftext|>",
|
| 26 |
+
"split_special_tokens": false,
|
| 27 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 28 |
+
"unk_token": null
|
| 29 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3112/trainer_state.json
ADDED
|
@@ -0,0 +1,742 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"best_global_step": null,
|
| 3 |
+
"best_metric": null,
|
| 4 |
+
"best_model_checkpoint": null,
|
| 5 |
+
"epoch": 8.0,
|
| 6 |
+
"eval_steps": 500,
|
| 7 |
+
"global_step": 3112,
|
| 8 |
+
"is_hyper_param_search": false,
|
| 9 |
+
"is_local_process_zero": true,
|
| 10 |
+
"is_world_process_zero": true,
|
| 11 |
+
"log_history": [
|
| 12 |
+
{
|
| 13 |
+
"entropy": 1.6110110306739807,
|
| 14 |
+
"epoch": 0.1287001287001287,
|
| 15 |
+
"grad_norm": 0.8611119389533997,
|
| 16 |
+
"learning_rate": 3.4130257962133866e-05,
|
| 17 |
+
"loss": 1.53956298828125,
|
| 18 |
+
"mean_token_accuracy": 0.6652342769503593,
|
| 19 |
+
"num_tokens": 73407.0,
|
| 20 |
+
"step": 50
|
| 21 |
+
},
|
| 22 |
+
{
|
| 23 |
+
"entropy": 0.8662820833921433,
|
| 24 |
+
"epoch": 0.2574002574002574,
|
| 25 |
+
"grad_norm": 0.7405035495758057,
|
| 26 |
+
"learning_rate": 6.895705180104598e-05,
|
| 27 |
+
"loss": 0.7929539489746094,
|
| 28 |
+
"mean_token_accuracy": 0.7786347842216492,
|
| 29 |
+
"num_tokens": 143994.0,
|
| 30 |
+
"step": 100
|
| 31 |
+
},
|
| 32 |
+
{
|
| 33 |
+
"entropy": 0.7912819278240204,
|
| 34 |
+
"epoch": 0.3861003861003861,
|
| 35 |
+
"grad_norm": 0.6105485558509827,
|
| 36 |
+
"learning_rate": 0.00010378384563995809,
|
| 37 |
+
"loss": 0.7236511993408203,
|
| 38 |
+
"mean_token_accuracy": 0.7925467795133591,
|
| 39 |
+
"num_tokens": 216171.0,
|
| 40 |
+
"step": 150
|
| 41 |
+
},
|
| 42 |
+
{
|
| 43 |
+
"entropy": 0.7777709531784057,
|
| 44 |
+
"epoch": 0.5148005148005148,
|
| 45 |
+
"grad_norm": 0.5781793594360352,
|
| 46 |
+
"learning_rate": 0.0001386106394788702,
|
| 47 |
+
"loss": 0.7024919891357422,
|
| 48 |
+
"mean_token_accuracy": 0.7969876372814179,
|
| 49 |
+
"num_tokens": 284702.0,
|
| 50 |
+
"step": 200
|
| 51 |
+
},
|
| 52 |
+
{
|
| 53 |
+
"entropy": 0.7632609683275223,
|
| 54 |
+
"epoch": 0.6435006435006435,
|
| 55 |
+
"grad_norm": 0.6553444862365723,
|
| 56 |
+
"learning_rate": 0.00017343743331778232,
|
| 57 |
+
"loss": 0.7000718688964844,
|
| 58 |
+
"mean_token_accuracy": 0.7994742071628571,
|
| 59 |
+
"num_tokens": 356393.0,
|
| 60 |
+
"step": 250
|
| 61 |
+
},
|
| 62 |
+
{
|
| 63 |
+
"entropy": 0.754165632724762,
|
| 64 |
+
"epoch": 0.7722007722007722,
|
| 65 |
+
"grad_norm": 0.634222149848938,
|
| 66 |
+
"learning_rate": 0.00020826422715669444,
|
| 67 |
+
"loss": 0.6909049987792969,
|
| 68 |
+
"mean_token_accuracy": 0.8016809666156769,
|
| 69 |
+
"num_tokens": 426916.0,
|
| 70 |
+
"step": 300
|
| 71 |
+
},
|
| 72 |
+
{
|
| 73 |
+
"entropy": 0.7530711203813553,
|
| 74 |
+
"epoch": 0.9009009009009009,
|
| 75 |
+
"grad_norm": 0.7020156383514404,
|
| 76 |
+
"learning_rate": 0.00024309102099560653,
|
| 77 |
+
"loss": 0.69554931640625,
|
| 78 |
+
"mean_token_accuracy": 0.8002807641029358,
|
| 79 |
+
"num_tokens": 499603.0,
|
| 80 |
+
"step": 350
|
| 81 |
+
},
|
| 82 |
+
{
|
| 83 |
+
"epoch": 1.0,
|
| 84 |
+
"eval_entropy": 0.5728044164242204,
|
| 85 |
+
"eval_loss": 0.6785851716995239,
|
| 86 |
+
"eval_mean_token_accuracy": 0.798433246993527,
|
| 87 |
+
"eval_num_tokens": 555493.0,
|
| 88 |
+
"eval_runtime": 162.4129,
|
| 89 |
+
"eval_samples_per_second": 9.513,
|
| 90 |
+
"eval_steps_per_second": 1.194,
|
| 91 |
+
"step": 389
|
| 92 |
+
},
|
| 93 |
+
{
|
| 94 |
+
"entropy": 0.7489313946829902,
|
| 95 |
+
"epoch": 1.0283140283140284,
|
| 96 |
+
"grad_norm": 0.7532815933227539,
|
| 97 |
+
"learning_rate": 0.0002709470016827303,
|
| 98 |
+
"loss": 0.6856581878662109,
|
| 99 |
+
"mean_token_accuracy": 0.8023027079273956,
|
| 100 |
+
"num_tokens": 570813.0,
|
| 101 |
+
"step": 400
|
| 102 |
+
},
|
| 103 |
+
{
|
| 104 |
+
"entropy": 0.7238409864902496,
|
| 105 |
+
"epoch": 1.157014157014157,
|
| 106 |
+
"grad_norm": 1.2016215324401855,
|
| 107 |
+
"learning_rate": 0.0002707561443541359,
|
| 108 |
+
"loss": 0.6699818420410156,
|
| 109 |
+
"mean_token_accuracy": 0.8070302194356919,
|
| 110 |
+
"num_tokens": 642956.0,
|
| 111 |
+
"step": 450
|
| 112 |
+
},
|
| 113 |
+
{
|
| 114 |
+
"entropy": 0.7388938587903976,
|
| 115 |
+
"epoch": 1.2857142857142856,
|
| 116 |
+
"grad_norm": 0.7279272079467773,
|
| 117 |
+
"learning_rate": 0.0002702930068622498,
|
| 118 |
+
"loss": 0.6728517150878907,
|
| 119 |
+
"mean_token_accuracy": 0.8049580943584442,
|
| 120 |
+
"num_tokens": 714498.0,
|
| 121 |
+
"step": 500
|
| 122 |
+
},
|
| 123 |
+
{
|
| 124 |
+
"entropy": 0.7284485149383545,
|
| 125 |
+
"epoch": 1.4144144144144144,
|
| 126 |
+
"grad_norm": 0.8563987016677856,
|
| 127 |
+
"learning_rate": 0.0002695585213716931,
|
| 128 |
+
"loss": 0.6657986450195312,
|
| 129 |
+
"mean_token_accuracy": 0.8085588800907135,
|
| 130 |
+
"num_tokens": 785595.0,
|
| 131 |
+
"step": 550
|
| 132 |
+
},
|
| 133 |
+
{
|
| 134 |
+
"entropy": 0.6960959500074386,
|
| 135 |
+
"epoch": 1.5431145431145432,
|
| 136 |
+
"grad_norm": 0.5145474672317505,
|
| 137 |
+
"learning_rate": 0.0002685541661937683,
|
| 138 |
+
"loss": 0.6358638763427734,
|
| 139 |
+
"mean_token_accuracy": 0.8131621342897415,
|
| 140 |
+
"num_tokens": 857551.0,
|
| 141 |
+
"step": 600
|
| 142 |
+
},
|
| 143 |
+
{
|
| 144 |
+
"entropy": 0.7220414417982102,
|
| 145 |
+
"epoch": 1.6718146718146718,
|
| 146 |
+
"grad_norm": 0.9381059408187866,
|
| 147 |
+
"learning_rate": 0.00026728196281103746,
|
| 148 |
+
"loss": 0.6531407928466797,
|
| 149 |
+
"mean_token_accuracy": 0.811619822382927,
|
| 150 |
+
"num_tokens": 926864.0,
|
| 151 |
+
"step": 650
|
| 152 |
+
},
|
| 153 |
+
{
|
| 154 |
+
"entropy": 0.6955459499359131,
|
| 155 |
+
"epoch": 1.8005148005148004,
|
| 156 |
+
"grad_norm": 0.6115108728408813,
|
| 157 |
+
"learning_rate": 0.0002657444718086503,
|
| 158 |
+
"loss": 0.6269588088989257,
|
| 159 |
+
"mean_token_accuracy": 0.8155373805761337,
|
| 160 |
+
"num_tokens": 998824.0,
|
| 161 |
+
"step": 700
|
| 162 |
+
},
|
| 163 |
+
{
|
| 164 |
+
"entropy": 0.6929812705516816,
|
| 165 |
+
"epoch": 1.9292149292149292,
|
| 166 |
+
"grad_norm": 0.749489426612854,
|
| 167 |
+
"learning_rate": 0.0002639447877206115,
|
| 168 |
+
"loss": 0.629054069519043,
|
| 169 |
+
"mean_token_accuracy": 0.8183021235466004,
|
| 170 |
+
"num_tokens": 1069332.0,
|
| 171 |
+
"step": 750
|
| 172 |
+
},
|
| 173 |
+
{
|
| 174 |
+
"epoch": 2.0,
|
| 175 |
+
"eval_entropy": 0.5833694102223387,
|
| 176 |
+
"eval_loss": 0.6328718662261963,
|
| 177 |
+
"eval_mean_token_accuracy": 0.8159044071571114,
|
| 178 |
+
"eval_num_tokens": 1110986.0,
|
| 179 |
+
"eval_runtime": 161.885,
|
| 180 |
+
"eval_samples_per_second": 9.544,
|
| 181 |
+
"eval_steps_per_second": 1.198,
|
| 182 |
+
"step": 778
|
| 183 |
+
},
|
| 184 |
+
{
|
| 185 |
+
"entropy": 0.647241060480927,
|
| 186 |
+
"epoch": 2.056628056628057,
|
| 187 |
+
"grad_norm": 0.7070767879486084,
|
| 188 |
+
"learning_rate": 0.00026188653280135975,
|
| 189 |
+
"loss": 0.5823922348022461,
|
| 190 |
+
"mean_token_accuracy": 0.8265365301960647,
|
| 191 |
+
"num_tokens": 1141195.0,
|
| 192 |
+
"step": 800
|
| 193 |
+
},
|
| 194 |
+
{
|
| 195 |
+
"entropy": 0.5995378407835961,
|
| 196 |
+
"epoch": 2.1853281853281854,
|
| 197 |
+
"grad_norm": 0.8090486526489258,
|
| 198 |
+
"learning_rate": 0.0002595738497351955,
|
| 199 |
+
"loss": 0.5325597763061524,
|
| 200 |
+
"mean_token_accuracy": 0.8369336777925491,
|
| 201 |
+
"num_tokens": 1210708.0,
|
| 202 |
+
"step": 850
|
| 203 |
+
},
|
| 204 |
+
{
|
| 205 |
+
"entropy": 0.6043145382404327,
|
| 206 |
+
"epoch": 2.314028314028314,
|
| 207 |
+
"grad_norm": 0.8279913067817688,
|
| 208 |
+
"learning_rate": 0.00025701139329823054,
|
| 209 |
+
"loss": 0.5414446258544922,
|
| 210 |
+
"mean_token_accuracy": 0.8361396533250809,
|
| 211 |
+
"num_tokens": 1283441.0,
|
| 212 |
+
"step": 900
|
| 213 |
+
},
|
| 214 |
+
{
|
| 215 |
+
"entropy": 0.5953224584460258,
|
| 216 |
+
"epoch": 2.4427284427284426,
|
| 217 |
+
"grad_norm": 0.6075023412704468,
|
| 218 |
+
"learning_rate": 0.00025420432098964183,
|
| 219 |
+
"loss": 0.536654167175293,
|
| 220 |
+
"mean_token_accuracy": 0.8340093129873276,
|
| 221 |
+
"num_tokens": 1356418.0,
|
| 222 |
+
"step": 950
|
| 223 |
+
},
|
| 224 |
+
{
|
| 225 |
+
"entropy": 0.5998479858040809,
|
| 226 |
+
"epoch": 2.571428571428571,
|
| 227 |
+
"grad_norm": 1.0311471223831177,
|
| 228 |
+
"learning_rate": 0.0002511582826510862,
|
| 229 |
+
"loss": 0.5372924423217773,
|
| 230 |
+
"mean_token_accuracy": 0.8366045409440994,
|
| 231 |
+
"num_tokens": 1427797.0,
|
| 232 |
+
"step": 1000
|
| 233 |
+
},
|
| 234 |
+
{
|
| 235 |
+
"entropy": 0.5980158120393753,
|
| 236 |
+
"epoch": 2.7001287001287,
|
| 237 |
+
"grad_norm": 0.5971426367759705,
|
| 238 |
+
"learning_rate": 0.0002478794090951689,
|
| 239 |
+
"loss": 0.5392885208129883,
|
| 240 |
+
"mean_token_accuracy": 0.8347727072238922,
|
| 241 |
+
"num_tokens": 1498082.0,
|
| 242 |
+
"step": 1050
|
| 243 |
+
},
|
| 244 |
+
{
|
| 245 |
+
"entropy": 0.5990680930018425,
|
| 246 |
+
"epoch": 2.828828828828829,
|
| 247 |
+
"grad_norm": 0.5662627220153809,
|
| 248 |
+
"learning_rate": 0.0002443742997658538,
|
| 249 |
+
"loss": 0.5360498428344727,
|
| 250 |
+
"mean_token_accuracy": 0.8371847170591354,
|
| 251 |
+
"num_tokens": 1568798.0,
|
| 252 |
+
"step": 1100
|
| 253 |
+
},
|
| 254 |
+
{
|
| 255 |
+
"entropy": 0.5957039377093315,
|
| 256 |
+
"epoch": 2.9575289575289574,
|
| 257 |
+
"grad_norm": 0.5043798685073853,
|
| 258 |
+
"learning_rate": 0.00024065000945565205,
|
| 259 |
+
"loss": 0.5342231369018555,
|
| 260 |
+
"mean_token_accuracy": 0.8380735236406326,
|
| 261 |
+
"num_tokens": 1643449.0,
|
| 262 |
+
"step": 1150
|
| 263 |
+
},
|
| 264 |
+
{
|
| 265 |
+
"epoch": 3.0,
|
| 266 |
+
"eval_entropy": 0.5346978819861854,
|
| 267 |
+
"eval_loss": 0.6255015134811401,
|
| 268 |
+
"eval_mean_token_accuracy": 0.8202014476368108,
|
| 269 |
+
"eval_num_tokens": 1666479.0,
|
| 270 |
+
"eval_runtime": 161.6098,
|
| 271 |
+
"eval_samples_per_second": 9.56,
|
| 272 |
+
"eval_steps_per_second": 1.2,
|
| 273 |
+
"step": 1167
|
| 274 |
+
},
|
| 275 |
+
{
|
| 276 |
+
"entropy": 0.5212209137401196,
|
| 277 |
+
"epoch": 3.0849420849420848,
|
| 278 |
+
"grad_norm": 0.7095440626144409,
|
| 279 |
+
"learning_rate": 0.00023671403410632178,
|
| 280 |
+
"loss": 0.45311901092529294,
|
| 281 |
+
"mean_token_accuracy": 0.856362871449403,
|
| 282 |
+
"num_tokens": 1713536.0,
|
| 283 |
+
"step": 1200
|
| 284 |
+
},
|
| 285 |
+
{
|
| 286 |
+
"entropy": 0.4701593083143234,
|
| 287 |
+
"epoch": 3.213642213642214,
|
| 288 |
+
"grad_norm": 0.6408083438873291,
|
| 289 |
+
"learning_rate": 0.0002325742957216607,
|
| 290 |
+
"loss": 0.39916397094726563,
|
| 291 |
+
"mean_token_accuracy": 0.8698061722517013,
|
| 292 |
+
"num_tokens": 1785609.0,
|
| 293 |
+
"step": 1250
|
| 294 |
+
},
|
| 295 |
+
{
|
| 296 |
+
"entropy": 0.4766591975092888,
|
| 297 |
+
"epoch": 3.3423423423423424,
|
| 298 |
+
"grad_norm": 0.6415093541145325,
|
| 299 |
+
"learning_rate": 0.0002282391264227552,
|
| 300 |
+
"loss": 0.4116698455810547,
|
| 301 |
+
"mean_token_accuracy": 0.8679435575008392,
|
| 302 |
+
"num_tokens": 1858651.0,
|
| 303 |
+
"step": 1300
|
| 304 |
+
},
|
| 305 |
+
{
|
| 306 |
+
"entropy": 0.4937947469949722,
|
| 307 |
+
"epoch": 3.471042471042471,
|
| 308 |
+
"grad_norm": 0.6549825072288513,
|
| 309 |
+
"learning_rate": 0.00022371725167778054,
|
| 310 |
+
"loss": 0.4296376037597656,
|
| 311 |
+
"mean_token_accuracy": 0.8609357953071595,
|
| 312 |
+
"num_tokens": 1928692.0,
|
| 313 |
+
"step": 1350
|
| 314 |
+
},
|
| 315 |
+
{
|
| 316 |
+
"entropy": 0.4920153194665909,
|
| 317 |
+
"epoch": 3.5997425997425996,
|
| 318 |
+
"grad_norm": 0.6452126502990723,
|
| 319 |
+
"learning_rate": 0.00021901777274010406,
|
| 320 |
+
"loss": 0.4307489013671875,
|
| 321 |
+
"mean_token_accuracy": 0.8606827831268311,
|
| 322 |
+
"num_tokens": 1998668.0,
|
| 323 |
+
"step": 1400
|
| 324 |
+
},
|
| 325 |
+
{
|
| 326 |
+
"entropy": 0.490042342543602,
|
| 327 |
+
"epoch": 3.7284427284427286,
|
| 328 |
+
"grad_norm": 0.5727734565734863,
|
| 329 |
+
"learning_rate": 0.0002141501483300395,
|
| 330 |
+
"loss": 0.4295254135131836,
|
| 331 |
+
"mean_token_accuracy": 0.8616545403003693,
|
| 332 |
+
"num_tokens": 2072809.0,
|
| 333 |
+
"step": 1450
|
| 334 |
+
},
|
| 335 |
+
{
|
| 336 |
+
"entropy": 0.49857193052768706,
|
| 337 |
+
"epoch": 3.857142857142857,
|
| 338 |
+
"grad_norm": 0.7732954025268555,
|
| 339 |
+
"learning_rate": 0.00020912417559712133,
|
| 340 |
+
"loss": 0.4289303207397461,
|
| 341 |
+
"mean_token_accuracy": 0.8616443765163422,
|
| 342 |
+
"num_tokens": 2142475.0,
|
| 343 |
+
"step": 1500
|
| 344 |
+
},
|
| 345 |
+
{
|
| 346 |
+
"entropy": 0.4746784272789955,
|
| 347 |
+
"epoch": 3.985842985842986,
|
| 348 |
+
"grad_norm": 0.5791187882423401,
|
| 349 |
+
"learning_rate": 0.00020394997040121726,
|
| 350 |
+
"loss": 0.4180263900756836,
|
| 351 |
+
"mean_token_accuracy": 0.866080379486084,
|
| 352 |
+
"num_tokens": 2214044.0,
|
| 353 |
+
"step": 1550
|
| 354 |
+
},
|
| 355 |
+
{
|
| 356 |
+
"epoch": 4.0,
|
| 357 |
+
"eval_entropy": 0.4793926059585257,
|
| 358 |
+
"eval_loss": 0.627627968788147,
|
| 359 |
+
"eval_mean_token_accuracy": 0.8233870095813397,
|
| 360 |
+
"eval_num_tokens": 2221972.0,
|
| 361 |
+
"eval_runtime": 162.0225,
|
| 362 |
+
"eval_samples_per_second": 9.536,
|
| 363 |
+
"eval_steps_per_second": 1.197,
|
| 364 |
+
"step": 1556
|
| 365 |
+
},
|
| 366 |
+
{
|
| 367 |
+
"entropy": 0.38195489000792454,
|
| 368 |
+
"epoch": 4.113256113256114,
|
| 369 |
+
"grad_norm": 0.6002617478370667,
|
| 370 |
+
"learning_rate": 0.0001986379469521669,
|
| 371 |
+
"loss": 0.30819049835205076,
|
| 372 |
+
"mean_token_accuracy": 0.8977848634575353,
|
| 373 |
+
"num_tokens": 2282164.0,
|
| 374 |
+
"step": 1600
|
| 375 |
+
},
|
| 376 |
+
{
|
| 377 |
+
"entropy": 0.3655787402391434,
|
| 378 |
+
"epoch": 4.241956241956242,
|
| 379 |
+
"grad_norm": 0.7100041508674622,
|
| 380 |
+
"learning_rate": 0.00019319879684892634,
|
| 381 |
+
"loss": 0.29959835052490236,
|
| 382 |
+
"mean_token_accuracy": 0.8991208010911942,
|
| 383 |
+
"num_tokens": 2353213.0,
|
| 384 |
+
"step": 1650
|
| 385 |
+
},
|
| 386 |
+
{
|
| 387 |
+
"entropy": 0.3821141055226326,
|
| 388 |
+
"epoch": 4.370656370656371,
|
| 389 |
+
"grad_norm": 0.5848307013511658,
|
| 390 |
+
"learning_rate": 0.00018764346756040715,
|
| 391 |
+
"loss": 0.313802490234375,
|
| 392 |
+
"mean_token_accuracy": 0.895167955160141,
|
| 393 |
+
"num_tokens": 2425068.0,
|
| 394 |
+
"step": 1700
|
| 395 |
+
},
|
| 396 |
+
{
|
| 397 |
+
"entropy": 0.37083797007799146,
|
| 398 |
+
"epoch": 4.499356499356499,
|
| 399 |
+
"grad_norm": 0.6447024941444397,
|
| 400 |
+
"learning_rate": 0.00018198314039132143,
|
| 401 |
+
"loss": 0.30583988189697264,
|
| 402 |
+
"mean_token_accuracy": 0.8961733293533325,
|
| 403 |
+
"num_tokens": 2498321.0,
|
| 404 |
+
"step": 1750
|
| 405 |
+
},
|
| 406 |
+
{
|
| 407 |
+
"entropy": 0.3791545969247818,
|
| 408 |
+
"epoch": 4.628056628056628,
|
| 409 |
+
"grad_norm": 0.6575382351875305,
|
| 410 |
+
"learning_rate": 0.00017622920797738184,
|
| 411 |
+
"loss": 0.3088031005859375,
|
| 412 |
+
"mean_token_accuracy": 0.8960050916671753,
|
| 413 |
+
"num_tokens": 2570321.0,
|
| 414 |
+
"step": 1800
|
| 415 |
+
},
|
| 416 |
+
{
|
| 417 |
+
"entropy": 0.3946831756830215,
|
| 418 |
+
"epoch": 4.756756756756757,
|
| 419 |
+
"grad_norm": 0.5351552963256836,
|
| 420 |
+
"learning_rate": 0.00017039325135515207,
|
| 421 |
+
"loss": 0.3229162979125977,
|
| 422 |
+
"mean_token_accuracy": 0.8920552498102188,
|
| 423 |
+
"num_tokens": 2642851.0,
|
| 424 |
+
"step": 1850
|
| 425 |
+
},
|
| 426 |
+
{
|
| 427 |
+
"entropy": 0.37298239797353744,
|
| 428 |
+
"epoch": 4.885456885456885,
|
| 429 |
+
"grad_norm": 0.7624587416648865,
|
| 430 |
+
"learning_rate": 0.00016448701665269964,
|
| 431 |
+
"loss": 0.3067934799194336,
|
| 432 |
+
"mean_token_accuracy": 0.8951873427629471,
|
| 433 |
+
"num_tokens": 2715629.0,
|
| 434 |
+
"step": 1900
|
| 435 |
+
},
|
| 436 |
+
{
|
| 437 |
+
"epoch": 5.0,
|
| 438 |
+
"eval_entropy": 0.3738150204887095,
|
| 439 |
+
"eval_loss": 0.7131896615028381,
|
| 440 |
+
"eval_mean_token_accuracy": 0.8207752468045225,
|
| 441 |
+
"eval_num_tokens": 2777465.0,
|
| 442 |
+
"eval_runtime": 162.1244,
|
| 443 |
+
"eval_samples_per_second": 9.53,
|
| 444 |
+
"eval_steps_per_second": 1.197,
|
| 445 |
+
"step": 1945
|
| 446 |
+
},
|
| 447 |
+
{
|
| 448 |
+
"entropy": 0.377820266617669,
|
| 449 |
+
"epoch": 5.012870012870013,
|
| 450 |
+
"grad_norm": 0.41990190744400024,
|
| 451 |
+
"learning_rate": 0.00015852239144796624,
|
| 452 |
+
"loss": 0.3058685111999512,
|
| 453 |
+
"mean_token_accuracy": 0.8964343480389527,
|
| 454 |
+
"num_tokens": 2784896.0,
|
| 455 |
+
"step": 1950
|
| 456 |
+
},
|
| 457 |
+
{
|
| 458 |
+
"entropy": 0.2714502356946468,
|
| 459 |
+
"epoch": 5.141570141570142,
|
| 460 |
+
"grad_norm": 0.422568678855896,
|
| 461 |
+
"learning_rate": 0.00015251138084243995,
|
| 462 |
+
"loss": 0.2093442153930664,
|
| 463 |
+
"mean_token_accuracy": 0.9311346983909607,
|
| 464 |
+
"num_tokens": 2854374.0,
|
| 465 |
+
"step": 2000
|
| 466 |
+
},
|
| 467 |
+
{
|
| 468 |
+
"entropy": 0.268475965410471,
|
| 469 |
+
"epoch": 5.27027027027027,
|
| 470 |
+
"grad_norm": 0.6637414693832397,
|
| 471 |
+
"learning_rate": 0.0001464660832982852,
|
| 472 |
+
"loss": 0.20736080169677734,
|
| 473 |
+
"mean_token_accuracy": 0.9289199805259705,
|
| 474 |
+
"num_tokens": 2927362.0,
|
| 475 |
+
"step": 2050
|
| 476 |
+
},
|
| 477 |
+
{
|
| 478 |
+
"entropy": 0.2644876340031624,
|
| 479 |
+
"epoch": 5.398970398970399,
|
| 480 |
+
"grad_norm": 0.47317707538604736,
|
| 481 |
+
"learning_rate": 0.00014039866628756467,
|
| 482 |
+
"loss": 0.20464908599853515,
|
| 483 |
+
"mean_token_accuracy": 0.9300856202840805,
|
| 484 |
+
"num_tokens": 3000143.0,
|
| 485 |
+
"step": 2100
|
| 486 |
+
},
|
| 487 |
+
{
|
| 488 |
+
"entropy": 0.2675253136456013,
|
| 489 |
+
"epoch": 5.527670527670527,
|
| 490 |
+
"grad_norm": 0.5253982543945312,
|
| 491 |
+
"learning_rate": 0.00013432134180256338,
|
| 492 |
+
"loss": 0.21154335021972656,
|
| 493 |
+
"mean_token_accuracy": 0.9283734840154648,
|
| 494 |
+
"num_tokens": 3072561.0,
|
| 495 |
+
"step": 2150
|
| 496 |
+
},
|
| 497 |
+
{
|
| 498 |
+
"entropy": 0.27213907435536383,
|
| 499 |
+
"epoch": 5.656370656370656,
|
| 500 |
+
"grad_norm": 0.46738553047180176,
|
| 501 |
+
"learning_rate": 0.00012824634177650664,
|
| 502 |
+
"loss": 0.21339216232299804,
|
| 503 |
+
"mean_token_accuracy": 0.9272083270549775,
|
| 504 |
+
"num_tokens": 3144831.0,
|
| 505 |
+
"step": 2200
|
| 506 |
+
},
|
| 507 |
+
{
|
| 508 |
+
"entropy": 0.2785488124191761,
|
| 509 |
+
"epoch": 5.785070785070785,
|
| 510 |
+
"grad_norm": 0.4469502866268158,
|
| 511 |
+
"learning_rate": 0.00012218589346414205,
|
| 512 |
+
"loss": 0.21601097106933595,
|
| 513 |
+
"mean_token_accuracy": 0.9255663657188415,
|
| 514 |
+
"num_tokens": 3215960.0,
|
| 515 |
+
"step": 2250
|
| 516 |
+
},
|
| 517 |
+
{
|
| 518 |
+
"entropy": 0.2699935150146484,
|
| 519 |
+
"epoch": 5.913770913770914,
|
| 520 |
+
"grad_norm": 0.7359778881072998,
|
| 521 |
+
"learning_rate": 0.00011615219483173828,
|
| 522 |
+
"loss": 0.20725584030151367,
|
| 523 |
+
"mean_token_accuracy": 0.9286630594730377,
|
| 524 |
+
"num_tokens": 3287499.0,
|
| 525 |
+
"step": 2300
|
| 526 |
+
},
|
| 527 |
+
{
|
| 528 |
+
"epoch": 6.0,
|
| 529 |
+
"eval_entropy": 0.26417383682174783,
|
| 530 |
+
"eval_loss": 0.8880229592323303,
|
| 531 |
+
"eval_mean_token_accuracy": 0.8159987201395723,
|
| 532 |
+
"eval_num_tokens": 3332958.0,
|
| 533 |
+
"eval_runtime": 162.0991,
|
| 534 |
+
"eval_samples_per_second": 9.531,
|
| 535 |
+
"eval_steps_per_second": 1.197,
|
| 536 |
+
"step": 2334
|
| 537 |
+
},
|
| 538 |
+
{
|
| 539 |
+
"entropy": 0.24814540704693458,
|
| 540 |
+
"epoch": 6.041184041184041,
|
| 541 |
+
"grad_norm": 0.4953760802745819,
|
| 542 |
+
"learning_rate": 0.00011015739000603316,
|
| 543 |
+
"loss": 0.18749794006347656,
|
| 544 |
+
"mean_token_accuracy": 0.9370789509831052,
|
| 545 |
+
"num_tokens": 3356879.0,
|
| 546 |
+
"step": 2350
|
| 547 |
+
},
|
| 548 |
+
{
|
| 549 |
+
"entropy": 0.19976271741092205,
|
| 550 |
+
"epoch": 6.1698841698841695,
|
| 551 |
+
"grad_norm": 0.4834803342819214,
|
| 552 |
+
"learning_rate": 0.00010421354483154553,
|
| 553 |
+
"loss": 0.14283526420593262,
|
| 554 |
+
"mean_token_accuracy": 0.9521516615152359,
|
| 555 |
+
"num_tokens": 3427587.0,
|
| 556 |
+
"step": 2400
|
| 557 |
+
},
|
| 558 |
+
{
|
| 559 |
+
"entropy": 0.2060488449037075,
|
| 560 |
+
"epoch": 6.298584298584299,
|
| 561 |
+
"grad_norm": 0.4888673722743988,
|
| 562 |
+
"learning_rate": 9.8332622585447e-05,
|
| 563 |
+
"loss": 0.14414511680603026,
|
| 564 |
+
"mean_token_accuracy": 0.9510996866226197,
|
| 565 |
+
"num_tokens": 3498688.0,
|
| 566 |
+
"step": 2450
|
| 567 |
+
},
|
| 568 |
+
{
|
| 569 |
+
"entropy": 0.2059111550450325,
|
| 570 |
+
"epoch": 6.427284427284428,
|
| 571 |
+
"grad_norm": 0.4064404368400574,
|
| 572 |
+
"learning_rate": 9.252645989887253e-05,
|
| 573 |
+
"loss": 0.14820143699645996,
|
| 574 |
+
"mean_token_accuracy": 0.9507584601640702,
|
| 575 |
+
"num_tokens": 3566137.0,
|
| 576 |
+
"step": 2500
|
| 577 |
+
},
|
| 578 |
+
{
|
| 579 |
+
"entropy": 0.19700154662132263,
|
| 580 |
+
"epoch": 6.555984555984556,
|
| 581 |
+
"grad_norm": 0.467965304851532,
|
| 582 |
+
"learning_rate": 8.680674293313417e-05,
|
| 583 |
+
"loss": 0.14303470611572267,
|
| 584 |
+
"mean_token_accuracy": 0.9515972435474396,
|
| 585 |
+
"num_tokens": 3639573.0,
|
| 586 |
+
"step": 2550
|
| 587 |
+
},
|
| 588 |
+
{
|
| 589 |
+
"entropy": 0.20180423602461814,
|
| 590 |
+
"epoch": 6.684684684684685,
|
| 591 |
+
"grad_norm": 0.36836138367652893,
|
| 592 |
+
"learning_rate": 8.118498385878736e-05,
|
| 593 |
+
"loss": 0.14280882835388184,
|
| 594 |
+
"mean_token_accuracy": 0.9515993863344192,
|
| 595 |
+
"num_tokens": 3710433.0,
|
| 596 |
+
"step": 2600
|
| 597 |
+
},
|
| 598 |
+
{
|
| 599 |
+
"entropy": 0.20024395987391472,
|
| 600 |
+
"epoch": 6.813384813384813,
|
| 601 |
+
"grad_norm": 0.38375866413116455,
|
| 602 |
+
"learning_rate": 7.567249768489171e-05,
|
| 603 |
+
"loss": 0.1427844524383545,
|
| 604 |
+
"mean_token_accuracy": 0.9524166631698608,
|
| 605 |
+
"num_tokens": 3781550.0,
|
| 606 |
+
"step": 2650
|
| 607 |
+
},
|
| 608 |
+
{
|
| 609 |
+
"entropy": 0.19561587080359458,
|
| 610 |
+
"epoch": 6.942084942084942,
|
| 611 |
+
"grad_norm": 0.41185441613197327,
|
| 612 |
+
"learning_rate": 7.028037948510187e-05,
|
| 613 |
+
"loss": 0.13993803024291993,
|
| 614 |
+
"mean_token_accuracy": 0.9522478264570237,
|
| 615 |
+
"num_tokens": 3854952.0,
|
| 616 |
+
"step": 2700
|
| 617 |
+
},
|
| 618 |
+
{
|
| 619 |
+
"epoch": 7.0,
|
| 620 |
+
"eval_entropy": 0.19501976062034823,
|
| 621 |
+
"eval_loss": 1.0653952360153198,
|
| 622 |
+
"eval_mean_token_accuracy": 0.8205490803595671,
|
| 623 |
+
"eval_num_tokens": 3888451.0,
|
| 624 |
+
"eval_runtime": 161.8533,
|
| 625 |
+
"eval_samples_per_second": 9.546,
|
| 626 |
+
"eval_steps_per_second": 1.199,
|
| 627 |
+
"step": 2723
|
| 628 |
+
},
|
| 629 |
+
{
|
| 630 |
+
"entropy": 0.17789882526855277,
|
| 631 |
+
"epoch": 7.06949806949807,
|
| 632 |
+
"grad_norm": 0.41413992643356323,
|
| 633 |
+
"learning_rate": 6.50194820664261e-05,
|
| 634 |
+
"loss": 0.12078390121459961,
|
| 635 |
+
"mean_token_accuracy": 0.9589925727458916,
|
| 636 |
+
"num_tokens": 3928354.0,
|
| 637 |
+
"step": 2750
|
| 638 |
+
},
|
| 639 |
+
{
|
| 640 |
+
"entropy": 0.16781829454004765,
|
| 641 |
+
"epoch": 7.198198198198198,
|
| 642 |
+
"grad_norm": 0.25806066393852234,
|
| 643 |
+
"learning_rate": 5.990039412559906e-05,
|
| 644 |
+
"loss": 0.10963023185729981,
|
| 645 |
+
"mean_token_accuracy": 0.9617267113924026,
|
| 646 |
+
"num_tokens": 4000113.0,
|
| 647 |
+
"step": 2800
|
| 648 |
+
},
|
| 649 |
+
{
|
| 650 |
+
"entropy": 0.1649068508297205,
|
| 651 |
+
"epoch": 7.326898326898327,
|
| 652 |
+
"grad_norm": 0.27411890029907227,
|
| 653 |
+
"learning_rate": 5.493341893703393e-05,
|
| 654 |
+
"loss": 0.11152458190917969,
|
| 655 |
+
"mean_token_accuracy": 0.9620639663934708,
|
| 656 |
+
"num_tokens": 4071032.0,
|
| 657 |
+
"step": 2850
|
| 658 |
+
},
|
| 659 |
+
{
|
| 660 |
+
"entropy": 0.161333369910717,
|
| 661 |
+
"epoch": 7.455598455598455,
|
| 662 |
+
"grad_norm": 0.24944494664669037,
|
| 663 |
+
"learning_rate": 5.0128553615248396e-05,
|
| 664 |
+
"loss": 0.1094522476196289,
|
| 665 |
+
"mean_token_accuracy": 0.962428919672966,
|
| 666 |
+
"num_tokens": 4143616.0,
|
| 667 |
+
"step": 2900
|
| 668 |
+
},
|
| 669 |
+
{
|
| 670 |
+
"entropy": 0.15613057143986225,
|
| 671 |
+
"epoch": 7.584298584298584,
|
| 672 |
+
"grad_norm": 0.1455036848783493,
|
| 673 |
+
"learning_rate": 4.549546899350423e-05,
|
| 674 |
+
"loss": 0.11092090606689453,
|
| 675 |
+
"mean_token_accuracy": 0.9620462411642074,
|
| 676 |
+
"num_tokens": 4215664.0,
|
| 677 |
+
"step": 2950
|
| 678 |
+
},
|
| 679 |
+
{
|
| 680 |
+
"entropy": 0.1631234459578991,
|
| 681 |
+
"epoch": 7.712998712998713,
|
| 682 |
+
"grad_norm": 0.2129560261964798,
|
| 683 |
+
"learning_rate": 4.104349015915862e-05,
|
| 684 |
+
"loss": 0.1141857624053955,
|
| 685 |
+
"mean_token_accuracy": 0.9613765001296997,
|
| 686 |
+
"num_tokens": 4286387.0,
|
| 687 |
+
"step": 3000
|
| 688 |
+
},
|
| 689 |
+
{
|
| 690 |
+
"entropy": 0.1680422095954418,
|
| 691 |
+
"epoch": 7.841698841698841,
|
| 692 |
+
"grad_norm": 0.24886097013950348,
|
| 693 |
+
"learning_rate": 3.678157768490372e-05,
|
| 694 |
+
"loss": 0.11513191223144531,
|
| 695 |
+
"mean_token_accuracy": 0.9615794748067856,
|
| 696 |
+
"num_tokens": 4355875.0,
|
| 697 |
+
"step": 3050
|
| 698 |
+
},
|
| 699 |
+
{
|
| 700 |
+
"entropy": 0.16354035697877406,
|
| 701 |
+
"epoch": 7.97039897039897,
|
| 702 |
+
"grad_norm": 0.27600204944610596,
|
| 703 |
+
"learning_rate": 3.27183095936714e-05,
|
| 704 |
+
"loss": 0.1118631362915039,
|
| 705 |
+
"mean_token_accuracy": 0.9623224419355393,
|
| 706 |
+
"num_tokens": 4427088.0,
|
| 707 |
+
"step": 3100
|
| 708 |
+
},
|
| 709 |
+
{
|
| 710 |
+
"epoch": 8.0,
|
| 711 |
+
"eval_entropy": 0.16586771200305409,
|
| 712 |
+
"eval_loss": 1.204746961593628,
|
| 713 |
+
"eval_mean_token_accuracy": 0.8229975042883882,
|
| 714 |
+
"eval_num_tokens": 4443944.0,
|
| 715 |
+
"eval_runtime": 162.0251,
|
| 716 |
+
"eval_samples_per_second": 9.536,
|
| 717 |
+
"eval_steps_per_second": 1.197,
|
| 718 |
+
"step": 3112
|
| 719 |
+
}
|
| 720 |
+
],
|
| 721 |
+
"logging_steps": 50,
|
| 722 |
+
"max_steps": 3890,
|
| 723 |
+
"num_input_tokens_seen": 0,
|
| 724 |
+
"num_train_epochs": 10,
|
| 725 |
+
"save_steps": 500,
|
| 726 |
+
"stateful_callbacks": {
|
| 727 |
+
"TrainerControl": {
|
| 728 |
+
"args": {
|
| 729 |
+
"should_epoch_stop": false,
|
| 730 |
+
"should_evaluate": false,
|
| 731 |
+
"should_log": false,
|
| 732 |
+
"should_save": true,
|
| 733 |
+
"should_training_stop": false
|
| 734 |
+
},
|
| 735 |
+
"attributes": {}
|
| 736 |
+
}
|
| 737 |
+
},
|
| 738 |
+
"total_flos": 7.445001770940518e+17,
|
| 739 |
+
"train_batch_size": 8,
|
| 740 |
+
"trial_name": null,
|
| 741 |
+
"trial_params": null
|
| 742 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/README.md
ADDED
|
@@ -0,0 +1,209 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen3-14B-Base
|
| 3 |
+
library_name: peft
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:Qwen/Qwen3-14B-Base
|
| 7 |
+
- lora
|
| 8 |
+
- sft
|
| 9 |
+
- transformers
|
| 10 |
+
- trl
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Model Card for Model ID
|
| 14 |
+
|
| 15 |
+
<!-- Provide a quick summary of what the model is/does. -->
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
## Model Details
|
| 20 |
+
|
| 21 |
+
### Model Description
|
| 22 |
+
|
| 23 |
+
<!-- Provide a longer summary of what this model is. -->
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
- **Developed by:** [More Information Needed]
|
| 28 |
+
- **Funded by [optional]:** [More Information Needed]
|
| 29 |
+
- **Shared by [optional]:** [More Information Needed]
|
| 30 |
+
- **Model type:** [More Information Needed]
|
| 31 |
+
- **Language(s) (NLP):** [More Information Needed]
|
| 32 |
+
- **License:** [More Information Needed]
|
| 33 |
+
- **Finetuned from model [optional]:** [More Information Needed]
|
| 34 |
+
|
| 35 |
+
### Model Sources [optional]
|
| 36 |
+
|
| 37 |
+
<!-- Provide the basic links for the model. -->
|
| 38 |
+
|
| 39 |
+
- **Repository:** [More Information Needed]
|
| 40 |
+
- **Paper [optional]:** [More Information Needed]
|
| 41 |
+
- **Demo [optional]:** [More Information Needed]
|
| 42 |
+
|
| 43 |
+
## Uses
|
| 44 |
+
|
| 45 |
+
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
| 46 |
+
|
| 47 |
+
### Direct Use
|
| 48 |
+
|
| 49 |
+
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
| 50 |
+
|
| 51 |
+
[More Information Needed]
|
| 52 |
+
|
| 53 |
+
### Downstream Use [optional]
|
| 54 |
+
|
| 55 |
+
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
| 56 |
+
|
| 57 |
+
[More Information Needed]
|
| 58 |
+
|
| 59 |
+
### Out-of-Scope Use
|
| 60 |
+
|
| 61 |
+
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
| 62 |
+
|
| 63 |
+
[More Information Needed]
|
| 64 |
+
|
| 65 |
+
## Bias, Risks, and Limitations
|
| 66 |
+
|
| 67 |
+
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
| 68 |
+
|
| 69 |
+
[More Information Needed]
|
| 70 |
+
|
| 71 |
+
### Recommendations
|
| 72 |
+
|
| 73 |
+
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
| 74 |
+
|
| 75 |
+
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
| 76 |
+
|
| 77 |
+
## How to Get Started with the Model
|
| 78 |
+
|
| 79 |
+
Use the code below to get started with the model.
|
| 80 |
+
|
| 81 |
+
[More Information Needed]
|
| 82 |
+
|
| 83 |
+
## Training Details
|
| 84 |
+
|
| 85 |
+
### Training Data
|
| 86 |
+
|
| 87 |
+
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
| 88 |
+
|
| 89 |
+
[More Information Needed]
|
| 90 |
+
|
| 91 |
+
### Training Procedure
|
| 92 |
+
|
| 93 |
+
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
| 94 |
+
|
| 95 |
+
#### Preprocessing [optional]
|
| 96 |
+
|
| 97 |
+
[More Information Needed]
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
#### Training Hyperparameters
|
| 101 |
+
|
| 102 |
+
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
| 103 |
+
|
| 104 |
+
#### Speeds, Sizes, Times [optional]
|
| 105 |
+
|
| 106 |
+
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
| 107 |
+
|
| 108 |
+
[More Information Needed]
|
| 109 |
+
|
| 110 |
+
## Evaluation
|
| 111 |
+
|
| 112 |
+
<!-- This section describes the evaluation protocols and provides the results. -->
|
| 113 |
+
|
| 114 |
+
### Testing Data, Factors & Metrics
|
| 115 |
+
|
| 116 |
+
#### Testing Data
|
| 117 |
+
|
| 118 |
+
<!-- This should link to a Dataset Card if possible. -->
|
| 119 |
+
|
| 120 |
+
[More Information Needed]
|
| 121 |
+
|
| 122 |
+
#### Factors
|
| 123 |
+
|
| 124 |
+
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
| 125 |
+
|
| 126 |
+
[More Information Needed]
|
| 127 |
+
|
| 128 |
+
#### Metrics
|
| 129 |
+
|
| 130 |
+
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
| 131 |
+
|
| 132 |
+
[More Information Needed]
|
| 133 |
+
|
| 134 |
+
### Results
|
| 135 |
+
|
| 136 |
+
[More Information Needed]
|
| 137 |
+
|
| 138 |
+
#### Summary
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
## Model Examination [optional]
|
| 143 |
+
|
| 144 |
+
<!-- Relevant interpretability work for the model goes here -->
|
| 145 |
+
|
| 146 |
+
[More Information Needed]
|
| 147 |
+
|
| 148 |
+
## Environmental Impact
|
| 149 |
+
|
| 150 |
+
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
| 151 |
+
|
| 152 |
+
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
| 153 |
+
|
| 154 |
+
- **Hardware Type:** [More Information Needed]
|
| 155 |
+
- **Hours used:** [More Information Needed]
|
| 156 |
+
- **Cloud Provider:** [More Information Needed]
|
| 157 |
+
- **Compute Region:** [More Information Needed]
|
| 158 |
+
- **Carbon Emitted:** [More Information Needed]
|
| 159 |
+
|
| 160 |
+
## Technical Specifications [optional]
|
| 161 |
+
|
| 162 |
+
### Model Architecture and Objective
|
| 163 |
+
|
| 164 |
+
[More Information Needed]
|
| 165 |
+
|
| 166 |
+
### Compute Infrastructure
|
| 167 |
+
|
| 168 |
+
[More Information Needed]
|
| 169 |
+
|
| 170 |
+
#### Hardware
|
| 171 |
+
|
| 172 |
+
[More Information Needed]
|
| 173 |
+
|
| 174 |
+
#### Software
|
| 175 |
+
|
| 176 |
+
[More Information Needed]
|
| 177 |
+
|
| 178 |
+
## Citation [optional]
|
| 179 |
+
|
| 180 |
+
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
| 181 |
+
|
| 182 |
+
**BibTeX:**
|
| 183 |
+
|
| 184 |
+
[More Information Needed]
|
| 185 |
+
|
| 186 |
+
**APA:**
|
| 187 |
+
|
| 188 |
+
[More Information Needed]
|
| 189 |
+
|
| 190 |
+
## Glossary [optional]
|
| 191 |
+
|
| 192 |
+
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
| 193 |
+
|
| 194 |
+
[More Information Needed]
|
| 195 |
+
|
| 196 |
+
## More Information [optional]
|
| 197 |
+
|
| 198 |
+
[More Information Needed]
|
| 199 |
+
|
| 200 |
+
## Model Card Authors [optional]
|
| 201 |
+
|
| 202 |
+
[More Information Needed]
|
| 203 |
+
|
| 204 |
+
## Model Card Contact
|
| 205 |
+
|
| 206 |
+
[More Information Needed]
|
| 207 |
+
### Framework versions
|
| 208 |
+
|
| 209 |
+
- PEFT 0.19.1
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/adapter_config.json
ADDED
|
@@ -0,0 +1,48 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "Qwen/Qwen3-14B-Base",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"layer_replication": null,
|
| 16 |
+
"layers_pattern": null,
|
| 17 |
+
"layers_to_transform": null,
|
| 18 |
+
"loftq_config": {},
|
| 19 |
+
"lora_alpha": 256,
|
| 20 |
+
"lora_bias": false,
|
| 21 |
+
"lora_dropout": 0.00237968804112545,
|
| 22 |
+
"lora_ga_config": null,
|
| 23 |
+
"megatron_config": null,
|
| 24 |
+
"megatron_core": "megatron.core",
|
| 25 |
+
"modules_to_save": null,
|
| 26 |
+
"peft_type": "LORA",
|
| 27 |
+
"peft_version": "0.19.1",
|
| 28 |
+
"qalora_group_size": 16,
|
| 29 |
+
"r": 128,
|
| 30 |
+
"rank_pattern": {},
|
| 31 |
+
"revision": null,
|
| 32 |
+
"target_modules": [
|
| 33 |
+
"gate_proj",
|
| 34 |
+
"q_proj",
|
| 35 |
+
"up_proj",
|
| 36 |
+
"o_proj",
|
| 37 |
+
"v_proj",
|
| 38 |
+
"k_proj",
|
| 39 |
+
"down_proj"
|
| 40 |
+
],
|
| 41 |
+
"target_parameters": null,
|
| 42 |
+
"task_type": "CAUSAL_LM",
|
| 43 |
+
"trainable_token_indices": null,
|
| 44 |
+
"use_bdlora": null,
|
| 45 |
+
"use_dora": false,
|
| 46 |
+
"use_qalora": false,
|
| 47 |
+
"use_rslora": false
|
| 48 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/chat_template.jinja
ADDED
|
@@ -0,0 +1,85 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- if tools %}
|
| 2 |
+
{{- '<|im_start|>system\n' }}
|
| 3 |
+
{%- if messages[0].role == 'system' %}
|
| 4 |
+
{{- messages[0].content + '\n\n' }}
|
| 5 |
+
{%- endif %}
|
| 6 |
+
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
| 7 |
+
{%- for tool in tools %}
|
| 8 |
+
{{- "\n" }}
|
| 9 |
+
{{- tool | tojson }}
|
| 10 |
+
{%- endfor %}
|
| 11 |
+
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
| 12 |
+
{%- else %}
|
| 13 |
+
{%- if messages[0].role == 'system' %}
|
| 14 |
+
{{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
|
| 15 |
+
{%- endif %}
|
| 16 |
+
{%- endif %}
|
| 17 |
+
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
| 18 |
+
{%- for message in messages[::-1] %}
|
| 19 |
+
{%- set index = (messages|length - 1) - loop.index0 %}
|
| 20 |
+
{%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
|
| 21 |
+
{%- set ns.multi_step_tool = false %}
|
| 22 |
+
{%- set ns.last_query_index = index %}
|
| 23 |
+
{%- endif %}
|
| 24 |
+
{%- endfor %}
|
| 25 |
+
{%- for message in messages %}
|
| 26 |
+
{%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
|
| 27 |
+
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
|
| 28 |
+
{%- elif message.role == "assistant" %}
|
| 29 |
+
{%- set content = message.content %}
|
| 30 |
+
{%- set reasoning_content = '' %}
|
| 31 |
+
{%- if message.reasoning_content is defined and message.reasoning_content is not none %}
|
| 32 |
+
{%- set reasoning_content = message.reasoning_content %}
|
| 33 |
+
{%- else %}
|
| 34 |
+
{%- if '</think>' in message.content %}
|
| 35 |
+
{%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
|
| 36 |
+
{%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
| 37 |
+
{%- endif %}
|
| 38 |
+
{%- endif %}
|
| 39 |
+
{%- if loop.index0 > ns.last_query_index %}
|
| 40 |
+
{%- if loop.last or (not loop.last and reasoning_content) %}
|
| 41 |
+
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
|
| 42 |
+
{%- else %}
|
| 43 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 44 |
+
{%- endif %}
|
| 45 |
+
{%- else %}
|
| 46 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 47 |
+
{%- endif %}
|
| 48 |
+
{%- if message.tool_calls %}
|
| 49 |
+
{%- for tool_call in message.tool_calls %}
|
| 50 |
+
{%- if (loop.first and content) or (not loop.first) %}
|
| 51 |
+
{{- '\n' }}
|
| 52 |
+
{%- endif %}
|
| 53 |
+
{%- if tool_call.function %}
|
| 54 |
+
{%- set tool_call = tool_call.function %}
|
| 55 |
+
{%- endif %}
|
| 56 |
+
{{- '<tool_call>\n{"name": "' }}
|
| 57 |
+
{{- tool_call.name }}
|
| 58 |
+
{{- '", "arguments": ' }}
|
| 59 |
+
{%- if tool_call.arguments is string %}
|
| 60 |
+
{{- tool_call.arguments }}
|
| 61 |
+
{%- else %}
|
| 62 |
+
{{- tool_call.arguments | tojson }}
|
| 63 |
+
{%- endif %}
|
| 64 |
+
{{- '}\n</tool_call>' }}
|
| 65 |
+
{%- endfor %}
|
| 66 |
+
{%- endif %}
|
| 67 |
+
{{- '<|im_end|>\n' }}
|
| 68 |
+
{%- elif message.role == "tool" %}
|
| 69 |
+
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
|
| 70 |
+
{{- '<|im_start|>user' }}
|
| 71 |
+
{%- endif %}
|
| 72 |
+
{{- '\n<tool_response>\n' }}
|
| 73 |
+
{{- message.content }}
|
| 74 |
+
{{- '\n</tool_response>' }}
|
| 75 |
+
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
| 76 |
+
{{- '<|im_end|>\n' }}
|
| 77 |
+
{%- endif %}
|
| 78 |
+
{%- endif %}
|
| 79 |
+
{%- endfor %}
|
| 80 |
+
{%- if add_generation_prompt %}
|
| 81 |
+
{{- '<|im_start|>assistant\n' }}
|
| 82 |
+
{%- if enable_thinking is defined and enable_thinking is false %}
|
| 83 |
+
{{- '<think>\n\n</think>\n\n' }}
|
| 84 |
+
{%- endif %}
|
| 85 |
+
{%- endif %}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/tokenizer_config.json
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"backend": "tokenizers",
|
| 4 |
+
"bos_token": null,
|
| 5 |
+
"clean_up_tokenization_spaces": false,
|
| 6 |
+
"eos_token": "<|endoftext|>",
|
| 7 |
+
"errors": "replace",
|
| 8 |
+
"extra_special_tokens": [
|
| 9 |
+
"<|im_start|>",
|
| 10 |
+
"<|im_end|>",
|
| 11 |
+
"<|object_ref_start|>",
|
| 12 |
+
"<|object_ref_end|>",
|
| 13 |
+
"<|box_start|>",
|
| 14 |
+
"<|box_end|>",
|
| 15 |
+
"<|quad_start|>",
|
| 16 |
+
"<|quad_end|>",
|
| 17 |
+
"<|vision_start|>",
|
| 18 |
+
"<|vision_end|>",
|
| 19 |
+
"<|vision_pad|>",
|
| 20 |
+
"<|image_pad|>",
|
| 21 |
+
"<|video_pad|>"
|
| 22 |
+
],
|
| 23 |
+
"is_local": false,
|
| 24 |
+
"model_max_length": 131072,
|
| 25 |
+
"pad_token": "<|endoftext|>",
|
| 26 |
+
"split_special_tokens": false,
|
| 27 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 28 |
+
"unk_token": null
|
| 29 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3501/trainer_state.json
ADDED
|
@@ -0,0 +1,833 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"best_global_step": null,
|
| 3 |
+
"best_metric": null,
|
| 4 |
+
"best_model_checkpoint": null,
|
| 5 |
+
"epoch": 9.0,
|
| 6 |
+
"eval_steps": 500,
|
| 7 |
+
"global_step": 3501,
|
| 8 |
+
"is_hyper_param_search": false,
|
| 9 |
+
"is_local_process_zero": true,
|
| 10 |
+
"is_world_process_zero": true,
|
| 11 |
+
"log_history": [
|
| 12 |
+
{
|
| 13 |
+
"entropy": 1.6110110306739807,
|
| 14 |
+
"epoch": 0.1287001287001287,
|
| 15 |
+
"grad_norm": 0.8611119389533997,
|
| 16 |
+
"learning_rate": 3.4130257962133866e-05,
|
| 17 |
+
"loss": 1.53956298828125,
|
| 18 |
+
"mean_token_accuracy": 0.6652342769503593,
|
| 19 |
+
"num_tokens": 73407.0,
|
| 20 |
+
"step": 50
|
| 21 |
+
},
|
| 22 |
+
{
|
| 23 |
+
"entropy": 0.8662820833921433,
|
| 24 |
+
"epoch": 0.2574002574002574,
|
| 25 |
+
"grad_norm": 0.7405035495758057,
|
| 26 |
+
"learning_rate": 6.895705180104598e-05,
|
| 27 |
+
"loss": 0.7929539489746094,
|
| 28 |
+
"mean_token_accuracy": 0.7786347842216492,
|
| 29 |
+
"num_tokens": 143994.0,
|
| 30 |
+
"step": 100
|
| 31 |
+
},
|
| 32 |
+
{
|
| 33 |
+
"entropy": 0.7912819278240204,
|
| 34 |
+
"epoch": 0.3861003861003861,
|
| 35 |
+
"grad_norm": 0.6105485558509827,
|
| 36 |
+
"learning_rate": 0.00010378384563995809,
|
| 37 |
+
"loss": 0.7236511993408203,
|
| 38 |
+
"mean_token_accuracy": 0.7925467795133591,
|
| 39 |
+
"num_tokens": 216171.0,
|
| 40 |
+
"step": 150
|
| 41 |
+
},
|
| 42 |
+
{
|
| 43 |
+
"entropy": 0.7777709531784057,
|
| 44 |
+
"epoch": 0.5148005148005148,
|
| 45 |
+
"grad_norm": 0.5781793594360352,
|
| 46 |
+
"learning_rate": 0.0001386106394788702,
|
| 47 |
+
"loss": 0.7024919891357422,
|
| 48 |
+
"mean_token_accuracy": 0.7969876372814179,
|
| 49 |
+
"num_tokens": 284702.0,
|
| 50 |
+
"step": 200
|
| 51 |
+
},
|
| 52 |
+
{
|
| 53 |
+
"entropy": 0.7632609683275223,
|
| 54 |
+
"epoch": 0.6435006435006435,
|
| 55 |
+
"grad_norm": 0.6553444862365723,
|
| 56 |
+
"learning_rate": 0.00017343743331778232,
|
| 57 |
+
"loss": 0.7000718688964844,
|
| 58 |
+
"mean_token_accuracy": 0.7994742071628571,
|
| 59 |
+
"num_tokens": 356393.0,
|
| 60 |
+
"step": 250
|
| 61 |
+
},
|
| 62 |
+
{
|
| 63 |
+
"entropy": 0.754165632724762,
|
| 64 |
+
"epoch": 0.7722007722007722,
|
| 65 |
+
"grad_norm": 0.634222149848938,
|
| 66 |
+
"learning_rate": 0.00020826422715669444,
|
| 67 |
+
"loss": 0.6909049987792969,
|
| 68 |
+
"mean_token_accuracy": 0.8016809666156769,
|
| 69 |
+
"num_tokens": 426916.0,
|
| 70 |
+
"step": 300
|
| 71 |
+
},
|
| 72 |
+
{
|
| 73 |
+
"entropy": 0.7530711203813553,
|
| 74 |
+
"epoch": 0.9009009009009009,
|
| 75 |
+
"grad_norm": 0.7020156383514404,
|
| 76 |
+
"learning_rate": 0.00024309102099560653,
|
| 77 |
+
"loss": 0.69554931640625,
|
| 78 |
+
"mean_token_accuracy": 0.8002807641029358,
|
| 79 |
+
"num_tokens": 499603.0,
|
| 80 |
+
"step": 350
|
| 81 |
+
},
|
| 82 |
+
{
|
| 83 |
+
"epoch": 1.0,
|
| 84 |
+
"eval_entropy": 0.5728044164242204,
|
| 85 |
+
"eval_loss": 0.6785851716995239,
|
| 86 |
+
"eval_mean_token_accuracy": 0.798433246993527,
|
| 87 |
+
"eval_num_tokens": 555493.0,
|
| 88 |
+
"eval_runtime": 162.4129,
|
| 89 |
+
"eval_samples_per_second": 9.513,
|
| 90 |
+
"eval_steps_per_second": 1.194,
|
| 91 |
+
"step": 389
|
| 92 |
+
},
|
| 93 |
+
{
|
| 94 |
+
"entropy": 0.7489313946829902,
|
| 95 |
+
"epoch": 1.0283140283140284,
|
| 96 |
+
"grad_norm": 0.7532815933227539,
|
| 97 |
+
"learning_rate": 0.0002709470016827303,
|
| 98 |
+
"loss": 0.6856581878662109,
|
| 99 |
+
"mean_token_accuracy": 0.8023027079273956,
|
| 100 |
+
"num_tokens": 570813.0,
|
| 101 |
+
"step": 400
|
| 102 |
+
},
|
| 103 |
+
{
|
| 104 |
+
"entropy": 0.7238409864902496,
|
| 105 |
+
"epoch": 1.157014157014157,
|
| 106 |
+
"grad_norm": 1.2016215324401855,
|
| 107 |
+
"learning_rate": 0.0002707561443541359,
|
| 108 |
+
"loss": 0.6699818420410156,
|
| 109 |
+
"mean_token_accuracy": 0.8070302194356919,
|
| 110 |
+
"num_tokens": 642956.0,
|
| 111 |
+
"step": 450
|
| 112 |
+
},
|
| 113 |
+
{
|
| 114 |
+
"entropy": 0.7388938587903976,
|
| 115 |
+
"epoch": 1.2857142857142856,
|
| 116 |
+
"grad_norm": 0.7279272079467773,
|
| 117 |
+
"learning_rate": 0.0002702930068622498,
|
| 118 |
+
"loss": 0.6728517150878907,
|
| 119 |
+
"mean_token_accuracy": 0.8049580943584442,
|
| 120 |
+
"num_tokens": 714498.0,
|
| 121 |
+
"step": 500
|
| 122 |
+
},
|
| 123 |
+
{
|
| 124 |
+
"entropy": 0.7284485149383545,
|
| 125 |
+
"epoch": 1.4144144144144144,
|
| 126 |
+
"grad_norm": 0.8563987016677856,
|
| 127 |
+
"learning_rate": 0.0002695585213716931,
|
| 128 |
+
"loss": 0.6657986450195312,
|
| 129 |
+
"mean_token_accuracy": 0.8085588800907135,
|
| 130 |
+
"num_tokens": 785595.0,
|
| 131 |
+
"step": 550
|
| 132 |
+
},
|
| 133 |
+
{
|
| 134 |
+
"entropy": 0.6960959500074386,
|
| 135 |
+
"epoch": 1.5431145431145432,
|
| 136 |
+
"grad_norm": 0.5145474672317505,
|
| 137 |
+
"learning_rate": 0.0002685541661937683,
|
| 138 |
+
"loss": 0.6358638763427734,
|
| 139 |
+
"mean_token_accuracy": 0.8131621342897415,
|
| 140 |
+
"num_tokens": 857551.0,
|
| 141 |
+
"step": 600
|
| 142 |
+
},
|
| 143 |
+
{
|
| 144 |
+
"entropy": 0.7220414417982102,
|
| 145 |
+
"epoch": 1.6718146718146718,
|
| 146 |
+
"grad_norm": 0.9381059408187866,
|
| 147 |
+
"learning_rate": 0.00026728196281103746,
|
| 148 |
+
"loss": 0.6531407928466797,
|
| 149 |
+
"mean_token_accuracy": 0.811619822382927,
|
| 150 |
+
"num_tokens": 926864.0,
|
| 151 |
+
"step": 650
|
| 152 |
+
},
|
| 153 |
+
{
|
| 154 |
+
"entropy": 0.6955459499359131,
|
| 155 |
+
"epoch": 1.8005148005148004,
|
| 156 |
+
"grad_norm": 0.6115108728408813,
|
| 157 |
+
"learning_rate": 0.0002657444718086503,
|
| 158 |
+
"loss": 0.6269588088989257,
|
| 159 |
+
"mean_token_accuracy": 0.8155373805761337,
|
| 160 |
+
"num_tokens": 998824.0,
|
| 161 |
+
"step": 700
|
| 162 |
+
},
|
| 163 |
+
{
|
| 164 |
+
"entropy": 0.6929812705516816,
|
| 165 |
+
"epoch": 1.9292149292149292,
|
| 166 |
+
"grad_norm": 0.749489426612854,
|
| 167 |
+
"learning_rate": 0.0002639447877206115,
|
| 168 |
+
"loss": 0.629054069519043,
|
| 169 |
+
"mean_token_accuracy": 0.8183021235466004,
|
| 170 |
+
"num_tokens": 1069332.0,
|
| 171 |
+
"step": 750
|
| 172 |
+
},
|
| 173 |
+
{
|
| 174 |
+
"epoch": 2.0,
|
| 175 |
+
"eval_entropy": 0.5833694102223387,
|
| 176 |
+
"eval_loss": 0.6328718662261963,
|
| 177 |
+
"eval_mean_token_accuracy": 0.8159044071571114,
|
| 178 |
+
"eval_num_tokens": 1110986.0,
|
| 179 |
+
"eval_runtime": 161.885,
|
| 180 |
+
"eval_samples_per_second": 9.544,
|
| 181 |
+
"eval_steps_per_second": 1.198,
|
| 182 |
+
"step": 778
|
| 183 |
+
},
|
| 184 |
+
{
|
| 185 |
+
"entropy": 0.647241060480927,
|
| 186 |
+
"epoch": 2.056628056628057,
|
| 187 |
+
"grad_norm": 0.7070767879486084,
|
| 188 |
+
"learning_rate": 0.00026188653280135975,
|
| 189 |
+
"loss": 0.5823922348022461,
|
| 190 |
+
"mean_token_accuracy": 0.8265365301960647,
|
| 191 |
+
"num_tokens": 1141195.0,
|
| 192 |
+
"step": 800
|
| 193 |
+
},
|
| 194 |
+
{
|
| 195 |
+
"entropy": 0.5995378407835961,
|
| 196 |
+
"epoch": 2.1853281853281854,
|
| 197 |
+
"grad_norm": 0.8090486526489258,
|
| 198 |
+
"learning_rate": 0.0002595738497351955,
|
| 199 |
+
"loss": 0.5325597763061524,
|
| 200 |
+
"mean_token_accuracy": 0.8369336777925491,
|
| 201 |
+
"num_tokens": 1210708.0,
|
| 202 |
+
"step": 850
|
| 203 |
+
},
|
| 204 |
+
{
|
| 205 |
+
"entropy": 0.6043145382404327,
|
| 206 |
+
"epoch": 2.314028314028314,
|
| 207 |
+
"grad_norm": 0.8279913067817688,
|
| 208 |
+
"learning_rate": 0.00025701139329823054,
|
| 209 |
+
"loss": 0.5414446258544922,
|
| 210 |
+
"mean_token_accuracy": 0.8361396533250809,
|
| 211 |
+
"num_tokens": 1283441.0,
|
| 212 |
+
"step": 900
|
| 213 |
+
},
|
| 214 |
+
{
|
| 215 |
+
"entropy": 0.5953224584460258,
|
| 216 |
+
"epoch": 2.4427284427284426,
|
| 217 |
+
"grad_norm": 0.6075023412704468,
|
| 218 |
+
"learning_rate": 0.00025420432098964183,
|
| 219 |
+
"loss": 0.536654167175293,
|
| 220 |
+
"mean_token_accuracy": 0.8340093129873276,
|
| 221 |
+
"num_tokens": 1356418.0,
|
| 222 |
+
"step": 950
|
| 223 |
+
},
|
| 224 |
+
{
|
| 225 |
+
"entropy": 0.5998479858040809,
|
| 226 |
+
"epoch": 2.571428571428571,
|
| 227 |
+
"grad_norm": 1.0311471223831177,
|
| 228 |
+
"learning_rate": 0.0002511582826510862,
|
| 229 |
+
"loss": 0.5372924423217773,
|
| 230 |
+
"mean_token_accuracy": 0.8366045409440994,
|
| 231 |
+
"num_tokens": 1427797.0,
|
| 232 |
+
"step": 1000
|
| 233 |
+
},
|
| 234 |
+
{
|
| 235 |
+
"entropy": 0.5980158120393753,
|
| 236 |
+
"epoch": 2.7001287001287,
|
| 237 |
+
"grad_norm": 0.5971426367759705,
|
| 238 |
+
"learning_rate": 0.0002478794090951689,
|
| 239 |
+
"loss": 0.5392885208129883,
|
| 240 |
+
"mean_token_accuracy": 0.8347727072238922,
|
| 241 |
+
"num_tokens": 1498082.0,
|
| 242 |
+
"step": 1050
|
| 243 |
+
},
|
| 244 |
+
{
|
| 245 |
+
"entropy": 0.5990680930018425,
|
| 246 |
+
"epoch": 2.828828828828829,
|
| 247 |
+
"grad_norm": 0.5662627220153809,
|
| 248 |
+
"learning_rate": 0.0002443742997658538,
|
| 249 |
+
"loss": 0.5360498428344727,
|
| 250 |
+
"mean_token_accuracy": 0.8371847170591354,
|
| 251 |
+
"num_tokens": 1568798.0,
|
| 252 |
+
"step": 1100
|
| 253 |
+
},
|
| 254 |
+
{
|
| 255 |
+
"entropy": 0.5957039377093315,
|
| 256 |
+
"epoch": 2.9575289575289574,
|
| 257 |
+
"grad_norm": 0.5043798685073853,
|
| 258 |
+
"learning_rate": 0.00024065000945565205,
|
| 259 |
+
"loss": 0.5342231369018555,
|
| 260 |
+
"mean_token_accuracy": 0.8380735236406326,
|
| 261 |
+
"num_tokens": 1643449.0,
|
| 262 |
+
"step": 1150
|
| 263 |
+
},
|
| 264 |
+
{
|
| 265 |
+
"epoch": 3.0,
|
| 266 |
+
"eval_entropy": 0.5346978819861854,
|
| 267 |
+
"eval_loss": 0.6255015134811401,
|
| 268 |
+
"eval_mean_token_accuracy": 0.8202014476368108,
|
| 269 |
+
"eval_num_tokens": 1666479.0,
|
| 270 |
+
"eval_runtime": 161.6098,
|
| 271 |
+
"eval_samples_per_second": 9.56,
|
| 272 |
+
"eval_steps_per_second": 1.2,
|
| 273 |
+
"step": 1167
|
| 274 |
+
},
|
| 275 |
+
{
|
| 276 |
+
"entropy": 0.5212209137401196,
|
| 277 |
+
"epoch": 3.0849420849420848,
|
| 278 |
+
"grad_norm": 0.7095440626144409,
|
| 279 |
+
"learning_rate": 0.00023671403410632178,
|
| 280 |
+
"loss": 0.45311901092529294,
|
| 281 |
+
"mean_token_accuracy": 0.856362871449403,
|
| 282 |
+
"num_tokens": 1713536.0,
|
| 283 |
+
"step": 1200
|
| 284 |
+
},
|
| 285 |
+
{
|
| 286 |
+
"entropy": 0.4701593083143234,
|
| 287 |
+
"epoch": 3.213642213642214,
|
| 288 |
+
"grad_norm": 0.6408083438873291,
|
| 289 |
+
"learning_rate": 0.0002325742957216607,
|
| 290 |
+
"loss": 0.39916397094726563,
|
| 291 |
+
"mean_token_accuracy": 0.8698061722517013,
|
| 292 |
+
"num_tokens": 1785609.0,
|
| 293 |
+
"step": 1250
|
| 294 |
+
},
|
| 295 |
+
{
|
| 296 |
+
"entropy": 0.4766591975092888,
|
| 297 |
+
"epoch": 3.3423423423423424,
|
| 298 |
+
"grad_norm": 0.6415093541145325,
|
| 299 |
+
"learning_rate": 0.0002282391264227552,
|
| 300 |
+
"loss": 0.4116698455810547,
|
| 301 |
+
"mean_token_accuracy": 0.8679435575008392,
|
| 302 |
+
"num_tokens": 1858651.0,
|
| 303 |
+
"step": 1300
|
| 304 |
+
},
|
| 305 |
+
{
|
| 306 |
+
"entropy": 0.4937947469949722,
|
| 307 |
+
"epoch": 3.471042471042471,
|
| 308 |
+
"grad_norm": 0.6549825072288513,
|
| 309 |
+
"learning_rate": 0.00022371725167778054,
|
| 310 |
+
"loss": 0.4296376037597656,
|
| 311 |
+
"mean_token_accuracy": 0.8609357953071595,
|
| 312 |
+
"num_tokens": 1928692.0,
|
| 313 |
+
"step": 1350
|
| 314 |
+
},
|
| 315 |
+
{
|
| 316 |
+
"entropy": 0.4920153194665909,
|
| 317 |
+
"epoch": 3.5997425997425996,
|
| 318 |
+
"grad_norm": 0.6452126502990723,
|
| 319 |
+
"learning_rate": 0.00021901777274010406,
|
| 320 |
+
"loss": 0.4307489013671875,
|
| 321 |
+
"mean_token_accuracy": 0.8606827831268311,
|
| 322 |
+
"num_tokens": 1998668.0,
|
| 323 |
+
"step": 1400
|
| 324 |
+
},
|
| 325 |
+
{
|
| 326 |
+
"entropy": 0.490042342543602,
|
| 327 |
+
"epoch": 3.7284427284427286,
|
| 328 |
+
"grad_norm": 0.5727734565734863,
|
| 329 |
+
"learning_rate": 0.0002141501483300395,
|
| 330 |
+
"loss": 0.4295254135131836,
|
| 331 |
+
"mean_token_accuracy": 0.8616545403003693,
|
| 332 |
+
"num_tokens": 2072809.0,
|
| 333 |
+
"step": 1450
|
| 334 |
+
},
|
| 335 |
+
{
|
| 336 |
+
"entropy": 0.49857193052768706,
|
| 337 |
+
"epoch": 3.857142857142857,
|
| 338 |
+
"grad_norm": 0.7732954025268555,
|
| 339 |
+
"learning_rate": 0.00020912417559712133,
|
| 340 |
+
"loss": 0.4289303207397461,
|
| 341 |
+
"mean_token_accuracy": 0.8616443765163422,
|
| 342 |
+
"num_tokens": 2142475.0,
|
| 343 |
+
"step": 1500
|
| 344 |
+
},
|
| 345 |
+
{
|
| 346 |
+
"entropy": 0.4746784272789955,
|
| 347 |
+
"epoch": 3.985842985842986,
|
| 348 |
+
"grad_norm": 0.5791187882423401,
|
| 349 |
+
"learning_rate": 0.00020394997040121726,
|
| 350 |
+
"loss": 0.4180263900756836,
|
| 351 |
+
"mean_token_accuracy": 0.866080379486084,
|
| 352 |
+
"num_tokens": 2214044.0,
|
| 353 |
+
"step": 1550
|
| 354 |
+
},
|
| 355 |
+
{
|
| 356 |
+
"epoch": 4.0,
|
| 357 |
+
"eval_entropy": 0.4793926059585257,
|
| 358 |
+
"eval_loss": 0.627627968788147,
|
| 359 |
+
"eval_mean_token_accuracy": 0.8233870095813397,
|
| 360 |
+
"eval_num_tokens": 2221972.0,
|
| 361 |
+
"eval_runtime": 162.0225,
|
| 362 |
+
"eval_samples_per_second": 9.536,
|
| 363 |
+
"eval_steps_per_second": 1.197,
|
| 364 |
+
"step": 1556
|
| 365 |
+
},
|
| 366 |
+
{
|
| 367 |
+
"entropy": 0.38195489000792454,
|
| 368 |
+
"epoch": 4.113256113256114,
|
| 369 |
+
"grad_norm": 0.6002617478370667,
|
| 370 |
+
"learning_rate": 0.0001986379469521669,
|
| 371 |
+
"loss": 0.30819049835205076,
|
| 372 |
+
"mean_token_accuracy": 0.8977848634575353,
|
| 373 |
+
"num_tokens": 2282164.0,
|
| 374 |
+
"step": 1600
|
| 375 |
+
},
|
| 376 |
+
{
|
| 377 |
+
"entropy": 0.3655787402391434,
|
| 378 |
+
"epoch": 4.241956241956242,
|
| 379 |
+
"grad_norm": 0.7100041508674622,
|
| 380 |
+
"learning_rate": 0.00019319879684892634,
|
| 381 |
+
"loss": 0.29959835052490236,
|
| 382 |
+
"mean_token_accuracy": 0.8991208010911942,
|
| 383 |
+
"num_tokens": 2353213.0,
|
| 384 |
+
"step": 1650
|
| 385 |
+
},
|
| 386 |
+
{
|
| 387 |
+
"entropy": 0.3821141055226326,
|
| 388 |
+
"epoch": 4.370656370656371,
|
| 389 |
+
"grad_norm": 0.5848307013511658,
|
| 390 |
+
"learning_rate": 0.00018764346756040715,
|
| 391 |
+
"loss": 0.313802490234375,
|
| 392 |
+
"mean_token_accuracy": 0.895167955160141,
|
| 393 |
+
"num_tokens": 2425068.0,
|
| 394 |
+
"step": 1700
|
| 395 |
+
},
|
| 396 |
+
{
|
| 397 |
+
"entropy": 0.37083797007799146,
|
| 398 |
+
"epoch": 4.499356499356499,
|
| 399 |
+
"grad_norm": 0.6447024941444397,
|
| 400 |
+
"learning_rate": 0.00018198314039132143,
|
| 401 |
+
"loss": 0.30583988189697264,
|
| 402 |
+
"mean_token_accuracy": 0.8961733293533325,
|
| 403 |
+
"num_tokens": 2498321.0,
|
| 404 |
+
"step": 1750
|
| 405 |
+
},
|
| 406 |
+
{
|
| 407 |
+
"entropy": 0.3791545969247818,
|
| 408 |
+
"epoch": 4.628056628056628,
|
| 409 |
+
"grad_norm": 0.6575382351875305,
|
| 410 |
+
"learning_rate": 0.00017622920797738184,
|
| 411 |
+
"loss": 0.3088031005859375,
|
| 412 |
+
"mean_token_accuracy": 0.8960050916671753,
|
| 413 |
+
"num_tokens": 2570321.0,
|
| 414 |
+
"step": 1800
|
| 415 |
+
},
|
| 416 |
+
{
|
| 417 |
+
"entropy": 0.3946831756830215,
|
| 418 |
+
"epoch": 4.756756756756757,
|
| 419 |
+
"grad_norm": 0.5351552963256836,
|
| 420 |
+
"learning_rate": 0.00017039325135515207,
|
| 421 |
+
"loss": 0.3229162979125977,
|
| 422 |
+
"mean_token_accuracy": 0.8920552498102188,
|
| 423 |
+
"num_tokens": 2642851.0,
|
| 424 |
+
"step": 1850
|
| 425 |
+
},
|
| 426 |
+
{
|
| 427 |
+
"entropy": 0.37298239797353744,
|
| 428 |
+
"epoch": 4.885456885456885,
|
| 429 |
+
"grad_norm": 0.7624587416648865,
|
| 430 |
+
"learning_rate": 0.00016448701665269964,
|
| 431 |
+
"loss": 0.3067934799194336,
|
| 432 |
+
"mean_token_accuracy": 0.8951873427629471,
|
| 433 |
+
"num_tokens": 2715629.0,
|
| 434 |
+
"step": 1900
|
| 435 |
+
},
|
| 436 |
+
{
|
| 437 |
+
"epoch": 5.0,
|
| 438 |
+
"eval_entropy": 0.3738150204887095,
|
| 439 |
+
"eval_loss": 0.7131896615028381,
|
| 440 |
+
"eval_mean_token_accuracy": 0.8207752468045225,
|
| 441 |
+
"eval_num_tokens": 2777465.0,
|
| 442 |
+
"eval_runtime": 162.1244,
|
| 443 |
+
"eval_samples_per_second": 9.53,
|
| 444 |
+
"eval_steps_per_second": 1.197,
|
| 445 |
+
"step": 1945
|
| 446 |
+
},
|
| 447 |
+
{
|
| 448 |
+
"entropy": 0.377820266617669,
|
| 449 |
+
"epoch": 5.012870012870013,
|
| 450 |
+
"grad_norm": 0.41990190744400024,
|
| 451 |
+
"learning_rate": 0.00015852239144796624,
|
| 452 |
+
"loss": 0.3058685111999512,
|
| 453 |
+
"mean_token_accuracy": 0.8964343480389527,
|
| 454 |
+
"num_tokens": 2784896.0,
|
| 455 |
+
"step": 1950
|
| 456 |
+
},
|
| 457 |
+
{
|
| 458 |
+
"entropy": 0.2714502356946468,
|
| 459 |
+
"epoch": 5.141570141570142,
|
| 460 |
+
"grad_norm": 0.422568678855896,
|
| 461 |
+
"learning_rate": 0.00015251138084243995,
|
| 462 |
+
"loss": 0.2093442153930664,
|
| 463 |
+
"mean_token_accuracy": 0.9311346983909607,
|
| 464 |
+
"num_tokens": 2854374.0,
|
| 465 |
+
"step": 2000
|
| 466 |
+
},
|
| 467 |
+
{
|
| 468 |
+
"entropy": 0.268475965410471,
|
| 469 |
+
"epoch": 5.27027027027027,
|
| 470 |
+
"grad_norm": 0.6637414693832397,
|
| 471 |
+
"learning_rate": 0.0001464660832982852,
|
| 472 |
+
"loss": 0.20736080169677734,
|
| 473 |
+
"mean_token_accuracy": 0.9289199805259705,
|
| 474 |
+
"num_tokens": 2927362.0,
|
| 475 |
+
"step": 2050
|
| 476 |
+
},
|
| 477 |
+
{
|
| 478 |
+
"entropy": 0.2644876340031624,
|
| 479 |
+
"epoch": 5.398970398970399,
|
| 480 |
+
"grad_norm": 0.47317707538604736,
|
| 481 |
+
"learning_rate": 0.00014039866628756467,
|
| 482 |
+
"loss": 0.20464908599853515,
|
| 483 |
+
"mean_token_accuracy": 0.9300856202840805,
|
| 484 |
+
"num_tokens": 3000143.0,
|
| 485 |
+
"step": 2100
|
| 486 |
+
},
|
| 487 |
+
{
|
| 488 |
+
"entropy": 0.2675253136456013,
|
| 489 |
+
"epoch": 5.527670527670527,
|
| 490 |
+
"grad_norm": 0.5253982543945312,
|
| 491 |
+
"learning_rate": 0.00013432134180256338,
|
| 492 |
+
"loss": 0.21154335021972656,
|
| 493 |
+
"mean_token_accuracy": 0.9283734840154648,
|
| 494 |
+
"num_tokens": 3072561.0,
|
| 495 |
+
"step": 2150
|
| 496 |
+
},
|
| 497 |
+
{
|
| 498 |
+
"entropy": 0.27213907435536383,
|
| 499 |
+
"epoch": 5.656370656370656,
|
| 500 |
+
"grad_norm": 0.46738553047180176,
|
| 501 |
+
"learning_rate": 0.00012824634177650664,
|
| 502 |
+
"loss": 0.21339216232299804,
|
| 503 |
+
"mean_token_accuracy": 0.9272083270549775,
|
| 504 |
+
"num_tokens": 3144831.0,
|
| 505 |
+
"step": 2200
|
| 506 |
+
},
|
| 507 |
+
{
|
| 508 |
+
"entropy": 0.2785488124191761,
|
| 509 |
+
"epoch": 5.785070785070785,
|
| 510 |
+
"grad_norm": 0.4469502866268158,
|
| 511 |
+
"learning_rate": 0.00012218589346414205,
|
| 512 |
+
"loss": 0.21601097106933595,
|
| 513 |
+
"mean_token_accuracy": 0.9255663657188415,
|
| 514 |
+
"num_tokens": 3215960.0,
|
| 515 |
+
"step": 2250
|
| 516 |
+
},
|
| 517 |
+
{
|
| 518 |
+
"entropy": 0.2699935150146484,
|
| 519 |
+
"epoch": 5.913770913770914,
|
| 520 |
+
"grad_norm": 0.7359778881072998,
|
| 521 |
+
"learning_rate": 0.00011615219483173828,
|
| 522 |
+
"loss": 0.20725584030151367,
|
| 523 |
+
"mean_token_accuracy": 0.9286630594730377,
|
| 524 |
+
"num_tokens": 3287499.0,
|
| 525 |
+
"step": 2300
|
| 526 |
+
},
|
| 527 |
+
{
|
| 528 |
+
"epoch": 6.0,
|
| 529 |
+
"eval_entropy": 0.26417383682174783,
|
| 530 |
+
"eval_loss": 0.8880229592323303,
|
| 531 |
+
"eval_mean_token_accuracy": 0.8159987201395723,
|
| 532 |
+
"eval_num_tokens": 3332958.0,
|
| 533 |
+
"eval_runtime": 162.0991,
|
| 534 |
+
"eval_samples_per_second": 9.531,
|
| 535 |
+
"eval_steps_per_second": 1.197,
|
| 536 |
+
"step": 2334
|
| 537 |
+
},
|
| 538 |
+
{
|
| 539 |
+
"entropy": 0.24814540704693458,
|
| 540 |
+
"epoch": 6.041184041184041,
|
| 541 |
+
"grad_norm": 0.4953760802745819,
|
| 542 |
+
"learning_rate": 0.00011015739000603316,
|
| 543 |
+
"loss": 0.18749794006347656,
|
| 544 |
+
"mean_token_accuracy": 0.9370789509831052,
|
| 545 |
+
"num_tokens": 3356879.0,
|
| 546 |
+
"step": 2350
|
| 547 |
+
},
|
| 548 |
+
{
|
| 549 |
+
"entropy": 0.19976271741092205,
|
| 550 |
+
"epoch": 6.1698841698841695,
|
| 551 |
+
"grad_norm": 0.4834803342819214,
|
| 552 |
+
"learning_rate": 0.00010421354483154553,
|
| 553 |
+
"loss": 0.14283526420593262,
|
| 554 |
+
"mean_token_accuracy": 0.9521516615152359,
|
| 555 |
+
"num_tokens": 3427587.0,
|
| 556 |
+
"step": 2400
|
| 557 |
+
},
|
| 558 |
+
{
|
| 559 |
+
"entropy": 0.2060488449037075,
|
| 560 |
+
"epoch": 6.298584298584299,
|
| 561 |
+
"grad_norm": 0.4888673722743988,
|
| 562 |
+
"learning_rate": 9.8332622585447e-05,
|
| 563 |
+
"loss": 0.14414511680603026,
|
| 564 |
+
"mean_token_accuracy": 0.9510996866226197,
|
| 565 |
+
"num_tokens": 3498688.0,
|
| 566 |
+
"step": 2450
|
| 567 |
+
},
|
| 568 |
+
{
|
| 569 |
+
"entropy": 0.2059111550450325,
|
| 570 |
+
"epoch": 6.427284427284428,
|
| 571 |
+
"grad_norm": 0.4064404368400574,
|
| 572 |
+
"learning_rate": 9.252645989887253e-05,
|
| 573 |
+
"loss": 0.14820143699645996,
|
| 574 |
+
"mean_token_accuracy": 0.9507584601640702,
|
| 575 |
+
"num_tokens": 3566137.0,
|
| 576 |
+
"step": 2500
|
| 577 |
+
},
|
| 578 |
+
{
|
| 579 |
+
"entropy": 0.19700154662132263,
|
| 580 |
+
"epoch": 6.555984555984556,
|
| 581 |
+
"grad_norm": 0.467965304851532,
|
| 582 |
+
"learning_rate": 8.680674293313417e-05,
|
| 583 |
+
"loss": 0.14303470611572267,
|
| 584 |
+
"mean_token_accuracy": 0.9515972435474396,
|
| 585 |
+
"num_tokens": 3639573.0,
|
| 586 |
+
"step": 2550
|
| 587 |
+
},
|
| 588 |
+
{
|
| 589 |
+
"entropy": 0.20180423602461814,
|
| 590 |
+
"epoch": 6.684684684684685,
|
| 591 |
+
"grad_norm": 0.36836138367652893,
|
| 592 |
+
"learning_rate": 8.118498385878736e-05,
|
| 593 |
+
"loss": 0.14280882835388184,
|
| 594 |
+
"mean_token_accuracy": 0.9515993863344192,
|
| 595 |
+
"num_tokens": 3710433.0,
|
| 596 |
+
"step": 2600
|
| 597 |
+
},
|
| 598 |
+
{
|
| 599 |
+
"entropy": 0.20024395987391472,
|
| 600 |
+
"epoch": 6.813384813384813,
|
| 601 |
+
"grad_norm": 0.38375866413116455,
|
| 602 |
+
"learning_rate": 7.567249768489171e-05,
|
| 603 |
+
"loss": 0.1427844524383545,
|
| 604 |
+
"mean_token_accuracy": 0.9524166631698608,
|
| 605 |
+
"num_tokens": 3781550.0,
|
| 606 |
+
"step": 2650
|
| 607 |
+
},
|
| 608 |
+
{
|
| 609 |
+
"entropy": 0.19561587080359458,
|
| 610 |
+
"epoch": 6.942084942084942,
|
| 611 |
+
"grad_norm": 0.41185441613197327,
|
| 612 |
+
"learning_rate": 7.028037948510187e-05,
|
| 613 |
+
"loss": 0.13993803024291993,
|
| 614 |
+
"mean_token_accuracy": 0.9522478264570237,
|
| 615 |
+
"num_tokens": 3854952.0,
|
| 616 |
+
"step": 2700
|
| 617 |
+
},
|
| 618 |
+
{
|
| 619 |
+
"epoch": 7.0,
|
| 620 |
+
"eval_entropy": 0.19501976062034823,
|
| 621 |
+
"eval_loss": 1.0653952360153198,
|
| 622 |
+
"eval_mean_token_accuracy": 0.8205490803595671,
|
| 623 |
+
"eval_num_tokens": 3888451.0,
|
| 624 |
+
"eval_runtime": 161.8533,
|
| 625 |
+
"eval_samples_per_second": 9.546,
|
| 626 |
+
"eval_steps_per_second": 1.199,
|
| 627 |
+
"step": 2723
|
| 628 |
+
},
|
| 629 |
+
{
|
| 630 |
+
"entropy": 0.17789882526855277,
|
| 631 |
+
"epoch": 7.06949806949807,
|
| 632 |
+
"grad_norm": 0.41413992643356323,
|
| 633 |
+
"learning_rate": 6.50194820664261e-05,
|
| 634 |
+
"loss": 0.12078390121459961,
|
| 635 |
+
"mean_token_accuracy": 0.9589925727458916,
|
| 636 |
+
"num_tokens": 3928354.0,
|
| 637 |
+
"step": 2750
|
| 638 |
+
},
|
| 639 |
+
{
|
| 640 |
+
"entropy": 0.16781829454004765,
|
| 641 |
+
"epoch": 7.198198198198198,
|
| 642 |
+
"grad_norm": 0.25806066393852234,
|
| 643 |
+
"learning_rate": 5.990039412559906e-05,
|
| 644 |
+
"loss": 0.10963023185729981,
|
| 645 |
+
"mean_token_accuracy": 0.9617267113924026,
|
| 646 |
+
"num_tokens": 4000113.0,
|
| 647 |
+
"step": 2800
|
| 648 |
+
},
|
| 649 |
+
{
|
| 650 |
+
"entropy": 0.1649068508297205,
|
| 651 |
+
"epoch": 7.326898326898327,
|
| 652 |
+
"grad_norm": 0.27411890029907227,
|
| 653 |
+
"learning_rate": 5.493341893703393e-05,
|
| 654 |
+
"loss": 0.11152458190917969,
|
| 655 |
+
"mean_token_accuracy": 0.9620639663934708,
|
| 656 |
+
"num_tokens": 4071032.0,
|
| 657 |
+
"step": 2850
|
| 658 |
+
},
|
| 659 |
+
{
|
| 660 |
+
"entropy": 0.161333369910717,
|
| 661 |
+
"epoch": 7.455598455598455,
|
| 662 |
+
"grad_norm": 0.24944494664669037,
|
| 663 |
+
"learning_rate": 5.0128553615248396e-05,
|
| 664 |
+
"loss": 0.1094522476196289,
|
| 665 |
+
"mean_token_accuracy": 0.962428919672966,
|
| 666 |
+
"num_tokens": 4143616.0,
|
| 667 |
+
"step": 2900
|
| 668 |
+
},
|
| 669 |
+
{
|
| 670 |
+
"entropy": 0.15613057143986225,
|
| 671 |
+
"epoch": 7.584298584298584,
|
| 672 |
+
"grad_norm": 0.1455036848783493,
|
| 673 |
+
"learning_rate": 4.549546899350423e-05,
|
| 674 |
+
"loss": 0.11092090606689453,
|
| 675 |
+
"mean_token_accuracy": 0.9620462411642074,
|
| 676 |
+
"num_tokens": 4215664.0,
|
| 677 |
+
"step": 2950
|
| 678 |
+
},
|
| 679 |
+
{
|
| 680 |
+
"entropy": 0.1631234459578991,
|
| 681 |
+
"epoch": 7.712998712998713,
|
| 682 |
+
"grad_norm": 0.2129560261964798,
|
| 683 |
+
"learning_rate": 4.104349015915862e-05,
|
| 684 |
+
"loss": 0.1141857624053955,
|
| 685 |
+
"mean_token_accuracy": 0.9613765001296997,
|
| 686 |
+
"num_tokens": 4286387.0,
|
| 687 |
+
"step": 3000
|
| 688 |
+
},
|
| 689 |
+
{
|
| 690 |
+
"entropy": 0.1680422095954418,
|
| 691 |
+
"epoch": 7.841698841698841,
|
| 692 |
+
"grad_norm": 0.24886097013950348,
|
| 693 |
+
"learning_rate": 3.678157768490372e-05,
|
| 694 |
+
"loss": 0.11513191223144531,
|
| 695 |
+
"mean_token_accuracy": 0.9615794748067856,
|
| 696 |
+
"num_tokens": 4355875.0,
|
| 697 |
+
"step": 3050
|
| 698 |
+
},
|
| 699 |
+
{
|
| 700 |
+
"entropy": 0.16354035697877406,
|
| 701 |
+
"epoch": 7.97039897039897,
|
| 702 |
+
"grad_norm": 0.27600204944610596,
|
| 703 |
+
"learning_rate": 3.27183095936714e-05,
|
| 704 |
+
"loss": 0.1118631362915039,
|
| 705 |
+
"mean_token_accuracy": 0.9623224419355393,
|
| 706 |
+
"num_tokens": 4427088.0,
|
| 707 |
+
"step": 3100
|
| 708 |
+
},
|
| 709 |
+
{
|
| 710 |
+
"epoch": 8.0,
|
| 711 |
+
"eval_entropy": 0.16586771200305409,
|
| 712 |
+
"eval_loss": 1.204746961593628,
|
| 713 |
+
"eval_mean_token_accuracy": 0.8229975042883882,
|
| 714 |
+
"eval_num_tokens": 4443944.0,
|
| 715 |
+
"eval_runtime": 162.0251,
|
| 716 |
+
"eval_samples_per_second": 9.536,
|
| 717 |
+
"eval_steps_per_second": 1.197,
|
| 718 |
+
"step": 3112
|
| 719 |
+
},
|
| 720 |
+
{
|
| 721 |
+
"entropy": 0.1515902608934075,
|
| 722 |
+
"epoch": 8.097812097812097,
|
| 723 |
+
"grad_norm": 0.14122211933135986,
|
| 724 |
+
"learning_rate": 2.88618640935022e-05,
|
| 725 |
+
"loss": 0.09900871276855469,
|
| 726 |
+
"mean_token_accuracy": 0.9665110737386376,
|
| 727 |
+
"num_tokens": 4497867.0,
|
| 728 |
+
"step": 3150
|
| 729 |
+
},
|
| 730 |
+
{
|
| 731 |
+
"entropy": 0.14593622356653213,
|
| 732 |
+
"epoch": 8.226512226512227,
|
| 733 |
+
"grad_norm": 0.20527532696723938,
|
| 734 |
+
"learning_rate": 2.5220003117128462e-05,
|
| 735 |
+
"loss": 0.09842084884643555,
|
| 736 |
+
"mean_token_accuracy": 0.9655911487340927,
|
| 737 |
+
"num_tokens": 4568534.0,
|
| 738 |
+
"step": 3200
|
| 739 |
+
},
|
| 740 |
+
{
|
| 741 |
+
"entropy": 0.1482392605394125,
|
| 742 |
+
"epoch": 8.355212355212355,
|
| 743 |
+
"grad_norm": 0.13207173347473145,
|
| 744 |
+
"learning_rate": 2.1800056699401584e-05,
|
| 745 |
+
"loss": 0.09551989555358886,
|
| 746 |
+
"mean_token_accuracy": 0.9650829958915711,
|
| 747 |
+
"num_tokens": 4642530.0,
|
| 748 |
+
"step": 3250
|
| 749 |
+
},
|
| 750 |
+
{
|
| 751 |
+
"entropy": 0.15284131653606892,
|
| 752 |
+
"epoch": 8.483912483912484,
|
| 753 |
+
"grad_norm": 0.1777282953262329,
|
| 754 |
+
"learning_rate": 1.860890822400777e-05,
|
| 755 |
+
"loss": 0.10169261932373047,
|
| 756 |
+
"mean_token_accuracy": 0.9635212075710297,
|
| 757 |
+
"num_tokens": 4711573.0,
|
| 758 |
+
"step": 3300
|
| 759 |
+
},
|
| 760 |
+
{
|
| 761 |
+
"entropy": 0.15095721945166587,
|
| 762 |
+
"epoch": 8.612612612612612,
|
| 763 |
+
"grad_norm": 0.14988408982753754,
|
| 764 |
+
"learning_rate": 1.5652980569165692e-05,
|
| 765 |
+
"loss": 0.10045011520385742,
|
| 766 |
+
"mean_token_accuracy": 0.96439110994339,
|
| 767 |
+
"num_tokens": 4782666.0,
|
| 768 |
+
"step": 3350
|
| 769 |
+
},
|
| 770 |
+
{
|
| 771 |
+
"entropy": 0.15411154814064504,
|
| 772 |
+
"epoch": 8.741312741312742,
|
| 773 |
+
"grad_norm": 0.14055995643138885,
|
| 774 |
+
"learning_rate": 1.2938223180191691e-05,
|
| 775 |
+
"loss": 0.1034860897064209,
|
| 776 |
+
"mean_token_accuracy": 0.963447842001915,
|
| 777 |
+
"num_tokens": 4852180.0,
|
| 778 |
+
"step": 3400
|
| 779 |
+
},
|
| 780 |
+
{
|
| 781 |
+
"entropy": 0.14739766091108322,
|
| 782 |
+
"epoch": 8.87001287001287,
|
| 783 |
+
"grad_norm": 0.15041407942771912,
|
| 784 |
+
"learning_rate": 1.0470100094950792e-05,
|
| 785 |
+
"loss": 0.09690508842468262,
|
| 786 |
+
"mean_token_accuracy": 0.96561603307724,
|
| 787 |
+
"num_tokens": 4926402.0,
|
| 788 |
+
"step": 3450
|
| 789 |
+
},
|
| 790 |
+
{
|
| 791 |
+
"entropy": 0.1498453303426504,
|
| 792 |
+
"epoch": 8.998712998712998,
|
| 793 |
+
"grad_norm": 0.1293368935585022,
|
| 794 |
+
"learning_rate": 8.253578946296125e-06,
|
| 795 |
+
"loss": 0.09874271392822266,
|
| 796 |
+
"mean_token_accuracy": 0.9647125631570816,
|
| 797 |
+
"num_tokens": 4998841.0,
|
| 798 |
+
"step": 3500
|
| 799 |
+
},
|
| 800 |
+
{
|
| 801 |
+
"epoch": 9.0,
|
| 802 |
+
"eval_entropy": 0.151532097729211,
|
| 803 |
+
"eval_loss": 1.302620768547058,
|
| 804 |
+
"eval_mean_token_accuracy": 0.8230925079473516,
|
| 805 |
+
"eval_num_tokens": 4999437.0,
|
| 806 |
+
"eval_runtime": 161.5816,
|
| 807 |
+
"eval_samples_per_second": 9.562,
|
| 808 |
+
"eval_steps_per_second": 1.201,
|
| 809 |
+
"step": 3501
|
| 810 |
+
}
|
| 811 |
+
],
|
| 812 |
+
"logging_steps": 50,
|
| 813 |
+
"max_steps": 3890,
|
| 814 |
+
"num_input_tokens_seen": 0,
|
| 815 |
+
"num_train_epochs": 10,
|
| 816 |
+
"save_steps": 500,
|
| 817 |
+
"stateful_callbacks": {
|
| 818 |
+
"TrainerControl": {
|
| 819 |
+
"args": {
|
| 820 |
+
"should_epoch_stop": false,
|
| 821 |
+
"should_evaluate": false,
|
| 822 |
+
"should_log": false,
|
| 823 |
+
"should_save": true,
|
| 824 |
+
"should_training_stop": false
|
| 825 |
+
},
|
| 826 |
+
"attributes": {}
|
| 827 |
+
}
|
| 828 |
+
},
|
| 829 |
+
"total_flos": 8.377780321673011e+17,
|
| 830 |
+
"train_batch_size": 8,
|
| 831 |
+
"trial_name": null,
|
| 832 |
+
"trial_params": null
|
| 833 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/README.md
ADDED
|
@@ -0,0 +1,209 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen3-14B-Base
|
| 3 |
+
library_name: peft
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:Qwen/Qwen3-14B-Base
|
| 7 |
+
- lora
|
| 8 |
+
- sft
|
| 9 |
+
- transformers
|
| 10 |
+
- trl
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Model Card for Model ID
|
| 14 |
+
|
| 15 |
+
<!-- Provide a quick summary of what the model is/does. -->
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
## Model Details
|
| 20 |
+
|
| 21 |
+
### Model Description
|
| 22 |
+
|
| 23 |
+
<!-- Provide a longer summary of what this model is. -->
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
- **Developed by:** [More Information Needed]
|
| 28 |
+
- **Funded by [optional]:** [More Information Needed]
|
| 29 |
+
- **Shared by [optional]:** [More Information Needed]
|
| 30 |
+
- **Model type:** [More Information Needed]
|
| 31 |
+
- **Language(s) (NLP):** [More Information Needed]
|
| 32 |
+
- **License:** [More Information Needed]
|
| 33 |
+
- **Finetuned from model [optional]:** [More Information Needed]
|
| 34 |
+
|
| 35 |
+
### Model Sources [optional]
|
| 36 |
+
|
| 37 |
+
<!-- Provide the basic links for the model. -->
|
| 38 |
+
|
| 39 |
+
- **Repository:** [More Information Needed]
|
| 40 |
+
- **Paper [optional]:** [More Information Needed]
|
| 41 |
+
- **Demo [optional]:** [More Information Needed]
|
| 42 |
+
|
| 43 |
+
## Uses
|
| 44 |
+
|
| 45 |
+
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
| 46 |
+
|
| 47 |
+
### Direct Use
|
| 48 |
+
|
| 49 |
+
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
| 50 |
+
|
| 51 |
+
[More Information Needed]
|
| 52 |
+
|
| 53 |
+
### Downstream Use [optional]
|
| 54 |
+
|
| 55 |
+
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
| 56 |
+
|
| 57 |
+
[More Information Needed]
|
| 58 |
+
|
| 59 |
+
### Out-of-Scope Use
|
| 60 |
+
|
| 61 |
+
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
| 62 |
+
|
| 63 |
+
[More Information Needed]
|
| 64 |
+
|
| 65 |
+
## Bias, Risks, and Limitations
|
| 66 |
+
|
| 67 |
+
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
| 68 |
+
|
| 69 |
+
[More Information Needed]
|
| 70 |
+
|
| 71 |
+
### Recommendations
|
| 72 |
+
|
| 73 |
+
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
| 74 |
+
|
| 75 |
+
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
| 76 |
+
|
| 77 |
+
## How to Get Started with the Model
|
| 78 |
+
|
| 79 |
+
Use the code below to get started with the model.
|
| 80 |
+
|
| 81 |
+
[More Information Needed]
|
| 82 |
+
|
| 83 |
+
## Training Details
|
| 84 |
+
|
| 85 |
+
### Training Data
|
| 86 |
+
|
| 87 |
+
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
| 88 |
+
|
| 89 |
+
[More Information Needed]
|
| 90 |
+
|
| 91 |
+
### Training Procedure
|
| 92 |
+
|
| 93 |
+
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
| 94 |
+
|
| 95 |
+
#### Preprocessing [optional]
|
| 96 |
+
|
| 97 |
+
[More Information Needed]
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
#### Training Hyperparameters
|
| 101 |
+
|
| 102 |
+
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
| 103 |
+
|
| 104 |
+
#### Speeds, Sizes, Times [optional]
|
| 105 |
+
|
| 106 |
+
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
| 107 |
+
|
| 108 |
+
[More Information Needed]
|
| 109 |
+
|
| 110 |
+
## Evaluation
|
| 111 |
+
|
| 112 |
+
<!-- This section describes the evaluation protocols and provides the results. -->
|
| 113 |
+
|
| 114 |
+
### Testing Data, Factors & Metrics
|
| 115 |
+
|
| 116 |
+
#### Testing Data
|
| 117 |
+
|
| 118 |
+
<!-- This should link to a Dataset Card if possible. -->
|
| 119 |
+
|
| 120 |
+
[More Information Needed]
|
| 121 |
+
|
| 122 |
+
#### Factors
|
| 123 |
+
|
| 124 |
+
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
| 125 |
+
|
| 126 |
+
[More Information Needed]
|
| 127 |
+
|
| 128 |
+
#### Metrics
|
| 129 |
+
|
| 130 |
+
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
| 131 |
+
|
| 132 |
+
[More Information Needed]
|
| 133 |
+
|
| 134 |
+
### Results
|
| 135 |
+
|
| 136 |
+
[More Information Needed]
|
| 137 |
+
|
| 138 |
+
#### Summary
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
## Model Examination [optional]
|
| 143 |
+
|
| 144 |
+
<!-- Relevant interpretability work for the model goes here -->
|
| 145 |
+
|
| 146 |
+
[More Information Needed]
|
| 147 |
+
|
| 148 |
+
## Environmental Impact
|
| 149 |
+
|
| 150 |
+
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
| 151 |
+
|
| 152 |
+
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
| 153 |
+
|
| 154 |
+
- **Hardware Type:** [More Information Needed]
|
| 155 |
+
- **Hours used:** [More Information Needed]
|
| 156 |
+
- **Cloud Provider:** [More Information Needed]
|
| 157 |
+
- **Compute Region:** [More Information Needed]
|
| 158 |
+
- **Carbon Emitted:** [More Information Needed]
|
| 159 |
+
|
| 160 |
+
## Technical Specifications [optional]
|
| 161 |
+
|
| 162 |
+
### Model Architecture and Objective
|
| 163 |
+
|
| 164 |
+
[More Information Needed]
|
| 165 |
+
|
| 166 |
+
### Compute Infrastructure
|
| 167 |
+
|
| 168 |
+
[More Information Needed]
|
| 169 |
+
|
| 170 |
+
#### Hardware
|
| 171 |
+
|
| 172 |
+
[More Information Needed]
|
| 173 |
+
|
| 174 |
+
#### Software
|
| 175 |
+
|
| 176 |
+
[More Information Needed]
|
| 177 |
+
|
| 178 |
+
## Citation [optional]
|
| 179 |
+
|
| 180 |
+
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
| 181 |
+
|
| 182 |
+
**BibTeX:**
|
| 183 |
+
|
| 184 |
+
[More Information Needed]
|
| 185 |
+
|
| 186 |
+
**APA:**
|
| 187 |
+
|
| 188 |
+
[More Information Needed]
|
| 189 |
+
|
| 190 |
+
## Glossary [optional]
|
| 191 |
+
|
| 192 |
+
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
| 193 |
+
|
| 194 |
+
[More Information Needed]
|
| 195 |
+
|
| 196 |
+
## More Information [optional]
|
| 197 |
+
|
| 198 |
+
[More Information Needed]
|
| 199 |
+
|
| 200 |
+
## Model Card Authors [optional]
|
| 201 |
+
|
| 202 |
+
[More Information Needed]
|
| 203 |
+
|
| 204 |
+
## Model Card Contact
|
| 205 |
+
|
| 206 |
+
[More Information Needed]
|
| 207 |
+
### Framework versions
|
| 208 |
+
|
| 209 |
+
- PEFT 0.19.1
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/adapter_config.json
ADDED
|
@@ -0,0 +1,48 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "Qwen/Qwen3-14B-Base",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"layer_replication": null,
|
| 16 |
+
"layers_pattern": null,
|
| 17 |
+
"layers_to_transform": null,
|
| 18 |
+
"loftq_config": {},
|
| 19 |
+
"lora_alpha": 256,
|
| 20 |
+
"lora_bias": false,
|
| 21 |
+
"lora_dropout": 0.00237968804112545,
|
| 22 |
+
"lora_ga_config": null,
|
| 23 |
+
"megatron_config": null,
|
| 24 |
+
"megatron_core": "megatron.core",
|
| 25 |
+
"modules_to_save": null,
|
| 26 |
+
"peft_type": "LORA",
|
| 27 |
+
"peft_version": "0.19.1",
|
| 28 |
+
"qalora_group_size": 16,
|
| 29 |
+
"r": 128,
|
| 30 |
+
"rank_pattern": {},
|
| 31 |
+
"revision": null,
|
| 32 |
+
"target_modules": [
|
| 33 |
+
"gate_proj",
|
| 34 |
+
"q_proj",
|
| 35 |
+
"up_proj",
|
| 36 |
+
"o_proj",
|
| 37 |
+
"v_proj",
|
| 38 |
+
"k_proj",
|
| 39 |
+
"down_proj"
|
| 40 |
+
],
|
| 41 |
+
"target_parameters": null,
|
| 42 |
+
"task_type": "CAUSAL_LM",
|
| 43 |
+
"trainable_token_indices": null,
|
| 44 |
+
"use_bdlora": null,
|
| 45 |
+
"use_dora": false,
|
| 46 |
+
"use_qalora": false,
|
| 47 |
+
"use_rslora": false
|
| 48 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/chat_template.jinja
ADDED
|
@@ -0,0 +1,85 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- if tools %}
|
| 2 |
+
{{- '<|im_start|>system\n' }}
|
| 3 |
+
{%- if messages[0].role == 'system' %}
|
| 4 |
+
{{- messages[0].content + '\n\n' }}
|
| 5 |
+
{%- endif %}
|
| 6 |
+
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
| 7 |
+
{%- for tool in tools %}
|
| 8 |
+
{{- "\n" }}
|
| 9 |
+
{{- tool | tojson }}
|
| 10 |
+
{%- endfor %}
|
| 11 |
+
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
| 12 |
+
{%- else %}
|
| 13 |
+
{%- if messages[0].role == 'system' %}
|
| 14 |
+
{{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
|
| 15 |
+
{%- endif %}
|
| 16 |
+
{%- endif %}
|
| 17 |
+
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
| 18 |
+
{%- for message in messages[::-1] %}
|
| 19 |
+
{%- set index = (messages|length - 1) - loop.index0 %}
|
| 20 |
+
{%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
|
| 21 |
+
{%- set ns.multi_step_tool = false %}
|
| 22 |
+
{%- set ns.last_query_index = index %}
|
| 23 |
+
{%- endif %}
|
| 24 |
+
{%- endfor %}
|
| 25 |
+
{%- for message in messages %}
|
| 26 |
+
{%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
|
| 27 |
+
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
|
| 28 |
+
{%- elif message.role == "assistant" %}
|
| 29 |
+
{%- set content = message.content %}
|
| 30 |
+
{%- set reasoning_content = '' %}
|
| 31 |
+
{%- if message.reasoning_content is defined and message.reasoning_content is not none %}
|
| 32 |
+
{%- set reasoning_content = message.reasoning_content %}
|
| 33 |
+
{%- else %}
|
| 34 |
+
{%- if '</think>' in message.content %}
|
| 35 |
+
{%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
|
| 36 |
+
{%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
| 37 |
+
{%- endif %}
|
| 38 |
+
{%- endif %}
|
| 39 |
+
{%- if loop.index0 > ns.last_query_index %}
|
| 40 |
+
{%- if loop.last or (not loop.last and reasoning_content) %}
|
| 41 |
+
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
|
| 42 |
+
{%- else %}
|
| 43 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 44 |
+
{%- endif %}
|
| 45 |
+
{%- else %}
|
| 46 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 47 |
+
{%- endif %}
|
| 48 |
+
{%- if message.tool_calls %}
|
| 49 |
+
{%- for tool_call in message.tool_calls %}
|
| 50 |
+
{%- if (loop.first and content) or (not loop.first) %}
|
| 51 |
+
{{- '\n' }}
|
| 52 |
+
{%- endif %}
|
| 53 |
+
{%- if tool_call.function %}
|
| 54 |
+
{%- set tool_call = tool_call.function %}
|
| 55 |
+
{%- endif %}
|
| 56 |
+
{{- '<tool_call>\n{"name": "' }}
|
| 57 |
+
{{- tool_call.name }}
|
| 58 |
+
{{- '", "arguments": ' }}
|
| 59 |
+
{%- if tool_call.arguments is string %}
|
| 60 |
+
{{- tool_call.arguments }}
|
| 61 |
+
{%- else %}
|
| 62 |
+
{{- tool_call.arguments | tojson }}
|
| 63 |
+
{%- endif %}
|
| 64 |
+
{{- '}\n</tool_call>' }}
|
| 65 |
+
{%- endfor %}
|
| 66 |
+
{%- endif %}
|
| 67 |
+
{{- '<|im_end|>\n' }}
|
| 68 |
+
{%- elif message.role == "tool" %}
|
| 69 |
+
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
|
| 70 |
+
{{- '<|im_start|>user' }}
|
| 71 |
+
{%- endif %}
|
| 72 |
+
{{- '\n<tool_response>\n' }}
|
| 73 |
+
{{- message.content }}
|
| 74 |
+
{{- '\n</tool_response>' }}
|
| 75 |
+
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
| 76 |
+
{{- '<|im_end|>\n' }}
|
| 77 |
+
{%- endif %}
|
| 78 |
+
{%- endif %}
|
| 79 |
+
{%- endfor %}
|
| 80 |
+
{%- if add_generation_prompt %}
|
| 81 |
+
{{- '<|im_start|>assistant\n' }}
|
| 82 |
+
{%- if enable_thinking is defined and enable_thinking is false %}
|
| 83 |
+
{{- '<think>\n\n</think>\n\n' }}
|
| 84 |
+
{%- endif %}
|
| 85 |
+
{%- endif %}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/tokenizer_config.json
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"backend": "tokenizers",
|
| 4 |
+
"bos_token": null,
|
| 5 |
+
"clean_up_tokenization_spaces": false,
|
| 6 |
+
"eos_token": "<|endoftext|>",
|
| 7 |
+
"errors": "replace",
|
| 8 |
+
"extra_special_tokens": [
|
| 9 |
+
"<|im_start|>",
|
| 10 |
+
"<|im_end|>",
|
| 11 |
+
"<|object_ref_start|>",
|
| 12 |
+
"<|object_ref_end|>",
|
| 13 |
+
"<|box_start|>",
|
| 14 |
+
"<|box_end|>",
|
| 15 |
+
"<|quad_start|>",
|
| 16 |
+
"<|quad_end|>",
|
| 17 |
+
"<|vision_start|>",
|
| 18 |
+
"<|vision_end|>",
|
| 19 |
+
"<|vision_pad|>",
|
| 20 |
+
"<|image_pad|>",
|
| 21 |
+
"<|video_pad|>"
|
| 22 |
+
],
|
| 23 |
+
"is_local": false,
|
| 24 |
+
"model_max_length": 131072,
|
| 25 |
+
"pad_token": "<|endoftext|>",
|
| 26 |
+
"split_special_tokens": false,
|
| 27 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 28 |
+
"unk_token": null
|
| 29 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-389/trainer_state.json
ADDED
|
@@ -0,0 +1,115 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"best_global_step": null,
|
| 3 |
+
"best_metric": null,
|
| 4 |
+
"best_model_checkpoint": null,
|
| 5 |
+
"epoch": 1.0,
|
| 6 |
+
"eval_steps": 500,
|
| 7 |
+
"global_step": 389,
|
| 8 |
+
"is_hyper_param_search": false,
|
| 9 |
+
"is_local_process_zero": true,
|
| 10 |
+
"is_world_process_zero": true,
|
| 11 |
+
"log_history": [
|
| 12 |
+
{
|
| 13 |
+
"entropy": 1.6110110306739807,
|
| 14 |
+
"epoch": 0.1287001287001287,
|
| 15 |
+
"grad_norm": 0.8611119389533997,
|
| 16 |
+
"learning_rate": 3.4130257962133866e-05,
|
| 17 |
+
"loss": 1.53956298828125,
|
| 18 |
+
"mean_token_accuracy": 0.6652342769503593,
|
| 19 |
+
"num_tokens": 73407.0,
|
| 20 |
+
"step": 50
|
| 21 |
+
},
|
| 22 |
+
{
|
| 23 |
+
"entropy": 0.8662820833921433,
|
| 24 |
+
"epoch": 0.2574002574002574,
|
| 25 |
+
"grad_norm": 0.7405035495758057,
|
| 26 |
+
"learning_rate": 6.895705180104598e-05,
|
| 27 |
+
"loss": 0.7929539489746094,
|
| 28 |
+
"mean_token_accuracy": 0.7786347842216492,
|
| 29 |
+
"num_tokens": 143994.0,
|
| 30 |
+
"step": 100
|
| 31 |
+
},
|
| 32 |
+
{
|
| 33 |
+
"entropy": 0.7912819278240204,
|
| 34 |
+
"epoch": 0.3861003861003861,
|
| 35 |
+
"grad_norm": 0.6105485558509827,
|
| 36 |
+
"learning_rate": 0.00010378384563995809,
|
| 37 |
+
"loss": 0.7236511993408203,
|
| 38 |
+
"mean_token_accuracy": 0.7925467795133591,
|
| 39 |
+
"num_tokens": 216171.0,
|
| 40 |
+
"step": 150
|
| 41 |
+
},
|
| 42 |
+
{
|
| 43 |
+
"entropy": 0.7777709531784057,
|
| 44 |
+
"epoch": 0.5148005148005148,
|
| 45 |
+
"grad_norm": 0.5781793594360352,
|
| 46 |
+
"learning_rate": 0.0001386106394788702,
|
| 47 |
+
"loss": 0.7024919891357422,
|
| 48 |
+
"mean_token_accuracy": 0.7969876372814179,
|
| 49 |
+
"num_tokens": 284702.0,
|
| 50 |
+
"step": 200
|
| 51 |
+
},
|
| 52 |
+
{
|
| 53 |
+
"entropy": 0.7632609683275223,
|
| 54 |
+
"epoch": 0.6435006435006435,
|
| 55 |
+
"grad_norm": 0.6553444862365723,
|
| 56 |
+
"learning_rate": 0.00017343743331778232,
|
| 57 |
+
"loss": 0.7000718688964844,
|
| 58 |
+
"mean_token_accuracy": 0.7994742071628571,
|
| 59 |
+
"num_tokens": 356393.0,
|
| 60 |
+
"step": 250
|
| 61 |
+
},
|
| 62 |
+
{
|
| 63 |
+
"entropy": 0.754165632724762,
|
| 64 |
+
"epoch": 0.7722007722007722,
|
| 65 |
+
"grad_norm": 0.634222149848938,
|
| 66 |
+
"learning_rate": 0.00020826422715669444,
|
| 67 |
+
"loss": 0.6909049987792969,
|
| 68 |
+
"mean_token_accuracy": 0.8016809666156769,
|
| 69 |
+
"num_tokens": 426916.0,
|
| 70 |
+
"step": 300
|
| 71 |
+
},
|
| 72 |
+
{
|
| 73 |
+
"entropy": 0.7530711203813553,
|
| 74 |
+
"epoch": 0.9009009009009009,
|
| 75 |
+
"grad_norm": 0.7020156383514404,
|
| 76 |
+
"learning_rate": 0.00024309102099560653,
|
| 77 |
+
"loss": 0.69554931640625,
|
| 78 |
+
"mean_token_accuracy": 0.8002807641029358,
|
| 79 |
+
"num_tokens": 499603.0,
|
| 80 |
+
"step": 350
|
| 81 |
+
},
|
| 82 |
+
{
|
| 83 |
+
"epoch": 1.0,
|
| 84 |
+
"eval_entropy": 0.5728044164242204,
|
| 85 |
+
"eval_loss": 0.6785851716995239,
|
| 86 |
+
"eval_mean_token_accuracy": 0.798433246993527,
|
| 87 |
+
"eval_num_tokens": 555493.0,
|
| 88 |
+
"eval_runtime": 162.4129,
|
| 89 |
+
"eval_samples_per_second": 9.513,
|
| 90 |
+
"eval_steps_per_second": 1.194,
|
| 91 |
+
"step": 389
|
| 92 |
+
}
|
| 93 |
+
],
|
| 94 |
+
"logging_steps": 50,
|
| 95 |
+
"max_steps": 3890,
|
| 96 |
+
"num_input_tokens_seen": 0,
|
| 97 |
+
"num_train_epochs": 10,
|
| 98 |
+
"save_steps": 500,
|
| 99 |
+
"stateful_callbacks": {
|
| 100 |
+
"TrainerControl": {
|
| 101 |
+
"args": {
|
| 102 |
+
"should_epoch_stop": false,
|
| 103 |
+
"should_evaluate": false,
|
| 104 |
+
"should_log": false,
|
| 105 |
+
"should_save": true,
|
| 106 |
+
"should_training_stop": false
|
| 107 |
+
},
|
| 108 |
+
"attributes": {}
|
| 109 |
+
}
|
| 110 |
+
},
|
| 111 |
+
"total_flos": 9.306612280369152e+16,
|
| 112 |
+
"train_batch_size": 8,
|
| 113 |
+
"trial_name": null,
|
| 114 |
+
"trial_params": null
|
| 115 |
+
}
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3890/README.md
ADDED
|
@@ -0,0 +1,209 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen3-14B-Base
|
| 3 |
+
library_name: peft
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:Qwen/Qwen3-14B-Base
|
| 7 |
+
- lora
|
| 8 |
+
- sft
|
| 9 |
+
- transformers
|
| 10 |
+
- trl
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Model Card for Model ID
|
| 14 |
+
|
| 15 |
+
<!-- Provide a quick summary of what the model is/does. -->
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
## Model Details
|
| 20 |
+
|
| 21 |
+
### Model Description
|
| 22 |
+
|
| 23 |
+
<!-- Provide a longer summary of what this model is. -->
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
- **Developed by:** [More Information Needed]
|
| 28 |
+
- **Funded by [optional]:** [More Information Needed]
|
| 29 |
+
- **Shared by [optional]:** [More Information Needed]
|
| 30 |
+
- **Model type:** [More Information Needed]
|
| 31 |
+
- **Language(s) (NLP):** [More Information Needed]
|
| 32 |
+
- **License:** [More Information Needed]
|
| 33 |
+
- **Finetuned from model [optional]:** [More Information Needed]
|
| 34 |
+
|
| 35 |
+
### Model Sources [optional]
|
| 36 |
+
|
| 37 |
+
<!-- Provide the basic links for the model. -->
|
| 38 |
+
|
| 39 |
+
- **Repository:** [More Information Needed]
|
| 40 |
+
- **Paper [optional]:** [More Information Needed]
|
| 41 |
+
- **Demo [optional]:** [More Information Needed]
|
| 42 |
+
|
| 43 |
+
## Uses
|
| 44 |
+
|
| 45 |
+
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
| 46 |
+
|
| 47 |
+
### Direct Use
|
| 48 |
+
|
| 49 |
+
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
| 50 |
+
|
| 51 |
+
[More Information Needed]
|
| 52 |
+
|
| 53 |
+
### Downstream Use [optional]
|
| 54 |
+
|
| 55 |
+
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
| 56 |
+
|
| 57 |
+
[More Information Needed]
|
| 58 |
+
|
| 59 |
+
### Out-of-Scope Use
|
| 60 |
+
|
| 61 |
+
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
| 62 |
+
|
| 63 |
+
[More Information Needed]
|
| 64 |
+
|
| 65 |
+
## Bias, Risks, and Limitations
|
| 66 |
+
|
| 67 |
+
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
| 68 |
+
|
| 69 |
+
[More Information Needed]
|
| 70 |
+
|
| 71 |
+
### Recommendations
|
| 72 |
+
|
| 73 |
+
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
| 74 |
+
|
| 75 |
+
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
| 76 |
+
|
| 77 |
+
## How to Get Started with the Model
|
| 78 |
+
|
| 79 |
+
Use the code below to get started with the model.
|
| 80 |
+
|
| 81 |
+
[More Information Needed]
|
| 82 |
+
|
| 83 |
+
## Training Details
|
| 84 |
+
|
| 85 |
+
### Training Data
|
| 86 |
+
|
| 87 |
+
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
| 88 |
+
|
| 89 |
+
[More Information Needed]
|
| 90 |
+
|
| 91 |
+
### Training Procedure
|
| 92 |
+
|
| 93 |
+
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
| 94 |
+
|
| 95 |
+
#### Preprocessing [optional]
|
| 96 |
+
|
| 97 |
+
[More Information Needed]
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
#### Training Hyperparameters
|
| 101 |
+
|
| 102 |
+
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
| 103 |
+
|
| 104 |
+
#### Speeds, Sizes, Times [optional]
|
| 105 |
+
|
| 106 |
+
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
| 107 |
+
|
| 108 |
+
[More Information Needed]
|
| 109 |
+
|
| 110 |
+
## Evaluation
|
| 111 |
+
|
| 112 |
+
<!-- This section describes the evaluation protocols and provides the results. -->
|
| 113 |
+
|
| 114 |
+
### Testing Data, Factors & Metrics
|
| 115 |
+
|
| 116 |
+
#### Testing Data
|
| 117 |
+
|
| 118 |
+
<!-- This should link to a Dataset Card if possible. -->
|
| 119 |
+
|
| 120 |
+
[More Information Needed]
|
| 121 |
+
|
| 122 |
+
#### Factors
|
| 123 |
+
|
| 124 |
+
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
| 125 |
+
|
| 126 |
+
[More Information Needed]
|
| 127 |
+
|
| 128 |
+
#### Metrics
|
| 129 |
+
|
| 130 |
+
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
| 131 |
+
|
| 132 |
+
[More Information Needed]
|
| 133 |
+
|
| 134 |
+
### Results
|
| 135 |
+
|
| 136 |
+
[More Information Needed]
|
| 137 |
+
|
| 138 |
+
#### Summary
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
## Model Examination [optional]
|
| 143 |
+
|
| 144 |
+
<!-- Relevant interpretability work for the model goes here -->
|
| 145 |
+
|
| 146 |
+
[More Information Needed]
|
| 147 |
+
|
| 148 |
+
## Environmental Impact
|
| 149 |
+
|
| 150 |
+
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
| 151 |
+
|
| 152 |
+
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
| 153 |
+
|
| 154 |
+
- **Hardware Type:** [More Information Needed]
|
| 155 |
+
- **Hours used:** [More Information Needed]
|
| 156 |
+
- **Cloud Provider:** [More Information Needed]
|
| 157 |
+
- **Compute Region:** [More Information Needed]
|
| 158 |
+
- **Carbon Emitted:** [More Information Needed]
|
| 159 |
+
|
| 160 |
+
## Technical Specifications [optional]
|
| 161 |
+
|
| 162 |
+
### Model Architecture and Objective
|
| 163 |
+
|
| 164 |
+
[More Information Needed]
|
| 165 |
+
|
| 166 |
+
### Compute Infrastructure
|
| 167 |
+
|
| 168 |
+
[More Information Needed]
|
| 169 |
+
|
| 170 |
+
#### Hardware
|
| 171 |
+
|
| 172 |
+
[More Information Needed]
|
| 173 |
+
|
| 174 |
+
#### Software
|
| 175 |
+
|
| 176 |
+
[More Information Needed]
|
| 177 |
+
|
| 178 |
+
## Citation [optional]
|
| 179 |
+
|
| 180 |
+
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
| 181 |
+
|
| 182 |
+
**BibTeX:**
|
| 183 |
+
|
| 184 |
+
[More Information Needed]
|
| 185 |
+
|
| 186 |
+
**APA:**
|
| 187 |
+
|
| 188 |
+
[More Information Needed]
|
| 189 |
+
|
| 190 |
+
## Glossary [optional]
|
| 191 |
+
|
| 192 |
+
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
| 193 |
+
|
| 194 |
+
[More Information Needed]
|
| 195 |
+
|
| 196 |
+
## More Information [optional]
|
| 197 |
+
|
| 198 |
+
[More Information Needed]
|
| 199 |
+
|
| 200 |
+
## Model Card Authors [optional]
|
| 201 |
+
|
| 202 |
+
[More Information Needed]
|
| 203 |
+
|
| 204 |
+
## Model Card Contact
|
| 205 |
+
|
| 206 |
+
[More Information Needed]
|
| 207 |
+
### Framework versions
|
| 208 |
+
|
| 209 |
+
- PEFT 0.19.1
|
productivity_original_Estonian/Qwen3-14B-Base_productivity_splits_original_features_train_productivity_splits_original_features_test2/checkpoint-3890/adapter_config.json
ADDED
|
@@ -0,0 +1,48 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "Qwen/Qwen3-14B-Base",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"layer_replication": null,
|
| 16 |
+
"layers_pattern": null,
|
| 17 |
+
"layers_to_transform": null,
|
| 18 |
+
"loftq_config": {},
|
| 19 |
+
"lora_alpha": 256,
|
| 20 |
+
"lora_bias": false,
|
| 21 |
+
"lora_dropout": 0.00237968804112545,
|
| 22 |
+
"lora_ga_config": null,
|
| 23 |
+
"megatron_config": null,
|
| 24 |
+
"megatron_core": "megatron.core",
|
| 25 |
+
"modules_to_save": null,
|
| 26 |
+
"peft_type": "LORA",
|
| 27 |
+
"peft_version": "0.19.1",
|
| 28 |
+
"qalora_group_size": 16,
|
| 29 |
+
"r": 128,
|
| 30 |
+
"rank_pattern": {},
|
| 31 |
+
"revision": null,
|
| 32 |
+
"target_modules": [
|
| 33 |
+
"gate_proj",
|
| 34 |
+
"q_proj",
|
| 35 |
+
"up_proj",
|
| 36 |
+
"o_proj",
|
| 37 |
+
"v_proj",
|
| 38 |
+
"k_proj",
|
| 39 |
+
"down_proj"
|
| 40 |
+
],
|
| 41 |
+
"target_parameters": null,
|
| 42 |
+
"task_type": "CAUSAL_LM",
|
| 43 |
+
"trainable_token_indices": null,
|
| 44 |
+
"use_bdlora": null,
|
| 45 |
+
"use_dora": false,
|
| 46 |
+
"use_qalora": false,
|
| 47 |
+
"use_rslora": false
|
| 48 |
+
}
|
random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/README.md
ADDED
|
@@ -0,0 +1,209 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen3-14B-Base
|
| 3 |
+
library_name: peft
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:Qwen/Qwen3-14B-Base
|
| 7 |
+
- lora
|
| 8 |
+
- sft
|
| 9 |
+
- transformers
|
| 10 |
+
- trl
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Model Card for Model ID
|
| 14 |
+
|
| 15 |
+
<!-- Provide a quick summary of what the model is/does. -->
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
## Model Details
|
| 20 |
+
|
| 21 |
+
### Model Description
|
| 22 |
+
|
| 23 |
+
<!-- Provide a longer summary of what this model is. -->
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
- **Developed by:** [More Information Needed]
|
| 28 |
+
- **Funded by [optional]:** [More Information Needed]
|
| 29 |
+
- **Shared by [optional]:** [More Information Needed]
|
| 30 |
+
- **Model type:** [More Information Needed]
|
| 31 |
+
- **Language(s) (NLP):** [More Information Needed]
|
| 32 |
+
- **License:** [More Information Needed]
|
| 33 |
+
- **Finetuned from model [optional]:** [More Information Needed]
|
| 34 |
+
|
| 35 |
+
### Model Sources [optional]
|
| 36 |
+
|
| 37 |
+
<!-- Provide the basic links for the model. -->
|
| 38 |
+
|
| 39 |
+
- **Repository:** [More Information Needed]
|
| 40 |
+
- **Paper [optional]:** [More Information Needed]
|
| 41 |
+
- **Demo [optional]:** [More Information Needed]
|
| 42 |
+
|
| 43 |
+
## Uses
|
| 44 |
+
|
| 45 |
+
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
| 46 |
+
|
| 47 |
+
### Direct Use
|
| 48 |
+
|
| 49 |
+
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
| 50 |
+
|
| 51 |
+
[More Information Needed]
|
| 52 |
+
|
| 53 |
+
### Downstream Use [optional]
|
| 54 |
+
|
| 55 |
+
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
| 56 |
+
|
| 57 |
+
[More Information Needed]
|
| 58 |
+
|
| 59 |
+
### Out-of-Scope Use
|
| 60 |
+
|
| 61 |
+
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
| 62 |
+
|
| 63 |
+
[More Information Needed]
|
| 64 |
+
|
| 65 |
+
## Bias, Risks, and Limitations
|
| 66 |
+
|
| 67 |
+
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
| 68 |
+
|
| 69 |
+
[More Information Needed]
|
| 70 |
+
|
| 71 |
+
### Recommendations
|
| 72 |
+
|
| 73 |
+
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
| 74 |
+
|
| 75 |
+
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
| 76 |
+
|
| 77 |
+
## How to Get Started with the Model
|
| 78 |
+
|
| 79 |
+
Use the code below to get started with the model.
|
| 80 |
+
|
| 81 |
+
[More Information Needed]
|
| 82 |
+
|
| 83 |
+
## Training Details
|
| 84 |
+
|
| 85 |
+
### Training Data
|
| 86 |
+
|
| 87 |
+
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
| 88 |
+
|
| 89 |
+
[More Information Needed]
|
| 90 |
+
|
| 91 |
+
### Training Procedure
|
| 92 |
+
|
| 93 |
+
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
| 94 |
+
|
| 95 |
+
#### Preprocessing [optional]
|
| 96 |
+
|
| 97 |
+
[More Information Needed]
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
#### Training Hyperparameters
|
| 101 |
+
|
| 102 |
+
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
| 103 |
+
|
| 104 |
+
#### Speeds, Sizes, Times [optional]
|
| 105 |
+
|
| 106 |
+
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
| 107 |
+
|
| 108 |
+
[More Information Needed]
|
| 109 |
+
|
| 110 |
+
## Evaluation
|
| 111 |
+
|
| 112 |
+
<!-- This section describes the evaluation protocols and provides the results. -->
|
| 113 |
+
|
| 114 |
+
### Testing Data, Factors & Metrics
|
| 115 |
+
|
| 116 |
+
#### Testing Data
|
| 117 |
+
|
| 118 |
+
<!-- This should link to a Dataset Card if possible. -->
|
| 119 |
+
|
| 120 |
+
[More Information Needed]
|
| 121 |
+
|
| 122 |
+
#### Factors
|
| 123 |
+
|
| 124 |
+
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
| 125 |
+
|
| 126 |
+
[More Information Needed]
|
| 127 |
+
|
| 128 |
+
#### Metrics
|
| 129 |
+
|
| 130 |
+
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
| 131 |
+
|
| 132 |
+
[More Information Needed]
|
| 133 |
+
|
| 134 |
+
### Results
|
| 135 |
+
|
| 136 |
+
[More Information Needed]
|
| 137 |
+
|
| 138 |
+
#### Summary
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
## Model Examination [optional]
|
| 143 |
+
|
| 144 |
+
<!-- Relevant interpretability work for the model goes here -->
|
| 145 |
+
|
| 146 |
+
[More Information Needed]
|
| 147 |
+
|
| 148 |
+
## Environmental Impact
|
| 149 |
+
|
| 150 |
+
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
| 151 |
+
|
| 152 |
+
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
| 153 |
+
|
| 154 |
+
- **Hardware Type:** [More Information Needed]
|
| 155 |
+
- **Hours used:** [More Information Needed]
|
| 156 |
+
- **Cloud Provider:** [More Information Needed]
|
| 157 |
+
- **Compute Region:** [More Information Needed]
|
| 158 |
+
- **Carbon Emitted:** [More Information Needed]
|
| 159 |
+
|
| 160 |
+
## Technical Specifications [optional]
|
| 161 |
+
|
| 162 |
+
### Model Architecture and Objective
|
| 163 |
+
|
| 164 |
+
[More Information Needed]
|
| 165 |
+
|
| 166 |
+
### Compute Infrastructure
|
| 167 |
+
|
| 168 |
+
[More Information Needed]
|
| 169 |
+
|
| 170 |
+
#### Hardware
|
| 171 |
+
|
| 172 |
+
[More Information Needed]
|
| 173 |
+
|
| 174 |
+
#### Software
|
| 175 |
+
|
| 176 |
+
[More Information Needed]
|
| 177 |
+
|
| 178 |
+
## Citation [optional]
|
| 179 |
+
|
| 180 |
+
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
| 181 |
+
|
| 182 |
+
**BibTeX:**
|
| 183 |
+
|
| 184 |
+
[More Information Needed]
|
| 185 |
+
|
| 186 |
+
**APA:**
|
| 187 |
+
|
| 188 |
+
[More Information Needed]
|
| 189 |
+
|
| 190 |
+
## Glossary [optional]
|
| 191 |
+
|
| 192 |
+
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
| 193 |
+
|
| 194 |
+
[More Information Needed]
|
| 195 |
+
|
| 196 |
+
## More Information [optional]
|
| 197 |
+
|
| 198 |
+
[More Information Needed]
|
| 199 |
+
|
| 200 |
+
## Model Card Authors [optional]
|
| 201 |
+
|
| 202 |
+
[More Information Needed]
|
| 203 |
+
|
| 204 |
+
## Model Card Contact
|
| 205 |
+
|
| 206 |
+
[More Information Needed]
|
| 207 |
+
### Framework versions
|
| 208 |
+
|
| 209 |
+
- PEFT 0.19.1
|
random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/adapter_config.json
ADDED
|
@@ -0,0 +1,48 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "Qwen/Qwen3-14B-Base",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"layer_replication": null,
|
| 16 |
+
"layers_pattern": null,
|
| 17 |
+
"layers_to_transform": null,
|
| 18 |
+
"loftq_config": {},
|
| 19 |
+
"lora_alpha": 64,
|
| 20 |
+
"lora_bias": false,
|
| 21 |
+
"lora_dropout": 0.09831666542701797,
|
| 22 |
+
"lora_ga_config": null,
|
| 23 |
+
"megatron_config": null,
|
| 24 |
+
"megatron_core": "megatron.core",
|
| 25 |
+
"modules_to_save": null,
|
| 26 |
+
"peft_type": "LORA",
|
| 27 |
+
"peft_version": "0.19.1",
|
| 28 |
+
"qalora_group_size": 16,
|
| 29 |
+
"r": 32,
|
| 30 |
+
"rank_pattern": {},
|
| 31 |
+
"revision": null,
|
| 32 |
+
"target_modules": [
|
| 33 |
+
"gate_proj",
|
| 34 |
+
"o_proj",
|
| 35 |
+
"k_proj",
|
| 36 |
+
"q_proj",
|
| 37 |
+
"up_proj",
|
| 38 |
+
"v_proj",
|
| 39 |
+
"down_proj"
|
| 40 |
+
],
|
| 41 |
+
"target_parameters": null,
|
| 42 |
+
"task_type": "CAUSAL_LM",
|
| 43 |
+
"trainable_token_indices": null,
|
| 44 |
+
"use_bdlora": null,
|
| 45 |
+
"use_dora": false,
|
| 46 |
+
"use_qalora": false,
|
| 47 |
+
"use_rslora": false
|
| 48 |
+
}
|
random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/chat_template.jinja
ADDED
|
@@ -0,0 +1,85 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- if tools %}
|
| 2 |
+
{{- '<|im_start|>system\n' }}
|
| 3 |
+
{%- if messages[0].role == 'system' %}
|
| 4 |
+
{{- messages[0].content + '\n\n' }}
|
| 5 |
+
{%- endif %}
|
| 6 |
+
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
| 7 |
+
{%- for tool in tools %}
|
| 8 |
+
{{- "\n" }}
|
| 9 |
+
{{- tool | tojson }}
|
| 10 |
+
{%- endfor %}
|
| 11 |
+
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
| 12 |
+
{%- else %}
|
| 13 |
+
{%- if messages[0].role == 'system' %}
|
| 14 |
+
{{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
|
| 15 |
+
{%- endif %}
|
| 16 |
+
{%- endif %}
|
| 17 |
+
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
| 18 |
+
{%- for message in messages[::-1] %}
|
| 19 |
+
{%- set index = (messages|length - 1) - loop.index0 %}
|
| 20 |
+
{%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
|
| 21 |
+
{%- set ns.multi_step_tool = false %}
|
| 22 |
+
{%- set ns.last_query_index = index %}
|
| 23 |
+
{%- endif %}
|
| 24 |
+
{%- endfor %}
|
| 25 |
+
{%- for message in messages %}
|
| 26 |
+
{%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
|
| 27 |
+
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
|
| 28 |
+
{%- elif message.role == "assistant" %}
|
| 29 |
+
{%- set content = message.content %}
|
| 30 |
+
{%- set reasoning_content = '' %}
|
| 31 |
+
{%- if message.reasoning_content is defined and message.reasoning_content is not none %}
|
| 32 |
+
{%- set reasoning_content = message.reasoning_content %}
|
| 33 |
+
{%- else %}
|
| 34 |
+
{%- if '</think>' in message.content %}
|
| 35 |
+
{%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
|
| 36 |
+
{%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
| 37 |
+
{%- endif %}
|
| 38 |
+
{%- endif %}
|
| 39 |
+
{%- if loop.index0 > ns.last_query_index %}
|
| 40 |
+
{%- if loop.last or (not loop.last and reasoning_content) %}
|
| 41 |
+
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
|
| 42 |
+
{%- else %}
|
| 43 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 44 |
+
{%- endif %}
|
| 45 |
+
{%- else %}
|
| 46 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 47 |
+
{%- endif %}
|
| 48 |
+
{%- if message.tool_calls %}
|
| 49 |
+
{%- for tool_call in message.tool_calls %}
|
| 50 |
+
{%- if (loop.first and content) or (not loop.first) %}
|
| 51 |
+
{{- '\n' }}
|
| 52 |
+
{%- endif %}
|
| 53 |
+
{%- if tool_call.function %}
|
| 54 |
+
{%- set tool_call = tool_call.function %}
|
| 55 |
+
{%- endif %}
|
| 56 |
+
{{- '<tool_call>\n{"name": "' }}
|
| 57 |
+
{{- tool_call.name }}
|
| 58 |
+
{{- '", "arguments": ' }}
|
| 59 |
+
{%- if tool_call.arguments is string %}
|
| 60 |
+
{{- tool_call.arguments }}
|
| 61 |
+
{%- else %}
|
| 62 |
+
{{- tool_call.arguments | tojson }}
|
| 63 |
+
{%- endif %}
|
| 64 |
+
{{- '}\n</tool_call>' }}
|
| 65 |
+
{%- endfor %}
|
| 66 |
+
{%- endif %}
|
| 67 |
+
{{- '<|im_end|>\n' }}
|
| 68 |
+
{%- elif message.role == "tool" %}
|
| 69 |
+
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
|
| 70 |
+
{{- '<|im_start|>user' }}
|
| 71 |
+
{%- endif %}
|
| 72 |
+
{{- '\n<tool_response>\n' }}
|
| 73 |
+
{{- message.content }}
|
| 74 |
+
{{- '\n</tool_response>' }}
|
| 75 |
+
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
| 76 |
+
{{- '<|im_end|>\n' }}
|
| 77 |
+
{%- endif %}
|
| 78 |
+
{%- endif %}
|
| 79 |
+
{%- endfor %}
|
| 80 |
+
{%- if add_generation_prompt %}
|
| 81 |
+
{{- '<|im_start|>assistant\n' }}
|
| 82 |
+
{%- if enable_thinking is defined and enable_thinking is false %}
|
| 83 |
+
{{- '<think>\n\n</think>\n\n' }}
|
| 84 |
+
{%- endif %}
|
| 85 |
+
{%- endif %}
|
random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/tokenizer_config.json
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"backend": "tokenizers",
|
| 4 |
+
"bos_token": null,
|
| 5 |
+
"clean_up_tokenization_spaces": false,
|
| 6 |
+
"eos_token": "<|endoftext|>",
|
| 7 |
+
"errors": "replace",
|
| 8 |
+
"extra_special_tokens": [
|
| 9 |
+
"<|im_start|>",
|
| 10 |
+
"<|im_end|>",
|
| 11 |
+
"<|object_ref_start|>",
|
| 12 |
+
"<|object_ref_end|>",
|
| 13 |
+
"<|box_start|>",
|
| 14 |
+
"<|box_end|>",
|
| 15 |
+
"<|quad_start|>",
|
| 16 |
+
"<|quad_end|>",
|
| 17 |
+
"<|vision_start|>",
|
| 18 |
+
"<|vision_end|>",
|
| 19 |
+
"<|vision_pad|>",
|
| 20 |
+
"<|image_pad|>",
|
| 21 |
+
"<|video_pad|>"
|
| 22 |
+
],
|
| 23 |
+
"is_local": false,
|
| 24 |
+
"model_max_length": 131072,
|
| 25 |
+
"pad_token": "<|endoftext|>",
|
| 26 |
+
"split_special_tokens": false,
|
| 27 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 28 |
+
"unk_token": null
|
| 29 |
+
}
|
random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1224/trainer_state.json
ADDED
|
@@ -0,0 +1,307 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"best_global_step": null,
|
| 3 |
+
"best_metric": null,
|
| 4 |
+
"best_model_checkpoint": null,
|
| 5 |
+
"epoch": 3.0,
|
| 6 |
+
"eval_steps": 500,
|
| 7 |
+
"global_step": 1224,
|
| 8 |
+
"is_hyper_param_search": false,
|
| 9 |
+
"is_local_process_zero": true,
|
| 10 |
+
"is_world_process_zero": true,
|
| 11 |
+
"log_history": [
|
| 12 |
+
{
|
| 13 |
+
"entropy": 1.6558652856945992,
|
| 14 |
+
"epoch": 0.12277470841006753,
|
| 15 |
+
"grad_norm": 0.8294193148612976,
|
| 16 |
+
"learning_rate": 2.604481765338724e-05,
|
| 17 |
+
"loss": 1.5730807495117187,
|
| 18 |
+
"mean_token_accuracy": 0.6558220008015633,
|
| 19 |
+
"num_tokens": 137019.0,
|
| 20 |
+
"step": 50
|
| 21 |
+
},
|
| 22 |
+
{
|
| 23 |
+
"entropy": 0.8554385647177696,
|
| 24 |
+
"epoch": 0.24554941682013506,
|
| 25 |
+
"grad_norm": 1.029894232749939,
|
| 26 |
+
"learning_rate": 5.2621162197659935e-05,
|
| 27 |
+
"loss": 0.7961322021484375,
|
| 28 |
+
"mean_token_accuracy": 0.7755533090233803,
|
| 29 |
+
"num_tokens": 267448.0,
|
| 30 |
+
"step": 100
|
| 31 |
+
},
|
| 32 |
+
{
|
| 33 |
+
"entropy": 0.7380270153284073,
|
| 34 |
+
"epoch": 0.3683241252302026,
|
| 35 |
+
"grad_norm": 0.6096898913383484,
|
| 36 |
+
"learning_rate": 7.919750674193263e-05,
|
| 37 |
+
"loss": 0.6843100738525391,
|
| 38 |
+
"mean_token_accuracy": 0.7998733684420586,
|
| 39 |
+
"num_tokens": 409071.0,
|
| 40 |
+
"step": 150
|
| 41 |
+
},
|
| 42 |
+
{
|
| 43 |
+
"entropy": 0.6932123881578446,
|
| 44 |
+
"epoch": 0.4910988336402701,
|
| 45 |
+
"grad_norm": 0.5934288501739502,
|
| 46 |
+
"learning_rate": 0.00010577385128620532,
|
| 47 |
+
"loss": 0.6458084106445312,
|
| 48 |
+
"mean_token_accuracy": 0.8083393195271492,
|
| 49 |
+
"num_tokens": 542207.0,
|
| 50 |
+
"step": 200
|
| 51 |
+
},
|
| 52 |
+
{
|
| 53 |
+
"entropy": 0.6840551143884659,
|
| 54 |
+
"epoch": 0.6138735420503376,
|
| 55 |
+
"grad_norm": 0.4665544331073761,
|
| 56 |
+
"learning_rate": 0.00013235019583047802,
|
| 57 |
+
"loss": 0.6333241653442383,
|
| 58 |
+
"mean_token_accuracy": 0.8122684139013291,
|
| 59 |
+
"num_tokens": 679619.0,
|
| 60 |
+
"step": 250
|
| 61 |
+
},
|
| 62 |
+
{
|
| 63 |
+
"entropy": 0.6649293206632138,
|
| 64 |
+
"epoch": 0.7366482504604052,
|
| 65 |
+
"grad_norm": 0.47893014550209045,
|
| 66 |
+
"learning_rate": 0.00015892654037475069,
|
| 67 |
+
"loss": 0.6107040786743164,
|
| 68 |
+
"mean_token_accuracy": 0.8165940269827843,
|
| 69 |
+
"num_tokens": 815397.0,
|
| 70 |
+
"step": 300
|
| 71 |
+
},
|
| 72 |
+
{
|
| 73 |
+
"entropy": 0.6492742404341698,
|
| 74 |
+
"epoch": 0.8594229588704727,
|
| 75 |
+
"grad_norm": 0.6059293746948242,
|
| 76 |
+
"learning_rate": 0.0001855028849190234,
|
| 77 |
+
"loss": 0.597186050415039,
|
| 78 |
+
"mean_token_accuracy": 0.8217281407117843,
|
| 79 |
+
"num_tokens": 949315.0,
|
| 80 |
+
"step": 350
|
| 81 |
+
},
|
| 82 |
+
{
|
| 83 |
+
"entropy": 0.642182088047266,
|
| 84 |
+
"epoch": 0.9821976672805403,
|
| 85 |
+
"grad_norm": 0.3793066143989563,
|
| 86 |
+
"learning_rate": 0.0002120792294632961,
|
| 87 |
+
"loss": 0.5949800491333008,
|
| 88 |
+
"mean_token_accuracy": 0.8229166463017463,
|
| 89 |
+
"num_tokens": 1082903.0,
|
| 90 |
+
"step": 400
|
| 91 |
+
},
|
| 92 |
+
{
|
| 93 |
+
"epoch": 1.0,
|
| 94 |
+
"eval_entropy": 0.6541631272860936,
|
| 95 |
+
"eval_loss": 0.5903413891792297,
|
| 96 |
+
"eval_mean_token_accuracy": 0.8252546640804835,
|
| 97 |
+
"eval_num_tokens": 1101768.0,
|
| 98 |
+
"eval_runtime": 108.9056,
|
| 99 |
+
"eval_samples_per_second": 12.818,
|
| 100 |
+
"eval_steps_per_second": 1.607,
|
| 101 |
+
"step": 408
|
| 102 |
+
},
|
| 103 |
+
{
|
| 104 |
+
"entropy": 0.6077755300829253,
|
| 105 |
+
"epoch": 1.1031307550644567,
|
| 106 |
+
"grad_norm": 0.5028847455978394,
|
| 107 |
+
"learning_rate": 0.00021679626884558217,
|
| 108 |
+
"loss": 0.5617346954345703,
|
| 109 |
+
"mean_token_accuracy": 0.8301527442665875,
|
| 110 |
+
"num_tokens": 1223073.0,
|
| 111 |
+
"step": 450
|
| 112 |
+
},
|
| 113 |
+
{
|
| 114 |
+
"entropy": 0.5860125370323658,
|
| 115 |
+
"epoch": 1.2259054634745243,
|
| 116 |
+
"grad_norm": 0.38494572043418884,
|
| 117 |
+
"learning_rate": 0.00021653451093163906,
|
| 118 |
+
"loss": 0.5437137985229492,
|
| 119 |
+
"mean_token_accuracy": 0.8346565261483192,
|
| 120 |
+
"num_tokens": 1360978.0,
|
| 121 |
+
"step": 500
|
| 122 |
+
},
|
| 123 |
+
{
|
| 124 |
+
"entropy": 0.5945393888652325,
|
| 125 |
+
"epoch": 1.3486801718845918,
|
| 126 |
+
"grad_norm": 0.5123576521873474,
|
| 127 |
+
"learning_rate": 0.00021607496224450087,
|
| 128 |
+
"loss": 0.54191650390625,
|
| 129 |
+
"mean_token_accuracy": 0.8335164493322372,
|
| 130 |
+
"num_tokens": 1492506.0,
|
| 131 |
+
"step": 550
|
| 132 |
+
},
|
| 133 |
+
{
|
| 134 |
+
"entropy": 0.6083934807777405,
|
| 135 |
+
"epoch": 1.4714548802946594,
|
| 136 |
+
"grad_norm": 0.46661439538002014,
|
| 137 |
+
"learning_rate": 0.000215418463597734,
|
| 138 |
+
"loss": 0.5550478744506836,
|
| 139 |
+
"mean_token_accuracy": 0.8311341696977615,
|
| 140 |
+
"num_tokens": 1622275.0,
|
| 141 |
+
"step": 600
|
| 142 |
+
},
|
| 143 |
+
{
|
| 144 |
+
"entropy": 0.5862340961396694,
|
| 145 |
+
"epoch": 1.5942295887047269,
|
| 146 |
+
"grad_norm": 0.49572765827178955,
|
| 147 |
+
"learning_rate": 0.00021456621615453177,
|
| 148 |
+
"loss": 0.5297146606445312,
|
| 149 |
+
"mean_token_accuracy": 0.8360210624337197,
|
| 150 |
+
"num_tokens": 1756450.0,
|
| 151 |
+
"step": 650
|
| 152 |
+
},
|
| 153 |
+
{
|
| 154 |
+
"entropy": 0.5832360745966434,
|
| 155 |
+
"epoch": 1.7170042971147943,
|
| 156 |
+
"grad_norm": 0.35520103573799133,
|
| 157 |
+
"learning_rate": 0.0002135197792300053,
|
| 158 |
+
"loss": 0.5287040328979492,
|
| 159 |
+
"mean_token_accuracy": 0.8368027776479721,
|
| 160 |
+
"num_tokens": 1890634.0,
|
| 161 |
+
"step": 700
|
| 162 |
+
},
|
| 163 |
+
{
|
| 164 |
+
"entropy": 0.5792283065617084,
|
| 165 |
+
"epoch": 1.839779005524862,
|
| 166 |
+
"grad_norm": 0.3548352122306824,
|
| 167 |
+
"learning_rate": 0.00021228106743818178,
|
| 168 |
+
"loss": 0.5251237487792969,
|
| 169 |
+
"mean_token_accuracy": 0.8396679371595382,
|
| 170 |
+
"num_tokens": 2028046.0,
|
| 171 |
+
"step": 750
|
| 172 |
+
},
|
| 173 |
+
{
|
| 174 |
+
"entropy": 0.5694268302619457,
|
| 175 |
+
"epoch": 1.9625537139349294,
|
| 176 |
+
"grad_norm": 0.3476282060146332,
|
| 177 |
+
"learning_rate": 0.00021085234718892933,
|
| 178 |
+
"loss": 0.5189918899536132,
|
| 179 |
+
"mean_token_accuracy": 0.8398650133609772,
|
| 180 |
+
"num_tokens": 2163703.0,
|
| 181 |
+
"step": 800
|
| 182 |
+
},
|
| 183 |
+
{
|
| 184 |
+
"epoch": 2.0,
|
| 185 |
+
"eval_entropy": 0.5552395452771868,
|
| 186 |
+
"eval_loss": 0.5418588519096375,
|
| 187 |
+
"eval_mean_token_accuracy": 0.8379638079234532,
|
| 188 |
+
"eval_num_tokens": 2203536.0,
|
| 189 |
+
"eval_runtime": 108.6649,
|
| 190 |
+
"eval_samples_per_second": 12.847,
|
| 191 |
+
"eval_steps_per_second": 1.61,
|
| 192 |
+
"step": 816
|
| 193 |
+
},
|
| 194 |
+
{
|
| 195 |
+
"entropy": 0.5162206395023365,
|
| 196 |
+
"epoch": 2.0834868017188457,
|
| 197 |
+
"grad_norm": 0.3772931396961212,
|
| 198 |
+
"learning_rate": 0.0002092362325412188,
|
| 199 |
+
"loss": 0.4619992446899414,
|
| 200 |
+
"mean_token_accuracy": 0.8528054483650904,
|
| 201 |
+
"num_tokens": 2298231.0,
|
| 202 |
+
"step": 850
|
| 203 |
+
},
|
| 204 |
+
{
|
| 205 |
+
"entropy": 0.5109021583199501,
|
| 206 |
+
"epoch": 2.2062615101289134,
|
| 207 |
+
"grad_norm": 0.4661090672016144,
|
| 208 |
+
"learning_rate": 0.000207435680420309,
|
| 209 |
+
"loss": 0.4571444702148437,
|
| 210 |
+
"mean_token_accuracy": 0.8552933797240257,
|
| 211 |
+
"num_tokens": 2427608.0,
|
| 212 |
+
"step": 900
|
| 213 |
+
},
|
| 214 |
+
{
|
| 215 |
+
"entropy": 0.5012516237795352,
|
| 216 |
+
"epoch": 2.329036218538981,
|
| 217 |
+
"grad_norm": 0.5193169713020325,
|
| 218 |
+
"learning_rate": 0.0002054539852076065,
|
| 219 |
+
"loss": 0.45432735443115235,
|
| 220 |
+
"mean_token_accuracy": 0.8564087572693825,
|
| 221 |
+
"num_tokens": 2565982.0,
|
| 222 |
+
"step": 950
|
| 223 |
+
},
|
| 224 |
+
{
|
| 225 |
+
"entropy": 0.5173932652175427,
|
| 226 |
+
"epoch": 2.4518109269490487,
|
| 227 |
+
"grad_norm": 0.4666334390640259,
|
| 228 |
+
"learning_rate": 0.00020329477271309812,
|
| 229 |
+
"loss": 0.4616986083984375,
|
| 230 |
+
"mean_token_accuracy": 0.8532203987240792,
|
| 231 |
+
"num_tokens": 2695545.0,
|
| 232 |
+
"step": 1000
|
| 233 |
+
},
|
| 234 |
+
{
|
| 235 |
+
"entropy": 0.5169526914507151,
|
| 236 |
+
"epoch": 2.574585635359116,
|
| 237 |
+
"grad_norm": 0.47921115159988403,
|
| 238 |
+
"learning_rate": 0.0002009619935413857,
|
| 239 |
+
"loss": 0.45737281799316404,
|
| 240 |
+
"mean_token_accuracy": 0.8535552659630775,
|
| 241 |
+
"num_tokens": 2829781.0,
|
| 242 |
+
"step": 1050
|
| 243 |
+
},
|
| 244 |
+
{
|
| 245 |
+
"entropy": 0.50618562489748,
|
| 246 |
+
"epoch": 2.6973603437691835,
|
| 247 |
+
"grad_norm": 0.3407684862613678,
|
| 248 |
+
"learning_rate": 0.00019845991586345935,
|
| 249 |
+
"loss": 0.45972068786621095,
|
| 250 |
+
"mean_token_accuracy": 0.8532172521948814,
|
| 251 |
+
"num_tokens": 2970162.0,
|
| 252 |
+
"step": 1100
|
| 253 |
+
},
|
| 254 |
+
{
|
| 255 |
+
"entropy": 0.5158357314765454,
|
| 256 |
+
"epoch": 2.820135052179251,
|
| 257 |
+
"grad_norm": 0.47901391983032227,
|
| 258 |
+
"learning_rate": 0.00019579311760743563,
|
| 259 |
+
"loss": 0.46119583129882813,
|
| 260 |
+
"mean_token_accuracy": 0.8536588314175606,
|
| 261 |
+
"num_tokens": 3103775.0,
|
| 262 |
+
"step": 1150
|
| 263 |
+
},
|
| 264 |
+
{
|
| 265 |
+
"entropy": 0.5069115920364857,
|
| 266 |
+
"epoch": 2.942909760589319,
|
| 267 |
+
"grad_norm": 0.4287651479244232,
|
| 268 |
+
"learning_rate": 0.00019296647808254838,
|
| 269 |
+
"loss": 0.45447597503662107,
|
| 270 |
+
"mean_token_accuracy": 0.856622197329998,
|
| 271 |
+
"num_tokens": 3241634.0,
|
| 272 |
+
"step": 1200
|
| 273 |
+
},
|
| 274 |
+
{
|
| 275 |
+
"epoch": 3.0,
|
| 276 |
+
"eval_entropy": 0.5097703012398311,
|
| 277 |
+
"eval_loss": 0.5291272401809692,
|
| 278 |
+
"eval_mean_token_accuracy": 0.8438761404582432,
|
| 279 |
+
"eval_num_tokens": 3305304.0,
|
| 280 |
+
"eval_runtime": 108.6946,
|
| 281 |
+
"eval_samples_per_second": 12.843,
|
| 282 |
+
"eval_steps_per_second": 1.61,
|
| 283 |
+
"step": 1224
|
| 284 |
+
}
|
| 285 |
+
],
|
| 286 |
+
"logging_steps": 50,
|
| 287 |
+
"max_steps": 4080,
|
| 288 |
+
"num_input_tokens_seen": 0,
|
| 289 |
+
"num_train_epochs": 10,
|
| 290 |
+
"save_steps": 500,
|
| 291 |
+
"stateful_callbacks": {
|
| 292 |
+
"TrainerControl": {
|
| 293 |
+
"args": {
|
| 294 |
+
"should_epoch_stop": false,
|
| 295 |
+
"should_evaluate": false,
|
| 296 |
+
"should_log": false,
|
| 297 |
+
"should_save": true,
|
| 298 |
+
"should_training_stop": false
|
| 299 |
+
},
|
| 300 |
+
"attributes": {}
|
| 301 |
+
}
|
| 302 |
+
},
|
| 303 |
+
"total_flos": 5.376308378051789e+17,
|
| 304 |
+
"train_batch_size": 4,
|
| 305 |
+
"trial_name": null,
|
| 306 |
+
"trial_params": null
|
| 307 |
+
}
|
random_original_Estonian/Qwen3-14B-Base_random_splits_original_features_train_random_splits_original_features_test1/checkpoint-1632/README.md
ADDED
|
@@ -0,0 +1,209 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen3-14B-Base
|
| 3 |
+
library_name: peft
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:Qwen/Qwen3-14B-Base
|
| 7 |
+
- lora
|
| 8 |
+
- sft
|
| 9 |
+
- transformers
|
| 10 |
+
- trl
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Model Card for Model ID
|
| 14 |
+
|
| 15 |
+
<!-- Provide a quick summary of what the model is/does. -->
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
## Model Details
|
| 20 |
+
|
| 21 |
+
### Model Description
|
| 22 |
+
|
| 23 |
+
<!-- Provide a longer summary of what this model is. -->
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
- **Developed by:** [More Information Needed]
|
| 28 |
+
- **Funded by [optional]:** [More Information Needed]
|
| 29 |
+
- **Shared by [optional]:** [More Information Needed]
|
| 30 |
+
- **Model type:** [More Information Needed]
|
| 31 |
+
- **Language(s) (NLP):** [More Information Needed]
|
| 32 |
+
- **License:** [More Information Needed]
|
| 33 |
+
- **Finetuned from model [optional]:** [More Information Needed]
|
| 34 |
+
|
| 35 |
+
### Model Sources [optional]
|
| 36 |
+
|
| 37 |
+
<!-- Provide the basic links for the model. -->
|
| 38 |
+
|
| 39 |
+
- **Repository:** [More Information Needed]
|
| 40 |
+
- **Paper [optional]:** [More Information Needed]
|
| 41 |
+
- **Demo [optional]:** [More Information Needed]
|
| 42 |
+
|
| 43 |
+
## Uses
|
| 44 |
+
|
| 45 |
+
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
| 46 |
+
|
| 47 |
+
### Direct Use
|
| 48 |
+
|
| 49 |
+
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
| 50 |
+
|
| 51 |
+
[More Information Needed]
|
| 52 |
+
|
| 53 |
+
### Downstream Use [optional]
|
| 54 |
+
|
| 55 |
+
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
| 56 |
+
|
| 57 |
+
[More Information Needed]
|
| 58 |
+
|
| 59 |
+
### Out-of-Scope Use
|
| 60 |
+
|
| 61 |
+
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
| 62 |
+
|
| 63 |
+
[More Information Needed]
|
| 64 |
+
|
| 65 |
+
## Bias, Risks, and Limitations
|
| 66 |
+
|
| 67 |
+
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
| 68 |
+
|
| 69 |
+
[More Information Needed]
|
| 70 |
+
|
| 71 |
+
### Recommendations
|
| 72 |
+
|
| 73 |
+
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
| 74 |
+
|
| 75 |
+
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
| 76 |
+
|
| 77 |
+
## How to Get Started with the Model
|
| 78 |
+
|
| 79 |
+
Use the code below to get started with the model.
|
| 80 |
+
|
| 81 |
+
[More Information Needed]
|
| 82 |
+
|
| 83 |
+
## Training Details
|
| 84 |
+
|
| 85 |
+
### Training Data
|
| 86 |
+
|
| 87 |
+
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
| 88 |
+
|
| 89 |
+
[More Information Needed]
|
| 90 |
+
|
| 91 |
+
### Training Procedure
|
| 92 |
+
|
| 93 |
+
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
| 94 |
+
|
| 95 |
+
#### Preprocessing [optional]
|
| 96 |
+
|
| 97 |
+
[More Information Needed]
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
#### Training Hyperparameters
|
| 101 |
+
|
| 102 |
+
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
| 103 |
+
|
| 104 |
+
#### Speeds, Sizes, Times [optional]
|
| 105 |
+
|
| 106 |
+
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
| 107 |
+
|
| 108 |
+
[More Information Needed]
|
| 109 |
+
|
| 110 |
+
## Evaluation
|
| 111 |
+
|
| 112 |
+
<!-- This section describes the evaluation protocols and provides the results. -->
|
| 113 |
+
|
| 114 |
+
### Testing Data, Factors & Metrics
|
| 115 |
+
|
| 116 |
+
#### Testing Data
|
| 117 |
+
|
| 118 |
+
<!-- This should link to a Dataset Card if possible. -->
|
| 119 |
+
|
| 120 |
+
[More Information Needed]
|
| 121 |
+
|
| 122 |
+
#### Factors
|
| 123 |
+
|
| 124 |
+
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
| 125 |
+
|
| 126 |
+
[More Information Needed]
|
| 127 |
+
|
| 128 |
+
#### Metrics
|
| 129 |
+
|
| 130 |
+
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
| 131 |
+
|
| 132 |
+
[More Information Needed]
|
| 133 |
+
|
| 134 |
+
### Results
|
| 135 |
+
|
| 136 |
+
[More Information Needed]
|
| 137 |
+
|
| 138 |
+
#### Summary
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
## Model Examination [optional]
|
| 143 |
+
|
| 144 |
+
<!-- Relevant interpretability work for the model goes here -->
|
| 145 |
+
|
| 146 |
+
[More Information Needed]
|
| 147 |
+
|
| 148 |
+
## Environmental Impact
|
| 149 |
+
|
| 150 |
+
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
| 151 |
+
|
| 152 |
+
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
| 153 |
+
|
| 154 |
+
- **Hardware Type:** [More Information Needed]
|
| 155 |
+
- **Hours used:** [More Information Needed]
|
| 156 |
+
- **Cloud Provider:** [More Information Needed]
|
| 157 |
+
- **Compute Region:** [More Information Needed]
|
| 158 |
+
- **Carbon Emitted:** [More Information Needed]
|
| 159 |
+
|
| 160 |
+
## Technical Specifications [optional]
|
| 161 |
+
|
| 162 |
+
### Model Architecture and Objective
|
| 163 |
+
|
| 164 |
+
[More Information Needed]
|
| 165 |
+
|
| 166 |
+
### Compute Infrastructure
|
| 167 |
+
|
| 168 |
+
[More Information Needed]
|
| 169 |
+
|
| 170 |
+
#### Hardware
|
| 171 |
+
|
| 172 |
+
[More Information Needed]
|
| 173 |
+
|
| 174 |
+
#### Software
|
| 175 |
+
|
| 176 |
+
[More Information Needed]
|
| 177 |
+
|
| 178 |
+
## Citation [optional]
|
| 179 |
+
|
| 180 |
+
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
| 181 |
+
|
| 182 |
+
**BibTeX:**
|
| 183 |
+
|
| 184 |
+
[More Information Needed]
|
| 185 |
+
|
| 186 |
+
**APA:**
|
| 187 |
+
|
| 188 |
+
[More Information Needed]
|
| 189 |
+
|
| 190 |
+
## Glossary [optional]
|
| 191 |
+
|
| 192 |
+
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
| 193 |
+
|
| 194 |
+
[More Information Needed]
|
| 195 |
+
|
| 196 |
+
## More Information [optional]
|
| 197 |
+
|
| 198 |
+
[More Information Needed]
|
| 199 |
+
|
| 200 |
+
## Model Card Authors [optional]
|
| 201 |
+
|
| 202 |
+
[More Information Needed]
|
| 203 |
+
|
| 204 |
+
## Model Card Contact
|
| 205 |
+
|
| 206 |
+
[More Information Needed]
|
| 207 |
+
### Framework versions
|
| 208 |
+
|
| 209 |
+
- PEFT 0.19.1
|