diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..6aeed1166009b0ead937d4123341c3fba6a3e5dc
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/README.md
@@ -0,0 +1,58 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: transformers
+model_name: Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1
+tags:
+- generated_from_trainer
+- trl
+- sft
+licence: license
+---
+
+# Model Card for Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1
+
+This model is a fine-tuned version of [Qwen/Qwen3.5-4B-Base](https://huggingface.co/Qwen/Qwen3.5-4B-Base).
+It has been trained using [TRL](https://github.com/huggingface/trl).
+
+## Quick start
+
+```python
+from transformers import pipeline
+
+question = "If you had a time machine, but could only go to the past or the future once and never return, which would you choose and why?"
+generator = pipeline("text-generation", model="None", device="cuda")
+output = generator([{"role": "user", "content": question}], max_new_tokens=128, return_full_text=False)[0]
+print(output["generated_text"])
+```
+
+## Training procedure
+
+[
](https://wandb.ai/katriin-kukk/Cross_lingual_morphological_generalization/runs/1e798q43)
+
+
+
+This model was trained with SFT.
+
+### Framework versions
+
+- TRL: 0.29.0
+- Transformers: 5.5.4
+- Pytorch: 2.10.0
+- Datasets: 4.6.1
+- Tokenizers: 0.22.2
+
+## Citations
+
+
+
+Cite TRL as:
+
+```bibtex
+@software{vonwerra2020trl,
+ title = {{TRL: Transformers Reinforcement Learning}},
+ author = {von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and Beeching, Edward and Thrush, Tristan and Lambert, Nathan and Huang, Shengyi and Rasul, Kashif and Gallouédec, Quentin},
+ license = {Apache-2.0},
+ url = {https://github.com/huggingface/trl},
+ year = {2020}
+}
+```
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1155/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1155/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1155/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1155/adapter_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1155/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..13015f013b98b31173d1f215c91d6b5ff9f679cc
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1155/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 32,
+ "lora_bias": false,
+ "lora_dropout": 0.08930249338906042,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 16,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "k_proj",
+ "down_proj",
+ "q_proj",
+ "v_proj",
+ "up_proj",
+ "gate_proj",
+ "o_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1155/chat_template.jinja b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1155/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1155/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1155/tokenizer_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1155/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1155/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1155/trainer_state.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1155/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..f8a816bb183a5bdc87ab5aba7092b09d33cc387e
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1155/trainer_state.json
@@ -0,0 +1,297 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 3.0,
+ "eval_steps": 500,
+ "global_step": 1155,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.8437769564986228,
+ "epoch": 0.1299545159194282,
+ "grad_norm": 1.6047701835632324,
+ "learning_rate": 4.2543894243185634e-05,
+ "loss": 1.6604945373535156,
+ "mean_token_accuracy": 0.6506646998226643,
+ "num_tokens": 127115.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.9294079938530921,
+ "epoch": 0.2599090318388564,
+ "grad_norm": 1.1633784770965576,
+ "learning_rate": 8.595603122602811e-05,
+ "loss": 0.8363616180419922,
+ "mean_token_accuracy": 0.7776120388507843,
+ "num_tokens": 251796.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.8483590838313103,
+ "epoch": 0.3898635477582846,
+ "grad_norm": 0.8963253498077393,
+ "learning_rate": 0.00012936816820887062,
+ "loss": 0.7476110076904297,
+ "mean_token_accuracy": 0.7931549972295762,
+ "num_tokens": 366505.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.7840312227606774,
+ "epoch": 0.5198180636777128,
+ "grad_norm": 0.7876676917076111,
+ "learning_rate": 0.00017278030519171308,
+ "loss": 0.6980364990234375,
+ "mean_token_accuracy": 0.8027918112277984,
+ "num_tokens": 493150.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.7967149233818054,
+ "epoch": 0.649772579597141,
+ "grad_norm": 0.7976633906364441,
+ "learning_rate": 0.00021619244217455558,
+ "loss": 0.6958013153076172,
+ "mean_token_accuracy": 0.8039823538064956,
+ "num_tokens": 607197.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.7755034640431404,
+ "epoch": 0.7797270955165692,
+ "grad_norm": 0.7086226940155029,
+ "learning_rate": 0.00025960457915739807,
+ "loss": 0.6792864990234375,
+ "mean_token_accuracy": 0.8074422472715378,
+ "num_tokens": 726579.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.7452501457929611,
+ "epoch": 0.9096816114359974,
+ "grad_norm": 0.6065998673439026,
+ "learning_rate": 0.00030301671614024056,
+ "loss": 0.6627902984619141,
+ "mean_token_accuracy": 0.8118429997563362,
+ "num_tokens": 848944.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.7490686719807295,
+ "eval_loss": 0.718986451625824,
+ "eval_mean_token_accuracy": 0.7978653156986604,
+ "eval_num_tokens": 931087.0,
+ "eval_runtime": 93.2605,
+ "eval_samples_per_second": 17.767,
+ "eval_steps_per_second": 2.23,
+ "step": 385
+ },
+ {
+ "entropy": 0.7291471499893534,
+ "epoch": 1.0389863547758285,
+ "grad_norm": 0.7773680090904236,
+ "learning_rate": 0.0003342599904174574,
+ "loss": 0.6407540893554687,
+ "mean_token_accuracy": 0.816381629387937,
+ "num_tokens": 968580.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.6813811221718789,
+ "epoch": 1.1689408706952567,
+ "grad_norm": 0.7393642067909241,
+ "learning_rate": 0.0003339921524880377,
+ "loss": 0.6079624938964844,
+ "mean_token_accuracy": 0.825104131102562,
+ "num_tokens": 1090002.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.682051141411066,
+ "epoch": 1.2988953866146848,
+ "grad_norm": 0.6833657026290894,
+ "learning_rate": 0.0003333814684455912,
+ "loss": 0.6068630218505859,
+ "mean_token_accuracy": 0.8236467111110687,
+ "num_tokens": 1212161.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.6863492175936698,
+ "epoch": 1.428849902534113,
+ "grad_norm": 0.6060423851013184,
+ "learning_rate": 0.0003324291930928916,
+ "loss": 0.6073496627807617,
+ "mean_token_accuracy": 0.8244831630587578,
+ "num_tokens": 1332852.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.6790362980961799,
+ "epoch": 1.5588044184535412,
+ "grad_norm": 0.7274985313415527,
+ "learning_rate": 0.00033113728311731064,
+ "loss": 0.5929213714599609,
+ "mean_token_accuracy": 0.8260384351015091,
+ "num_tokens": 1453569.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.6581329819560051,
+ "epoch": 1.6887589343729694,
+ "grad_norm": 0.6212530136108398,
+ "learning_rate": 0.00032950839307031545,
+ "loss": 0.5812822341918945,
+ "mean_token_accuracy": 0.8288319337368012,
+ "num_tokens": 1577152.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.6893622297048568,
+ "epoch": 1.8187134502923976,
+ "grad_norm": 0.5525115132331848,
+ "learning_rate": 0.0003275458699130301,
+ "loss": 0.6035601425170899,
+ "mean_token_accuracy": 0.8239563027024269,
+ "num_tokens": 1691926.0,
+ "step": 700
+ },
+ {
+ "entropy": 0.6713312044739723,
+ "epoch": 1.9486679662118258,
+ "grad_norm": 0.746279776096344,
+ "learning_rate": 0.0003252537461390686,
+ "loss": 0.59674560546875,
+ "mean_token_accuracy": 0.8267503470182419,
+ "num_tokens": 1816059.0,
+ "step": 750
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6752213362890941,
+ "eval_loss": 0.6740940809249878,
+ "eval_mean_token_accuracy": 0.8084349325643136,
+ "eval_num_tokens": 1862174.0,
+ "eval_runtime": 94.8197,
+ "eval_samples_per_second": 17.475,
+ "eval_steps_per_second": 2.194,
+ "step": 770
+ },
+ {
+ "entropy": 0.6279655678487902,
+ "epoch": 2.077972709551657,
+ "grad_norm": 0.7196288108825684,
+ "learning_rate": 0.00032263673148877055,
+ "loss": 0.5418233871459961,
+ "mean_token_accuracy": 0.8383135430177852,
+ "num_tokens": 1931345.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.5900149047374725,
+ "epoch": 2.207927225471085,
+ "grad_norm": 0.6000070571899414,
+ "learning_rate": 0.00031970020327186393,
+ "loss": 0.5065702056884765,
+ "mean_token_accuracy": 0.8456701630353928,
+ "num_tokens": 2057087.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.5950760287046433,
+ "epoch": 2.3378817413905133,
+ "grad_norm": 0.5828987956047058,
+ "learning_rate": 0.0003164501953184404,
+ "loss": 0.5115758514404297,
+ "mean_token_accuracy": 0.8449143546819687,
+ "num_tokens": 2179801.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.5921229894459248,
+ "epoch": 2.4678362573099415,
+ "grad_norm": 0.6523755788803101,
+ "learning_rate": 0.00031289338558094566,
+ "loss": 0.5099533462524414,
+ "mean_token_accuracy": 0.8458523535728455,
+ "num_tokens": 2300775.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.5937536461651325,
+ "epoch": 2.5977907732293697,
+ "grad_norm": 0.6264522075653076,
+ "learning_rate": 0.0003090370824126593,
+ "loss": 0.5148015594482422,
+ "mean_token_accuracy": 0.8433699604868888,
+ "num_tokens": 2425909.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.594799654185772,
+ "epoch": 2.727745289148798,
+ "grad_norm": 0.867798388004303,
+ "learning_rate": 0.000304889209550859,
+ "loss": 0.5112411499023437,
+ "mean_token_accuracy": 0.844317550957203,
+ "num_tokens": 2542506.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.5839655491709709,
+ "epoch": 2.857699805068226,
+ "grad_norm": 0.5512472987174988,
+ "learning_rate": 0.0003004582898355252,
+ "loss": 0.5124132537841797,
+ "mean_token_accuracy": 0.8445275443792343,
+ "num_tokens": 2666261.0,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6069950518012047,
+ "epoch": 2.9876543209876543,
+ "grad_norm": 0.6030759215354919,
+ "learning_rate": 0.00029575342769703865,
+ "loss": 0.521760368347168,
+ "mean_token_accuracy": 0.8433170795440674,
+ "num_tokens": 2781874.0,
+ "step": 1150
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.6072480179942571,
+ "eval_loss": 0.666907787322998,
+ "eval_mean_token_accuracy": 0.8143598649364251,
+ "eval_num_tokens": 2793261.0,
+ "eval_runtime": 95.3076,
+ "eval_samples_per_second": 17.386,
+ "eval_steps_per_second": 2.182,
+ "step": 1155
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3850,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.1758758062337024e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1540/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1540/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1540/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1540/adapter_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1540/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..13015f013b98b31173d1f215c91d6b5ff9f679cc
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1540/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 32,
+ "lora_bias": false,
+ "lora_dropout": 0.08930249338906042,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 16,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "k_proj",
+ "down_proj",
+ "q_proj",
+ "v_proj",
+ "up_proj",
+ "gate_proj",
+ "o_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1540/chat_template.jinja b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1540/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1540/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1540/tokenizer_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1540/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1540/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1540/trainer_state.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1540/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..a0e0fb6ae212cd76a1b68fc4a53b663dd06a6e31
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1540/trainer_state.json
@@ -0,0 +1,378 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 4.0,
+ "eval_steps": 500,
+ "global_step": 1540,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.8437769564986228,
+ "epoch": 0.1299545159194282,
+ "grad_norm": 1.6047701835632324,
+ "learning_rate": 4.2543894243185634e-05,
+ "loss": 1.6604945373535156,
+ "mean_token_accuracy": 0.6506646998226643,
+ "num_tokens": 127115.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.9294079938530921,
+ "epoch": 0.2599090318388564,
+ "grad_norm": 1.1633784770965576,
+ "learning_rate": 8.595603122602811e-05,
+ "loss": 0.8363616180419922,
+ "mean_token_accuracy": 0.7776120388507843,
+ "num_tokens": 251796.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.8483590838313103,
+ "epoch": 0.3898635477582846,
+ "grad_norm": 0.8963253498077393,
+ "learning_rate": 0.00012936816820887062,
+ "loss": 0.7476110076904297,
+ "mean_token_accuracy": 0.7931549972295762,
+ "num_tokens": 366505.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.7840312227606774,
+ "epoch": 0.5198180636777128,
+ "grad_norm": 0.7876676917076111,
+ "learning_rate": 0.00017278030519171308,
+ "loss": 0.6980364990234375,
+ "mean_token_accuracy": 0.8027918112277984,
+ "num_tokens": 493150.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.7967149233818054,
+ "epoch": 0.649772579597141,
+ "grad_norm": 0.7976633906364441,
+ "learning_rate": 0.00021619244217455558,
+ "loss": 0.6958013153076172,
+ "mean_token_accuracy": 0.8039823538064956,
+ "num_tokens": 607197.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.7755034640431404,
+ "epoch": 0.7797270955165692,
+ "grad_norm": 0.7086226940155029,
+ "learning_rate": 0.00025960457915739807,
+ "loss": 0.6792864990234375,
+ "mean_token_accuracy": 0.8074422472715378,
+ "num_tokens": 726579.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.7452501457929611,
+ "epoch": 0.9096816114359974,
+ "grad_norm": 0.6065998673439026,
+ "learning_rate": 0.00030301671614024056,
+ "loss": 0.6627902984619141,
+ "mean_token_accuracy": 0.8118429997563362,
+ "num_tokens": 848944.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.7490686719807295,
+ "eval_loss": 0.718986451625824,
+ "eval_mean_token_accuracy": 0.7978653156986604,
+ "eval_num_tokens": 931087.0,
+ "eval_runtime": 93.2605,
+ "eval_samples_per_second": 17.767,
+ "eval_steps_per_second": 2.23,
+ "step": 385
+ },
+ {
+ "entropy": 0.7291471499893534,
+ "epoch": 1.0389863547758285,
+ "grad_norm": 0.7773680090904236,
+ "learning_rate": 0.0003342599904174574,
+ "loss": 0.6407540893554687,
+ "mean_token_accuracy": 0.816381629387937,
+ "num_tokens": 968580.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.6813811221718789,
+ "epoch": 1.1689408706952567,
+ "grad_norm": 0.7393642067909241,
+ "learning_rate": 0.0003339921524880377,
+ "loss": 0.6079624938964844,
+ "mean_token_accuracy": 0.825104131102562,
+ "num_tokens": 1090002.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.682051141411066,
+ "epoch": 1.2988953866146848,
+ "grad_norm": 0.6833657026290894,
+ "learning_rate": 0.0003333814684455912,
+ "loss": 0.6068630218505859,
+ "mean_token_accuracy": 0.8236467111110687,
+ "num_tokens": 1212161.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.6863492175936698,
+ "epoch": 1.428849902534113,
+ "grad_norm": 0.6060423851013184,
+ "learning_rate": 0.0003324291930928916,
+ "loss": 0.6073496627807617,
+ "mean_token_accuracy": 0.8244831630587578,
+ "num_tokens": 1332852.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.6790362980961799,
+ "epoch": 1.5588044184535412,
+ "grad_norm": 0.7274985313415527,
+ "learning_rate": 0.00033113728311731064,
+ "loss": 0.5929213714599609,
+ "mean_token_accuracy": 0.8260384351015091,
+ "num_tokens": 1453569.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.6581329819560051,
+ "epoch": 1.6887589343729694,
+ "grad_norm": 0.6212530136108398,
+ "learning_rate": 0.00032950839307031545,
+ "loss": 0.5812822341918945,
+ "mean_token_accuracy": 0.8288319337368012,
+ "num_tokens": 1577152.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.6893622297048568,
+ "epoch": 1.8187134502923976,
+ "grad_norm": 0.5525115132331848,
+ "learning_rate": 0.0003275458699130301,
+ "loss": 0.6035601425170899,
+ "mean_token_accuracy": 0.8239563027024269,
+ "num_tokens": 1691926.0,
+ "step": 700
+ },
+ {
+ "entropy": 0.6713312044739723,
+ "epoch": 1.9486679662118258,
+ "grad_norm": 0.746279776096344,
+ "learning_rate": 0.0003252537461390686,
+ "loss": 0.59674560546875,
+ "mean_token_accuracy": 0.8267503470182419,
+ "num_tokens": 1816059.0,
+ "step": 750
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6752213362890941,
+ "eval_loss": 0.6740940809249878,
+ "eval_mean_token_accuracy": 0.8084349325643136,
+ "eval_num_tokens": 1862174.0,
+ "eval_runtime": 94.8197,
+ "eval_samples_per_second": 17.475,
+ "eval_steps_per_second": 2.194,
+ "step": 770
+ },
+ {
+ "entropy": 0.6279655678487902,
+ "epoch": 2.077972709551657,
+ "grad_norm": 0.7196288108825684,
+ "learning_rate": 0.00032263673148877055,
+ "loss": 0.5418233871459961,
+ "mean_token_accuracy": 0.8383135430177852,
+ "num_tokens": 1931345.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.5900149047374725,
+ "epoch": 2.207927225471085,
+ "grad_norm": 0.6000070571899414,
+ "learning_rate": 0.00031970020327186393,
+ "loss": 0.5065702056884765,
+ "mean_token_accuracy": 0.8456701630353928,
+ "num_tokens": 2057087.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.5950760287046433,
+ "epoch": 2.3378817413905133,
+ "grad_norm": 0.5828987956047058,
+ "learning_rate": 0.0003164501953184404,
+ "loss": 0.5115758514404297,
+ "mean_token_accuracy": 0.8449143546819687,
+ "num_tokens": 2179801.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.5921229894459248,
+ "epoch": 2.4678362573099415,
+ "grad_norm": 0.6523755788803101,
+ "learning_rate": 0.00031289338558094566,
+ "loss": 0.5099533462524414,
+ "mean_token_accuracy": 0.8458523535728455,
+ "num_tokens": 2300775.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.5937536461651325,
+ "epoch": 2.5977907732293697,
+ "grad_norm": 0.6264522075653076,
+ "learning_rate": 0.0003090370824126593,
+ "loss": 0.5148015594482422,
+ "mean_token_accuracy": 0.8433699604868888,
+ "num_tokens": 2425909.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.594799654185772,
+ "epoch": 2.727745289148798,
+ "grad_norm": 0.867798388004303,
+ "learning_rate": 0.000304889209550859,
+ "loss": 0.5112411499023437,
+ "mean_token_accuracy": 0.844317550957203,
+ "num_tokens": 2542506.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.5839655491709709,
+ "epoch": 2.857699805068226,
+ "grad_norm": 0.5512472987174988,
+ "learning_rate": 0.0003004582898355252,
+ "loss": 0.5124132537841797,
+ "mean_token_accuracy": 0.8445275443792343,
+ "num_tokens": 2666261.0,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6069950518012047,
+ "epoch": 2.9876543209876543,
+ "grad_norm": 0.6030759215354919,
+ "learning_rate": 0.00029575342769703865,
+ "loss": 0.521760368347168,
+ "mean_token_accuracy": 0.8433170795440674,
+ "num_tokens": 2781874.0,
+ "step": 1150
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.6072480179942571,
+ "eval_loss": 0.666907787322998,
+ "eval_mean_token_accuracy": 0.8143598649364251,
+ "eval_num_tokens": 2793261.0,
+ "eval_runtime": 95.3076,
+ "eval_samples_per_second": 17.386,
+ "eval_steps_per_second": 2.182,
+ "step": 1155
+ },
+ {
+ "entropy": 0.5205143204885512,
+ "epoch": 3.116959064327485,
+ "grad_norm": 0.663673996925354,
+ "learning_rate": 0.00029078429044885647,
+ "loss": 0.42832801818847654,
+ "mean_token_accuracy": 0.8651104230976584,
+ "num_tokens": 2896983.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.5157249197363853,
+ "epoch": 3.246913580246914,
+ "grad_norm": 0.8357914686203003,
+ "learning_rate": 0.00028556108842360404,
+ "loss": 0.4264420700073242,
+ "mean_token_accuracy": 0.8651721149682998,
+ "num_tokens": 3019005.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.5092187704145908,
+ "epoch": 3.3768680961663415,
+ "grad_norm": 0.7227477431297302,
+ "learning_rate": 0.00028009455399339797,
+ "loss": 0.4299094772338867,
+ "mean_token_accuracy": 0.8649037969112396,
+ "num_tokens": 3143415.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5237507157027721,
+ "epoch": 3.50682261208577,
+ "grad_norm": 0.6605198383331299,
+ "learning_rate": 0.00027439591951750953,
+ "loss": 0.44015567779541015,
+ "mean_token_accuracy": 0.8605978584289551,
+ "num_tokens": 3263845.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.5314217877388,
+ "epoch": 3.636777128005198,
+ "grad_norm": 0.9264621734619141,
+ "learning_rate": 0.00026847689426267907,
+ "loss": 0.4449501037597656,
+ "mean_token_accuracy": 0.862554576098919,
+ "num_tokens": 3383117.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.5156177791953087,
+ "epoch": 3.7667316439246266,
+ "grad_norm": 0.674964964389801,
+ "learning_rate": 0.0002623496403435053,
+ "loss": 0.4325400161743164,
+ "mean_token_accuracy": 0.8629470297694206,
+ "num_tokens": 3509724.0,
+ "step": 1450
+ },
+ {
+ "entropy": 0.5290651166439057,
+ "epoch": 3.8966861598440543,
+ "grad_norm": 0.8422386050224304,
+ "learning_rate": 0.00025602674773234535,
+ "loss": 0.44204349517822267,
+ "mean_token_accuracy": 0.8624594563245773,
+ "num_tokens": 3631110.0,
+ "step": 1500
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5855578513672719,
+ "eval_loss": 0.6782004833221436,
+ "eval_mean_token_accuracy": 0.8134025845390099,
+ "eval_num_tokens": 3724348.0,
+ "eval_runtime": 95.7097,
+ "eval_samples_per_second": 17.313,
+ "eval_steps_per_second": 2.173,
+ "step": 1540
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3850,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.5658655218065408e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1925/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1925/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1925/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1925/adapter_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1925/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..13015f013b98b31173d1f215c91d6b5ff9f679cc
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1925/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 32,
+ "lora_bias": false,
+ "lora_dropout": 0.08930249338906042,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 16,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "k_proj",
+ "down_proj",
+ "q_proj",
+ "v_proj",
+ "up_proj",
+ "gate_proj",
+ "o_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1925/chat_template.jinja b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1925/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1925/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1925/tokenizer_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1925/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1925/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1925/trainer_state.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1925/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..f3a3c1a3e1e3cf76be7acb14e204abf232333a1c
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1925/trainer_state.json
@@ -0,0 +1,469 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 5.0,
+ "eval_steps": 500,
+ "global_step": 1925,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.8437769564986228,
+ "epoch": 0.1299545159194282,
+ "grad_norm": 1.6047701835632324,
+ "learning_rate": 4.2543894243185634e-05,
+ "loss": 1.6604945373535156,
+ "mean_token_accuracy": 0.6506646998226643,
+ "num_tokens": 127115.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.9294079938530921,
+ "epoch": 0.2599090318388564,
+ "grad_norm": 1.1633784770965576,
+ "learning_rate": 8.595603122602811e-05,
+ "loss": 0.8363616180419922,
+ "mean_token_accuracy": 0.7776120388507843,
+ "num_tokens": 251796.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.8483590838313103,
+ "epoch": 0.3898635477582846,
+ "grad_norm": 0.8963253498077393,
+ "learning_rate": 0.00012936816820887062,
+ "loss": 0.7476110076904297,
+ "mean_token_accuracy": 0.7931549972295762,
+ "num_tokens": 366505.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.7840312227606774,
+ "epoch": 0.5198180636777128,
+ "grad_norm": 0.7876676917076111,
+ "learning_rate": 0.00017278030519171308,
+ "loss": 0.6980364990234375,
+ "mean_token_accuracy": 0.8027918112277984,
+ "num_tokens": 493150.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.7967149233818054,
+ "epoch": 0.649772579597141,
+ "grad_norm": 0.7976633906364441,
+ "learning_rate": 0.00021619244217455558,
+ "loss": 0.6958013153076172,
+ "mean_token_accuracy": 0.8039823538064956,
+ "num_tokens": 607197.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.7755034640431404,
+ "epoch": 0.7797270955165692,
+ "grad_norm": 0.7086226940155029,
+ "learning_rate": 0.00025960457915739807,
+ "loss": 0.6792864990234375,
+ "mean_token_accuracy": 0.8074422472715378,
+ "num_tokens": 726579.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.7452501457929611,
+ "epoch": 0.9096816114359974,
+ "grad_norm": 0.6065998673439026,
+ "learning_rate": 0.00030301671614024056,
+ "loss": 0.6627902984619141,
+ "mean_token_accuracy": 0.8118429997563362,
+ "num_tokens": 848944.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.7490686719807295,
+ "eval_loss": 0.718986451625824,
+ "eval_mean_token_accuracy": 0.7978653156986604,
+ "eval_num_tokens": 931087.0,
+ "eval_runtime": 93.2605,
+ "eval_samples_per_second": 17.767,
+ "eval_steps_per_second": 2.23,
+ "step": 385
+ },
+ {
+ "entropy": 0.7291471499893534,
+ "epoch": 1.0389863547758285,
+ "grad_norm": 0.7773680090904236,
+ "learning_rate": 0.0003342599904174574,
+ "loss": 0.6407540893554687,
+ "mean_token_accuracy": 0.816381629387937,
+ "num_tokens": 968580.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.6813811221718789,
+ "epoch": 1.1689408706952567,
+ "grad_norm": 0.7393642067909241,
+ "learning_rate": 0.0003339921524880377,
+ "loss": 0.6079624938964844,
+ "mean_token_accuracy": 0.825104131102562,
+ "num_tokens": 1090002.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.682051141411066,
+ "epoch": 1.2988953866146848,
+ "grad_norm": 0.6833657026290894,
+ "learning_rate": 0.0003333814684455912,
+ "loss": 0.6068630218505859,
+ "mean_token_accuracy": 0.8236467111110687,
+ "num_tokens": 1212161.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.6863492175936698,
+ "epoch": 1.428849902534113,
+ "grad_norm": 0.6060423851013184,
+ "learning_rate": 0.0003324291930928916,
+ "loss": 0.6073496627807617,
+ "mean_token_accuracy": 0.8244831630587578,
+ "num_tokens": 1332852.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.6790362980961799,
+ "epoch": 1.5588044184535412,
+ "grad_norm": 0.7274985313415527,
+ "learning_rate": 0.00033113728311731064,
+ "loss": 0.5929213714599609,
+ "mean_token_accuracy": 0.8260384351015091,
+ "num_tokens": 1453569.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.6581329819560051,
+ "epoch": 1.6887589343729694,
+ "grad_norm": 0.6212530136108398,
+ "learning_rate": 0.00032950839307031545,
+ "loss": 0.5812822341918945,
+ "mean_token_accuracy": 0.8288319337368012,
+ "num_tokens": 1577152.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.6893622297048568,
+ "epoch": 1.8187134502923976,
+ "grad_norm": 0.5525115132331848,
+ "learning_rate": 0.0003275458699130301,
+ "loss": 0.6035601425170899,
+ "mean_token_accuracy": 0.8239563027024269,
+ "num_tokens": 1691926.0,
+ "step": 700
+ },
+ {
+ "entropy": 0.6713312044739723,
+ "epoch": 1.9486679662118258,
+ "grad_norm": 0.746279776096344,
+ "learning_rate": 0.0003252537461390686,
+ "loss": 0.59674560546875,
+ "mean_token_accuracy": 0.8267503470182419,
+ "num_tokens": 1816059.0,
+ "step": 750
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6752213362890941,
+ "eval_loss": 0.6740940809249878,
+ "eval_mean_token_accuracy": 0.8084349325643136,
+ "eval_num_tokens": 1862174.0,
+ "eval_runtime": 94.8197,
+ "eval_samples_per_second": 17.475,
+ "eval_steps_per_second": 2.194,
+ "step": 770
+ },
+ {
+ "entropy": 0.6279655678487902,
+ "epoch": 2.077972709551657,
+ "grad_norm": 0.7196288108825684,
+ "learning_rate": 0.00032263673148877055,
+ "loss": 0.5418233871459961,
+ "mean_token_accuracy": 0.8383135430177852,
+ "num_tokens": 1931345.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.5900149047374725,
+ "epoch": 2.207927225471085,
+ "grad_norm": 0.6000070571899414,
+ "learning_rate": 0.00031970020327186393,
+ "loss": 0.5065702056884765,
+ "mean_token_accuracy": 0.8456701630353928,
+ "num_tokens": 2057087.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.5950760287046433,
+ "epoch": 2.3378817413905133,
+ "grad_norm": 0.5828987956047058,
+ "learning_rate": 0.0003164501953184404,
+ "loss": 0.5115758514404297,
+ "mean_token_accuracy": 0.8449143546819687,
+ "num_tokens": 2179801.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.5921229894459248,
+ "epoch": 2.4678362573099415,
+ "grad_norm": 0.6523755788803101,
+ "learning_rate": 0.00031289338558094566,
+ "loss": 0.5099533462524414,
+ "mean_token_accuracy": 0.8458523535728455,
+ "num_tokens": 2300775.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.5937536461651325,
+ "epoch": 2.5977907732293697,
+ "grad_norm": 0.6264522075653076,
+ "learning_rate": 0.0003090370824126593,
+ "loss": 0.5148015594482422,
+ "mean_token_accuracy": 0.8433699604868888,
+ "num_tokens": 2425909.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.594799654185772,
+ "epoch": 2.727745289148798,
+ "grad_norm": 0.867798388004303,
+ "learning_rate": 0.000304889209550859,
+ "loss": 0.5112411499023437,
+ "mean_token_accuracy": 0.844317550957203,
+ "num_tokens": 2542506.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.5839655491709709,
+ "epoch": 2.857699805068226,
+ "grad_norm": 0.5512472987174988,
+ "learning_rate": 0.0003004582898355252,
+ "loss": 0.5124132537841797,
+ "mean_token_accuracy": 0.8445275443792343,
+ "num_tokens": 2666261.0,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6069950518012047,
+ "epoch": 2.9876543209876543,
+ "grad_norm": 0.6030759215354919,
+ "learning_rate": 0.00029575342769703865,
+ "loss": 0.521760368347168,
+ "mean_token_accuracy": 0.8433170795440674,
+ "num_tokens": 2781874.0,
+ "step": 1150
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.6072480179942571,
+ "eval_loss": 0.666907787322998,
+ "eval_mean_token_accuracy": 0.8143598649364251,
+ "eval_num_tokens": 2793261.0,
+ "eval_runtime": 95.3076,
+ "eval_samples_per_second": 17.386,
+ "eval_steps_per_second": 2.182,
+ "step": 1155
+ },
+ {
+ "entropy": 0.5205143204885512,
+ "epoch": 3.116959064327485,
+ "grad_norm": 0.663673996925354,
+ "learning_rate": 0.00029078429044885647,
+ "loss": 0.42832801818847654,
+ "mean_token_accuracy": 0.8651104230976584,
+ "num_tokens": 2896983.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.5157249197363853,
+ "epoch": 3.246913580246914,
+ "grad_norm": 0.8357914686203003,
+ "learning_rate": 0.00028556108842360404,
+ "loss": 0.4264420700073242,
+ "mean_token_accuracy": 0.8651721149682998,
+ "num_tokens": 3019005.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.5092187704145908,
+ "epoch": 3.3768680961663415,
+ "grad_norm": 0.7227477431297302,
+ "learning_rate": 0.00028009455399339797,
+ "loss": 0.4299094772338867,
+ "mean_token_accuracy": 0.8649037969112396,
+ "num_tokens": 3143415.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5237507157027721,
+ "epoch": 3.50682261208577,
+ "grad_norm": 0.6605198383331299,
+ "learning_rate": 0.00027439591951750953,
+ "loss": 0.44015567779541015,
+ "mean_token_accuracy": 0.8605978584289551,
+ "num_tokens": 3263845.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.5314217877388,
+ "epoch": 3.636777128005198,
+ "grad_norm": 0.9264621734619141,
+ "learning_rate": 0.00026847689426267907,
+ "loss": 0.4449501037597656,
+ "mean_token_accuracy": 0.862554576098919,
+ "num_tokens": 3383117.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.5156177791953087,
+ "epoch": 3.7667316439246266,
+ "grad_norm": 0.674964964389801,
+ "learning_rate": 0.0002623496403435053,
+ "loss": 0.4325400161743164,
+ "mean_token_accuracy": 0.8629470297694206,
+ "num_tokens": 3509724.0,
+ "step": 1450
+ },
+ {
+ "entropy": 0.5290651166439057,
+ "epoch": 3.8966861598440543,
+ "grad_norm": 0.8422386050224304,
+ "learning_rate": 0.00025602674773234535,
+ "loss": 0.44204349517822267,
+ "mean_token_accuracy": 0.8624594563245773,
+ "num_tokens": 3631110.0,
+ "step": 1500
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5855578513672719,
+ "eval_loss": 0.6782004833221436,
+ "eval_mean_token_accuracy": 0.8134025845390099,
+ "eval_num_tokens": 3724348.0,
+ "eval_runtime": 95.7097,
+ "eval_samples_per_second": 17.313,
+ "eval_steps_per_second": 2.173,
+ "step": 1540
+ },
+ {
+ "entropy": 0.513688079526077,
+ "epoch": 4.025990903183885,
+ "grad_norm": 0.71764075756073,
+ "learning_rate": 0.0002495212083900749,
+ "loss": 0.4214688491821289,
+ "mean_token_accuracy": 0.8665371561170223,
+ "num_tokens": 3749929.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.4366507241129875,
+ "epoch": 4.155945419103314,
+ "grad_norm": 0.8995222449302673,
+ "learning_rate": 0.00024284638957086167,
+ "loss": 0.3381219482421875,
+ "mean_token_accuracy": 0.8888532662391663,
+ "num_tokens": 3871403.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.4323717120289803,
+ "epoch": 4.2858999350227425,
+ "grad_norm": 0.7425215840339661,
+ "learning_rate": 0.00023601600635580594,
+ "loss": 0.3422117233276367,
+ "mean_token_accuracy": 0.8885110357403755,
+ "num_tokens": 3990410.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.44047351568937304,
+ "epoch": 4.41585445094217,
+ "grad_norm": 0.7615042328834534,
+ "learning_rate": 0.0002290440934718835,
+ "loss": 0.352070198059082,
+ "mean_token_accuracy": 0.8850394701957702,
+ "num_tokens": 4107377.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.4474208961427212,
+ "epoch": 4.545808966861598,
+ "grad_norm": 0.7616956830024719,
+ "learning_rate": 0.00022194497645409647,
+ "loss": 0.35494945526123045,
+ "mean_token_accuracy": 0.8840466964244843,
+ "num_tokens": 4225238.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.4457523064315319,
+ "epoch": 4.675763482781027,
+ "grad_norm": 0.7743670344352722,
+ "learning_rate": 0.00021473324221008624,
+ "loss": 0.35500614166259764,
+ "mean_token_accuracy": 0.8841654139757157,
+ "num_tokens": 4348229.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4438612677156925,
+ "epoch": 4.805717998700455,
+ "grad_norm": 1.1020246744155884,
+ "learning_rate": 0.0002074237090476915,
+ "loss": 0.3541957092285156,
+ "mean_token_accuracy": 0.8849165737628937,
+ "num_tokens": 4471337.0,
+ "step": 1850
+ },
+ {
+ "entropy": 0.44938020199537276,
+ "epoch": 4.935672514619883,
+ "grad_norm": 0.8200531601905823,
+ "learning_rate": 0.00020003139622703643,
+ "loss": 0.35912376403808594,
+ "mean_token_accuracy": 0.8818741375207901,
+ "num_tokens": 4596516.0,
+ "step": 1900
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.5258857041883928,
+ "eval_loss": 0.7124063372612,
+ "eval_mean_token_accuracy": 0.8141827643490754,
+ "eval_num_tokens": 4655435.0,
+ "eval_runtime": 95.068,
+ "eval_samples_per_second": 17.43,
+ "eval_steps_per_second": 2.188,
+ "step": 1925
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3850,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.9560754550221824e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2310/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2310/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2310/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2310/adapter_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2310/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..13015f013b98b31173d1f215c91d6b5ff9f679cc
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2310/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 32,
+ "lora_bias": false,
+ "lora_dropout": 0.08930249338906042,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 16,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "k_proj",
+ "down_proj",
+ "q_proj",
+ "v_proj",
+ "up_proj",
+ "gate_proj",
+ "o_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2310/chat_template.jinja b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2310/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2310/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2310/tokenizer_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2310/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2310/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2310/trainer_state.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2310/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..9d144c537433cb7b4f0706f8cc399e915813661a
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2310/trainer_state.json
@@ -0,0 +1,560 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 6.0,
+ "eval_steps": 500,
+ "global_step": 2310,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.8437769564986228,
+ "epoch": 0.1299545159194282,
+ "grad_norm": 1.6047701835632324,
+ "learning_rate": 4.2543894243185634e-05,
+ "loss": 1.6604945373535156,
+ "mean_token_accuracy": 0.6506646998226643,
+ "num_tokens": 127115.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.9294079938530921,
+ "epoch": 0.2599090318388564,
+ "grad_norm": 1.1633784770965576,
+ "learning_rate": 8.595603122602811e-05,
+ "loss": 0.8363616180419922,
+ "mean_token_accuracy": 0.7776120388507843,
+ "num_tokens": 251796.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.8483590838313103,
+ "epoch": 0.3898635477582846,
+ "grad_norm": 0.8963253498077393,
+ "learning_rate": 0.00012936816820887062,
+ "loss": 0.7476110076904297,
+ "mean_token_accuracy": 0.7931549972295762,
+ "num_tokens": 366505.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.7840312227606774,
+ "epoch": 0.5198180636777128,
+ "grad_norm": 0.7876676917076111,
+ "learning_rate": 0.00017278030519171308,
+ "loss": 0.6980364990234375,
+ "mean_token_accuracy": 0.8027918112277984,
+ "num_tokens": 493150.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.7967149233818054,
+ "epoch": 0.649772579597141,
+ "grad_norm": 0.7976633906364441,
+ "learning_rate": 0.00021619244217455558,
+ "loss": 0.6958013153076172,
+ "mean_token_accuracy": 0.8039823538064956,
+ "num_tokens": 607197.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.7755034640431404,
+ "epoch": 0.7797270955165692,
+ "grad_norm": 0.7086226940155029,
+ "learning_rate": 0.00025960457915739807,
+ "loss": 0.6792864990234375,
+ "mean_token_accuracy": 0.8074422472715378,
+ "num_tokens": 726579.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.7452501457929611,
+ "epoch": 0.9096816114359974,
+ "grad_norm": 0.6065998673439026,
+ "learning_rate": 0.00030301671614024056,
+ "loss": 0.6627902984619141,
+ "mean_token_accuracy": 0.8118429997563362,
+ "num_tokens": 848944.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.7490686719807295,
+ "eval_loss": 0.718986451625824,
+ "eval_mean_token_accuracy": 0.7978653156986604,
+ "eval_num_tokens": 931087.0,
+ "eval_runtime": 93.2605,
+ "eval_samples_per_second": 17.767,
+ "eval_steps_per_second": 2.23,
+ "step": 385
+ },
+ {
+ "entropy": 0.7291471499893534,
+ "epoch": 1.0389863547758285,
+ "grad_norm": 0.7773680090904236,
+ "learning_rate": 0.0003342599904174574,
+ "loss": 0.6407540893554687,
+ "mean_token_accuracy": 0.816381629387937,
+ "num_tokens": 968580.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.6813811221718789,
+ "epoch": 1.1689408706952567,
+ "grad_norm": 0.7393642067909241,
+ "learning_rate": 0.0003339921524880377,
+ "loss": 0.6079624938964844,
+ "mean_token_accuracy": 0.825104131102562,
+ "num_tokens": 1090002.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.682051141411066,
+ "epoch": 1.2988953866146848,
+ "grad_norm": 0.6833657026290894,
+ "learning_rate": 0.0003333814684455912,
+ "loss": 0.6068630218505859,
+ "mean_token_accuracy": 0.8236467111110687,
+ "num_tokens": 1212161.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.6863492175936698,
+ "epoch": 1.428849902534113,
+ "grad_norm": 0.6060423851013184,
+ "learning_rate": 0.0003324291930928916,
+ "loss": 0.6073496627807617,
+ "mean_token_accuracy": 0.8244831630587578,
+ "num_tokens": 1332852.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.6790362980961799,
+ "epoch": 1.5588044184535412,
+ "grad_norm": 0.7274985313415527,
+ "learning_rate": 0.00033113728311731064,
+ "loss": 0.5929213714599609,
+ "mean_token_accuracy": 0.8260384351015091,
+ "num_tokens": 1453569.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.6581329819560051,
+ "epoch": 1.6887589343729694,
+ "grad_norm": 0.6212530136108398,
+ "learning_rate": 0.00032950839307031545,
+ "loss": 0.5812822341918945,
+ "mean_token_accuracy": 0.8288319337368012,
+ "num_tokens": 1577152.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.6893622297048568,
+ "epoch": 1.8187134502923976,
+ "grad_norm": 0.5525115132331848,
+ "learning_rate": 0.0003275458699130301,
+ "loss": 0.6035601425170899,
+ "mean_token_accuracy": 0.8239563027024269,
+ "num_tokens": 1691926.0,
+ "step": 700
+ },
+ {
+ "entropy": 0.6713312044739723,
+ "epoch": 1.9486679662118258,
+ "grad_norm": 0.746279776096344,
+ "learning_rate": 0.0003252537461390686,
+ "loss": 0.59674560546875,
+ "mean_token_accuracy": 0.8267503470182419,
+ "num_tokens": 1816059.0,
+ "step": 750
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6752213362890941,
+ "eval_loss": 0.6740940809249878,
+ "eval_mean_token_accuracy": 0.8084349325643136,
+ "eval_num_tokens": 1862174.0,
+ "eval_runtime": 94.8197,
+ "eval_samples_per_second": 17.475,
+ "eval_steps_per_second": 2.194,
+ "step": 770
+ },
+ {
+ "entropy": 0.6279655678487902,
+ "epoch": 2.077972709551657,
+ "grad_norm": 0.7196288108825684,
+ "learning_rate": 0.00032263673148877055,
+ "loss": 0.5418233871459961,
+ "mean_token_accuracy": 0.8383135430177852,
+ "num_tokens": 1931345.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.5900149047374725,
+ "epoch": 2.207927225471085,
+ "grad_norm": 0.6000070571899414,
+ "learning_rate": 0.00031970020327186393,
+ "loss": 0.5065702056884765,
+ "mean_token_accuracy": 0.8456701630353928,
+ "num_tokens": 2057087.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.5950760287046433,
+ "epoch": 2.3378817413905133,
+ "grad_norm": 0.5828987956047058,
+ "learning_rate": 0.0003164501953184404,
+ "loss": 0.5115758514404297,
+ "mean_token_accuracy": 0.8449143546819687,
+ "num_tokens": 2179801.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.5921229894459248,
+ "epoch": 2.4678362573099415,
+ "grad_norm": 0.6523755788803101,
+ "learning_rate": 0.00031289338558094566,
+ "loss": 0.5099533462524414,
+ "mean_token_accuracy": 0.8458523535728455,
+ "num_tokens": 2300775.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.5937536461651325,
+ "epoch": 2.5977907732293697,
+ "grad_norm": 0.6264522075653076,
+ "learning_rate": 0.0003090370824126593,
+ "loss": 0.5148015594482422,
+ "mean_token_accuracy": 0.8433699604868888,
+ "num_tokens": 2425909.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.594799654185772,
+ "epoch": 2.727745289148798,
+ "grad_norm": 0.867798388004303,
+ "learning_rate": 0.000304889209550859,
+ "loss": 0.5112411499023437,
+ "mean_token_accuracy": 0.844317550957203,
+ "num_tokens": 2542506.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.5839655491709709,
+ "epoch": 2.857699805068226,
+ "grad_norm": 0.5512472987174988,
+ "learning_rate": 0.0003004582898355252,
+ "loss": 0.5124132537841797,
+ "mean_token_accuracy": 0.8445275443792343,
+ "num_tokens": 2666261.0,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6069950518012047,
+ "epoch": 2.9876543209876543,
+ "grad_norm": 0.6030759215354919,
+ "learning_rate": 0.00029575342769703865,
+ "loss": 0.521760368347168,
+ "mean_token_accuracy": 0.8433170795440674,
+ "num_tokens": 2781874.0,
+ "step": 1150
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.6072480179942571,
+ "eval_loss": 0.666907787322998,
+ "eval_mean_token_accuracy": 0.8143598649364251,
+ "eval_num_tokens": 2793261.0,
+ "eval_runtime": 95.3076,
+ "eval_samples_per_second": 17.386,
+ "eval_steps_per_second": 2.182,
+ "step": 1155
+ },
+ {
+ "entropy": 0.5205143204885512,
+ "epoch": 3.116959064327485,
+ "grad_norm": 0.663673996925354,
+ "learning_rate": 0.00029078429044885647,
+ "loss": 0.42832801818847654,
+ "mean_token_accuracy": 0.8651104230976584,
+ "num_tokens": 2896983.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.5157249197363853,
+ "epoch": 3.246913580246914,
+ "grad_norm": 0.8357914686203003,
+ "learning_rate": 0.00028556108842360404,
+ "loss": 0.4264420700073242,
+ "mean_token_accuracy": 0.8651721149682998,
+ "num_tokens": 3019005.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.5092187704145908,
+ "epoch": 3.3768680961663415,
+ "grad_norm": 0.7227477431297302,
+ "learning_rate": 0.00028009455399339797,
+ "loss": 0.4299094772338867,
+ "mean_token_accuracy": 0.8649037969112396,
+ "num_tokens": 3143415.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5237507157027721,
+ "epoch": 3.50682261208577,
+ "grad_norm": 0.6605198383331299,
+ "learning_rate": 0.00027439591951750953,
+ "loss": 0.44015567779541015,
+ "mean_token_accuracy": 0.8605978584289551,
+ "num_tokens": 3263845.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.5314217877388,
+ "epoch": 3.636777128005198,
+ "grad_norm": 0.9264621734619141,
+ "learning_rate": 0.00026847689426267907,
+ "loss": 0.4449501037597656,
+ "mean_token_accuracy": 0.862554576098919,
+ "num_tokens": 3383117.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.5156177791953087,
+ "epoch": 3.7667316439246266,
+ "grad_norm": 0.674964964389801,
+ "learning_rate": 0.0002623496403435053,
+ "loss": 0.4325400161743164,
+ "mean_token_accuracy": 0.8629470297694206,
+ "num_tokens": 3509724.0,
+ "step": 1450
+ },
+ {
+ "entropy": 0.5290651166439057,
+ "epoch": 3.8966861598440543,
+ "grad_norm": 0.8422386050224304,
+ "learning_rate": 0.00025602674773234535,
+ "loss": 0.44204349517822267,
+ "mean_token_accuracy": 0.8624594563245773,
+ "num_tokens": 3631110.0,
+ "step": 1500
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5855578513672719,
+ "eval_loss": 0.6782004833221436,
+ "eval_mean_token_accuracy": 0.8134025845390099,
+ "eval_num_tokens": 3724348.0,
+ "eval_runtime": 95.7097,
+ "eval_samples_per_second": 17.313,
+ "eval_steps_per_second": 2.173,
+ "step": 1540
+ },
+ {
+ "entropy": 0.513688079526077,
+ "epoch": 4.025990903183885,
+ "grad_norm": 0.71764075756073,
+ "learning_rate": 0.0002495212083900749,
+ "loss": 0.4214688491821289,
+ "mean_token_accuracy": 0.8665371561170223,
+ "num_tokens": 3749929.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.4366507241129875,
+ "epoch": 4.155945419103314,
+ "grad_norm": 0.8995222449302673,
+ "learning_rate": 0.00024284638957086167,
+ "loss": 0.3381219482421875,
+ "mean_token_accuracy": 0.8888532662391663,
+ "num_tokens": 3871403.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.4323717120289803,
+ "epoch": 4.2858999350227425,
+ "grad_norm": 0.7425215840339661,
+ "learning_rate": 0.00023601600635580594,
+ "loss": 0.3422117233276367,
+ "mean_token_accuracy": 0.8885110357403755,
+ "num_tokens": 3990410.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.44047351568937304,
+ "epoch": 4.41585445094217,
+ "grad_norm": 0.7615042328834534,
+ "learning_rate": 0.0002290440934718835,
+ "loss": 0.352070198059082,
+ "mean_token_accuracy": 0.8850394701957702,
+ "num_tokens": 4107377.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.4474208961427212,
+ "epoch": 4.545808966861598,
+ "grad_norm": 0.7616956830024719,
+ "learning_rate": 0.00022194497645409647,
+ "loss": 0.35494945526123045,
+ "mean_token_accuracy": 0.8840466964244843,
+ "num_tokens": 4225238.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.4457523064315319,
+ "epoch": 4.675763482781027,
+ "grad_norm": 0.7743670344352722,
+ "learning_rate": 0.00021473324221008624,
+ "loss": 0.35500614166259764,
+ "mean_token_accuracy": 0.8841654139757157,
+ "num_tokens": 4348229.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4438612677156925,
+ "epoch": 4.805717998700455,
+ "grad_norm": 1.1020246744155884,
+ "learning_rate": 0.0002074237090476915,
+ "loss": 0.3541957092285156,
+ "mean_token_accuracy": 0.8849165737628937,
+ "num_tokens": 4471337.0,
+ "step": 1850
+ },
+ {
+ "entropy": 0.44938020199537276,
+ "epoch": 4.935672514619883,
+ "grad_norm": 0.8200531601905823,
+ "learning_rate": 0.00020003139622703643,
+ "loss": 0.35912376403808594,
+ "mean_token_accuracy": 0.8818741375207901,
+ "num_tokens": 4596516.0,
+ "step": 1900
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.5258857041883928,
+ "eval_loss": 0.7124063372612,
+ "eval_mean_token_accuracy": 0.8141827643490754,
+ "eval_num_tokens": 4655435.0,
+ "eval_runtime": 95.068,
+ "eval_samples_per_second": 17.43,
+ "eval_steps_per_second": 2.188,
+ "step": 1925
+ },
+ {
+ "entropy": 0.3827864169774942,
+ "epoch": 5.064977257959714,
+ "grad_norm": 0.7284737825393677,
+ "learning_rate": 0.00019257149309971248,
+ "loss": 0.2970101737976074,
+ "mean_token_accuracy": 0.9038334503844755,
+ "num_tokens": 4720113.0,
+ "step": 1950
+ },
+ {
+ "entropy": 0.3430165499448776,
+ "epoch": 5.1949317738791425,
+ "grad_norm": 0.8974202275276184,
+ "learning_rate": 0.00018505932789846454,
+ "loss": 0.2502327537536621,
+ "mean_token_accuracy": 0.9151435106992721,
+ "num_tokens": 4842234.0,
+ "step": 2000
+ },
+ {
+ "entropy": 0.3617735780030489,
+ "epoch": 5.32488628979857,
+ "grad_norm": 0.7926186323165894,
+ "learning_rate": 0.00017751033624151158,
+ "loss": 0.25911506652832034,
+ "mean_token_accuracy": 0.9122897359728813,
+ "num_tokens": 4959991.0,
+ "step": 2050
+ },
+ {
+ "entropy": 0.3655681725591421,
+ "epoch": 5.454840805717999,
+ "grad_norm": 0.9221174716949463,
+ "learning_rate": 0.00016994002941621663,
+ "loss": 0.2651521873474121,
+ "mean_token_accuracy": 0.9111330035328865,
+ "num_tokens": 5082208.0,
+ "step": 2100
+ },
+ {
+ "entropy": 0.3543636628985405,
+ "epoch": 5.584795321637427,
+ "grad_norm": 0.7796305418014526,
+ "learning_rate": 0.00016236396250727552,
+ "loss": 0.2593145751953125,
+ "mean_token_accuracy": 0.912629722058773,
+ "num_tokens": 5199699.0,
+ "step": 2150
+ },
+ {
+ "entropy": 0.3563989091664553,
+ "epoch": 5.714749837556855,
+ "grad_norm": 0.8301284313201904,
+ "learning_rate": 0.00015479770243491373,
+ "loss": 0.26270965576171873,
+ "mean_token_accuracy": 0.9109204006195069,
+ "num_tokens": 5320501.0,
+ "step": 2200
+ },
+ {
+ "entropy": 0.3632403892278671,
+ "epoch": 5.844704353476283,
+ "grad_norm": 1.0204681158065796,
+ "learning_rate": 0.00014725679596876322,
+ "loss": 0.2649832725524902,
+ "mean_token_accuracy": 0.9091781166195869,
+ "num_tokens": 5442291.0,
+ "step": 2250
+ },
+ {
+ "entropy": 0.3748371870070696,
+ "epoch": 5.974658869395712,
+ "grad_norm": 0.9021655917167664,
+ "learning_rate": 0.00013975673778314469,
+ "loss": 0.2670526313781738,
+ "mean_token_accuracy": 0.9082769411802292,
+ "num_tokens": 5560531.0,
+ "step": 2300
+ },
+ {
+ "epoch": 6.0,
+ "eval_entropy": 0.4775148034095764,
+ "eval_loss": 0.7729372978210449,
+ "eval_mean_token_accuracy": 0.8107542619109154,
+ "eval_num_tokens": 5586522.0,
+ "eval_runtime": 95.0564,
+ "eval_samples_per_second": 17.432,
+ "eval_steps_per_second": 2.188,
+ "step": 2310
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3850,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.3465872717188096e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2695/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2695/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2695/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2695/adapter_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2695/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..13015f013b98b31173d1f215c91d6b5ff9f679cc
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2695/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 32,
+ "lora_bias": false,
+ "lora_dropout": 0.08930249338906042,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 16,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "k_proj",
+ "down_proj",
+ "q_proj",
+ "v_proj",
+ "up_proj",
+ "gate_proj",
+ "o_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2695/chat_template.jinja b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2695/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2695/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2695/tokenizer_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2695/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2695/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2695/trainer_state.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2695/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..978f8398dd30b13d74781e43b50c3265b1a786b1
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2695/trainer_state.json
@@ -0,0 +1,641 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 7.0,
+ "eval_steps": 500,
+ "global_step": 2695,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.8437769564986228,
+ "epoch": 0.1299545159194282,
+ "grad_norm": 1.6047701835632324,
+ "learning_rate": 4.2543894243185634e-05,
+ "loss": 1.6604945373535156,
+ "mean_token_accuracy": 0.6506646998226643,
+ "num_tokens": 127115.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.9294079938530921,
+ "epoch": 0.2599090318388564,
+ "grad_norm": 1.1633784770965576,
+ "learning_rate": 8.595603122602811e-05,
+ "loss": 0.8363616180419922,
+ "mean_token_accuracy": 0.7776120388507843,
+ "num_tokens": 251796.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.8483590838313103,
+ "epoch": 0.3898635477582846,
+ "grad_norm": 0.8963253498077393,
+ "learning_rate": 0.00012936816820887062,
+ "loss": 0.7476110076904297,
+ "mean_token_accuracy": 0.7931549972295762,
+ "num_tokens": 366505.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.7840312227606774,
+ "epoch": 0.5198180636777128,
+ "grad_norm": 0.7876676917076111,
+ "learning_rate": 0.00017278030519171308,
+ "loss": 0.6980364990234375,
+ "mean_token_accuracy": 0.8027918112277984,
+ "num_tokens": 493150.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.7967149233818054,
+ "epoch": 0.649772579597141,
+ "grad_norm": 0.7976633906364441,
+ "learning_rate": 0.00021619244217455558,
+ "loss": 0.6958013153076172,
+ "mean_token_accuracy": 0.8039823538064956,
+ "num_tokens": 607197.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.7755034640431404,
+ "epoch": 0.7797270955165692,
+ "grad_norm": 0.7086226940155029,
+ "learning_rate": 0.00025960457915739807,
+ "loss": 0.6792864990234375,
+ "mean_token_accuracy": 0.8074422472715378,
+ "num_tokens": 726579.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.7452501457929611,
+ "epoch": 0.9096816114359974,
+ "grad_norm": 0.6065998673439026,
+ "learning_rate": 0.00030301671614024056,
+ "loss": 0.6627902984619141,
+ "mean_token_accuracy": 0.8118429997563362,
+ "num_tokens": 848944.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.7490686719807295,
+ "eval_loss": 0.718986451625824,
+ "eval_mean_token_accuracy": 0.7978653156986604,
+ "eval_num_tokens": 931087.0,
+ "eval_runtime": 93.2605,
+ "eval_samples_per_second": 17.767,
+ "eval_steps_per_second": 2.23,
+ "step": 385
+ },
+ {
+ "entropy": 0.7291471499893534,
+ "epoch": 1.0389863547758285,
+ "grad_norm": 0.7773680090904236,
+ "learning_rate": 0.0003342599904174574,
+ "loss": 0.6407540893554687,
+ "mean_token_accuracy": 0.816381629387937,
+ "num_tokens": 968580.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.6813811221718789,
+ "epoch": 1.1689408706952567,
+ "grad_norm": 0.7393642067909241,
+ "learning_rate": 0.0003339921524880377,
+ "loss": 0.6079624938964844,
+ "mean_token_accuracy": 0.825104131102562,
+ "num_tokens": 1090002.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.682051141411066,
+ "epoch": 1.2988953866146848,
+ "grad_norm": 0.6833657026290894,
+ "learning_rate": 0.0003333814684455912,
+ "loss": 0.6068630218505859,
+ "mean_token_accuracy": 0.8236467111110687,
+ "num_tokens": 1212161.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.6863492175936698,
+ "epoch": 1.428849902534113,
+ "grad_norm": 0.6060423851013184,
+ "learning_rate": 0.0003324291930928916,
+ "loss": 0.6073496627807617,
+ "mean_token_accuracy": 0.8244831630587578,
+ "num_tokens": 1332852.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.6790362980961799,
+ "epoch": 1.5588044184535412,
+ "grad_norm": 0.7274985313415527,
+ "learning_rate": 0.00033113728311731064,
+ "loss": 0.5929213714599609,
+ "mean_token_accuracy": 0.8260384351015091,
+ "num_tokens": 1453569.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.6581329819560051,
+ "epoch": 1.6887589343729694,
+ "grad_norm": 0.6212530136108398,
+ "learning_rate": 0.00032950839307031545,
+ "loss": 0.5812822341918945,
+ "mean_token_accuracy": 0.8288319337368012,
+ "num_tokens": 1577152.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.6893622297048568,
+ "epoch": 1.8187134502923976,
+ "grad_norm": 0.5525115132331848,
+ "learning_rate": 0.0003275458699130301,
+ "loss": 0.6035601425170899,
+ "mean_token_accuracy": 0.8239563027024269,
+ "num_tokens": 1691926.0,
+ "step": 700
+ },
+ {
+ "entropy": 0.6713312044739723,
+ "epoch": 1.9486679662118258,
+ "grad_norm": 0.746279776096344,
+ "learning_rate": 0.0003252537461390686,
+ "loss": 0.59674560546875,
+ "mean_token_accuracy": 0.8267503470182419,
+ "num_tokens": 1816059.0,
+ "step": 750
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6752213362890941,
+ "eval_loss": 0.6740940809249878,
+ "eval_mean_token_accuracy": 0.8084349325643136,
+ "eval_num_tokens": 1862174.0,
+ "eval_runtime": 94.8197,
+ "eval_samples_per_second": 17.475,
+ "eval_steps_per_second": 2.194,
+ "step": 770
+ },
+ {
+ "entropy": 0.6279655678487902,
+ "epoch": 2.077972709551657,
+ "grad_norm": 0.7196288108825684,
+ "learning_rate": 0.00032263673148877055,
+ "loss": 0.5418233871459961,
+ "mean_token_accuracy": 0.8383135430177852,
+ "num_tokens": 1931345.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.5900149047374725,
+ "epoch": 2.207927225471085,
+ "grad_norm": 0.6000070571899414,
+ "learning_rate": 0.00031970020327186393,
+ "loss": 0.5065702056884765,
+ "mean_token_accuracy": 0.8456701630353928,
+ "num_tokens": 2057087.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.5950760287046433,
+ "epoch": 2.3378817413905133,
+ "grad_norm": 0.5828987956047058,
+ "learning_rate": 0.0003164501953184404,
+ "loss": 0.5115758514404297,
+ "mean_token_accuracy": 0.8449143546819687,
+ "num_tokens": 2179801.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.5921229894459248,
+ "epoch": 2.4678362573099415,
+ "grad_norm": 0.6523755788803101,
+ "learning_rate": 0.00031289338558094566,
+ "loss": 0.5099533462524414,
+ "mean_token_accuracy": 0.8458523535728455,
+ "num_tokens": 2300775.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.5937536461651325,
+ "epoch": 2.5977907732293697,
+ "grad_norm": 0.6264522075653076,
+ "learning_rate": 0.0003090370824126593,
+ "loss": 0.5148015594482422,
+ "mean_token_accuracy": 0.8433699604868888,
+ "num_tokens": 2425909.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.594799654185772,
+ "epoch": 2.727745289148798,
+ "grad_norm": 0.867798388004303,
+ "learning_rate": 0.000304889209550859,
+ "loss": 0.5112411499023437,
+ "mean_token_accuracy": 0.844317550957203,
+ "num_tokens": 2542506.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.5839655491709709,
+ "epoch": 2.857699805068226,
+ "grad_norm": 0.5512472987174988,
+ "learning_rate": 0.0003004582898355252,
+ "loss": 0.5124132537841797,
+ "mean_token_accuracy": 0.8445275443792343,
+ "num_tokens": 2666261.0,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6069950518012047,
+ "epoch": 2.9876543209876543,
+ "grad_norm": 0.6030759215354919,
+ "learning_rate": 0.00029575342769703865,
+ "loss": 0.521760368347168,
+ "mean_token_accuracy": 0.8433170795440674,
+ "num_tokens": 2781874.0,
+ "step": 1150
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.6072480179942571,
+ "eval_loss": 0.666907787322998,
+ "eval_mean_token_accuracy": 0.8143598649364251,
+ "eval_num_tokens": 2793261.0,
+ "eval_runtime": 95.3076,
+ "eval_samples_per_second": 17.386,
+ "eval_steps_per_second": 2.182,
+ "step": 1155
+ },
+ {
+ "entropy": 0.5205143204885512,
+ "epoch": 3.116959064327485,
+ "grad_norm": 0.663673996925354,
+ "learning_rate": 0.00029078429044885647,
+ "loss": 0.42832801818847654,
+ "mean_token_accuracy": 0.8651104230976584,
+ "num_tokens": 2896983.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.5157249197363853,
+ "epoch": 3.246913580246914,
+ "grad_norm": 0.8357914686203003,
+ "learning_rate": 0.00028556108842360404,
+ "loss": 0.4264420700073242,
+ "mean_token_accuracy": 0.8651721149682998,
+ "num_tokens": 3019005.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.5092187704145908,
+ "epoch": 3.3768680961663415,
+ "grad_norm": 0.7227477431297302,
+ "learning_rate": 0.00028009455399339797,
+ "loss": 0.4299094772338867,
+ "mean_token_accuracy": 0.8649037969112396,
+ "num_tokens": 3143415.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5237507157027721,
+ "epoch": 3.50682261208577,
+ "grad_norm": 0.6605198383331299,
+ "learning_rate": 0.00027439591951750953,
+ "loss": 0.44015567779541015,
+ "mean_token_accuracy": 0.8605978584289551,
+ "num_tokens": 3263845.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.5314217877388,
+ "epoch": 3.636777128005198,
+ "grad_norm": 0.9264621734619141,
+ "learning_rate": 0.00026847689426267907,
+ "loss": 0.4449501037597656,
+ "mean_token_accuracy": 0.862554576098919,
+ "num_tokens": 3383117.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.5156177791953087,
+ "epoch": 3.7667316439246266,
+ "grad_norm": 0.674964964389801,
+ "learning_rate": 0.0002623496403435053,
+ "loss": 0.4325400161743164,
+ "mean_token_accuracy": 0.8629470297694206,
+ "num_tokens": 3509724.0,
+ "step": 1450
+ },
+ {
+ "entropy": 0.5290651166439057,
+ "epoch": 3.8966861598440543,
+ "grad_norm": 0.8422386050224304,
+ "learning_rate": 0.00025602674773234535,
+ "loss": 0.44204349517822267,
+ "mean_token_accuracy": 0.8624594563245773,
+ "num_tokens": 3631110.0,
+ "step": 1500
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5855578513672719,
+ "eval_loss": 0.6782004833221436,
+ "eval_mean_token_accuracy": 0.8134025845390099,
+ "eval_num_tokens": 3724348.0,
+ "eval_runtime": 95.7097,
+ "eval_samples_per_second": 17.313,
+ "eval_steps_per_second": 2.173,
+ "step": 1540
+ },
+ {
+ "entropy": 0.513688079526077,
+ "epoch": 4.025990903183885,
+ "grad_norm": 0.71764075756073,
+ "learning_rate": 0.0002495212083900749,
+ "loss": 0.4214688491821289,
+ "mean_token_accuracy": 0.8665371561170223,
+ "num_tokens": 3749929.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.4366507241129875,
+ "epoch": 4.155945419103314,
+ "grad_norm": 0.8995222449302673,
+ "learning_rate": 0.00024284638957086167,
+ "loss": 0.3381219482421875,
+ "mean_token_accuracy": 0.8888532662391663,
+ "num_tokens": 3871403.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.4323717120289803,
+ "epoch": 4.2858999350227425,
+ "grad_norm": 0.7425215840339661,
+ "learning_rate": 0.00023601600635580594,
+ "loss": 0.3422117233276367,
+ "mean_token_accuracy": 0.8885110357403755,
+ "num_tokens": 3990410.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.44047351568937304,
+ "epoch": 4.41585445094217,
+ "grad_norm": 0.7615042328834534,
+ "learning_rate": 0.0002290440934718835,
+ "loss": 0.352070198059082,
+ "mean_token_accuracy": 0.8850394701957702,
+ "num_tokens": 4107377.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.4474208961427212,
+ "epoch": 4.545808966861598,
+ "grad_norm": 0.7616956830024719,
+ "learning_rate": 0.00022194497645409647,
+ "loss": 0.35494945526123045,
+ "mean_token_accuracy": 0.8840466964244843,
+ "num_tokens": 4225238.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.4457523064315319,
+ "epoch": 4.675763482781027,
+ "grad_norm": 0.7743670344352722,
+ "learning_rate": 0.00021473324221008624,
+ "loss": 0.35500614166259764,
+ "mean_token_accuracy": 0.8841654139757157,
+ "num_tokens": 4348229.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4438612677156925,
+ "epoch": 4.805717998700455,
+ "grad_norm": 1.1020246744155884,
+ "learning_rate": 0.0002074237090476915,
+ "loss": 0.3541957092285156,
+ "mean_token_accuracy": 0.8849165737628937,
+ "num_tokens": 4471337.0,
+ "step": 1850
+ },
+ {
+ "entropy": 0.44938020199537276,
+ "epoch": 4.935672514619883,
+ "grad_norm": 0.8200531601905823,
+ "learning_rate": 0.00020003139622703643,
+ "loss": 0.35912376403808594,
+ "mean_token_accuracy": 0.8818741375207901,
+ "num_tokens": 4596516.0,
+ "step": 1900
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.5258857041883928,
+ "eval_loss": 0.7124063372612,
+ "eval_mean_token_accuracy": 0.8141827643490754,
+ "eval_num_tokens": 4655435.0,
+ "eval_runtime": 95.068,
+ "eval_samples_per_second": 17.43,
+ "eval_steps_per_second": 2.188,
+ "step": 1925
+ },
+ {
+ "entropy": 0.3827864169774942,
+ "epoch": 5.064977257959714,
+ "grad_norm": 0.7284737825393677,
+ "learning_rate": 0.00019257149309971248,
+ "loss": 0.2970101737976074,
+ "mean_token_accuracy": 0.9038334503844755,
+ "num_tokens": 4720113.0,
+ "step": 1950
+ },
+ {
+ "entropy": 0.3430165499448776,
+ "epoch": 5.1949317738791425,
+ "grad_norm": 0.8974202275276184,
+ "learning_rate": 0.00018505932789846454,
+ "loss": 0.2502327537536621,
+ "mean_token_accuracy": 0.9151435106992721,
+ "num_tokens": 4842234.0,
+ "step": 2000
+ },
+ {
+ "entropy": 0.3617735780030489,
+ "epoch": 5.32488628979857,
+ "grad_norm": 0.7926186323165894,
+ "learning_rate": 0.00017751033624151158,
+ "loss": 0.25911506652832034,
+ "mean_token_accuracy": 0.9122897359728813,
+ "num_tokens": 4959991.0,
+ "step": 2050
+ },
+ {
+ "entropy": 0.3655681725591421,
+ "epoch": 5.454840805717999,
+ "grad_norm": 0.9221174716949463,
+ "learning_rate": 0.00016994002941621663,
+ "loss": 0.2651521873474121,
+ "mean_token_accuracy": 0.9111330035328865,
+ "num_tokens": 5082208.0,
+ "step": 2100
+ },
+ {
+ "entropy": 0.3543636628985405,
+ "epoch": 5.584795321637427,
+ "grad_norm": 0.7796305418014526,
+ "learning_rate": 0.00016236396250727552,
+ "loss": 0.2593145751953125,
+ "mean_token_accuracy": 0.912629722058773,
+ "num_tokens": 5199699.0,
+ "step": 2150
+ },
+ {
+ "entropy": 0.3563989091664553,
+ "epoch": 5.714749837556855,
+ "grad_norm": 0.8301284313201904,
+ "learning_rate": 0.00015479770243491373,
+ "loss": 0.26270965576171873,
+ "mean_token_accuracy": 0.9109204006195069,
+ "num_tokens": 5320501.0,
+ "step": 2200
+ },
+ {
+ "entropy": 0.3632403892278671,
+ "epoch": 5.844704353476283,
+ "grad_norm": 1.0204681158065796,
+ "learning_rate": 0.00014725679596876322,
+ "loss": 0.2649832725524902,
+ "mean_token_accuracy": 0.9091781166195869,
+ "num_tokens": 5442291.0,
+ "step": 2250
+ },
+ {
+ "entropy": 0.3748371870070696,
+ "epoch": 5.974658869395712,
+ "grad_norm": 0.9021655917167664,
+ "learning_rate": 0.00013975673778314469,
+ "loss": 0.2670526313781738,
+ "mean_token_accuracy": 0.9082769411802292,
+ "num_tokens": 5560531.0,
+ "step": 2300
+ },
+ {
+ "epoch": 6.0,
+ "eval_entropy": 0.4775148034095764,
+ "eval_loss": 0.7729372978210449,
+ "eval_mean_token_accuracy": 0.8107542619109154,
+ "eval_num_tokens": 5586522.0,
+ "eval_runtime": 95.0564,
+ "eval_samples_per_second": 17.432,
+ "eval_steps_per_second": 2.188,
+ "step": 2310
+ },
+ {
+ "entropy": 0.29565749216319326,
+ "epoch": 6.1039636127355426,
+ "grad_norm": 1.0129460096359253,
+ "learning_rate": 0.00013231293861939197,
+ "loss": 0.18774839401245116,
+ "mean_token_accuracy": 0.9347463528714587,
+ "num_tokens": 5682578.0,
+ "step": 2350
+ },
+ {
+ "entropy": 0.2680290696769953,
+ "epoch": 6.23391812865497,
+ "grad_norm": 0.7047284841537476,
+ "learning_rate": 0.00012494069362063769,
+ "loss": 0.1707102584838867,
+ "mean_token_accuracy": 0.9414009821414947,
+ "num_tokens": 5803409.0,
+ "step": 2400
+ },
+ {
+ "entropy": 0.2615845823287964,
+ "epoch": 6.363872644574399,
+ "grad_norm": 0.8583500981330872,
+ "learning_rate": 0.00011765515090412491,
+ "loss": 0.1720401382446289,
+ "mean_token_accuracy": 0.941322917342186,
+ "num_tokens": 5927439.0,
+ "step": 2450
+ },
+ {
+ "entropy": 0.26200685687363146,
+ "epoch": 6.493827160493828,
+ "grad_norm": 0.9596366286277771,
+ "learning_rate": 0.00011047128043561942,
+ "loss": 0.16701702117919923,
+ "mean_token_accuracy": 0.9415252614021301,
+ "num_tokens": 6052633.0,
+ "step": 2500
+ },
+ {
+ "entropy": 0.2842141987383366,
+ "epoch": 6.623781676413255,
+ "grad_norm": 1.1430736780166626,
+ "learning_rate": 0.00010340384326987922,
+ "loss": 0.18171667098999023,
+ "mean_token_accuracy": 0.935710808634758,
+ "num_tokens": 6167503.0,
+ "step": 2550
+ },
+ {
+ "entropy": 0.27876456022262575,
+ "epoch": 6.753736192332683,
+ "grad_norm": 1.0698399543762207,
+ "learning_rate": 9.646736122038319e-05,
+ "loss": 0.17995393753051758,
+ "mean_token_accuracy": 0.9379886907339096,
+ "num_tokens": 6284023.0,
+ "step": 2600
+ },
+ {
+ "entropy": 0.2657020591944456,
+ "epoch": 6.883690708252112,
+ "grad_norm": 0.8577425479888916,
+ "learning_rate": 8.96760870206418e-05,
+ "loss": 0.17407890319824218,
+ "mean_token_accuracy": 0.9403641560673713,
+ "num_tokens": 6408327.0,
+ "step": 2650
+ },
+ {
+ "epoch": 7.0,
+ "eval_entropy": 0.4080844047264411,
+ "eval_loss": 0.8632635474205017,
+ "eval_mean_token_accuracy": 0.8096621922002389,
+ "eval_num_tokens": 6517609.0,
+ "eval_runtime": 95.784,
+ "eval_samples_per_second": 17.299,
+ "eval_steps_per_second": 2.172,
+ "step": 2695
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3850,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.737490610388992e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3080/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3080/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3080/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3850/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3850/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3850/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3850/adapter_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3850/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..13015f013b98b31173d1f215c91d6b5ff9f679cc
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3850/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 32,
+ "lora_bias": false,
+ "lora_dropout": 0.08930249338906042,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 16,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "k_proj",
+ "down_proj",
+ "q_proj",
+ "v_proj",
+ "up_proj",
+ "gate_proj",
+ "o_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3850/chat_template.jinja b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3850/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3850/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3850/tokenizer_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3850/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3850/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3850/trainer_state.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3850/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..4f3d94bdfd7daf18129c9159aa84f7fc7625b535
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3850/trainer_state.json
@@ -0,0 +1,914 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 10.0,
+ "eval_steps": 500,
+ "global_step": 3850,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.8437769564986228,
+ "epoch": 0.1299545159194282,
+ "grad_norm": 1.6047701835632324,
+ "learning_rate": 4.2543894243185634e-05,
+ "loss": 1.6604945373535156,
+ "mean_token_accuracy": 0.6506646998226643,
+ "num_tokens": 127115.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.9294079938530921,
+ "epoch": 0.2599090318388564,
+ "grad_norm": 1.1633784770965576,
+ "learning_rate": 8.595603122602811e-05,
+ "loss": 0.8363616180419922,
+ "mean_token_accuracy": 0.7776120388507843,
+ "num_tokens": 251796.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.8483590838313103,
+ "epoch": 0.3898635477582846,
+ "grad_norm": 0.8963253498077393,
+ "learning_rate": 0.00012936816820887062,
+ "loss": 0.7476110076904297,
+ "mean_token_accuracy": 0.7931549972295762,
+ "num_tokens": 366505.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.7840312227606774,
+ "epoch": 0.5198180636777128,
+ "grad_norm": 0.7876676917076111,
+ "learning_rate": 0.00017278030519171308,
+ "loss": 0.6980364990234375,
+ "mean_token_accuracy": 0.8027918112277984,
+ "num_tokens": 493150.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.7967149233818054,
+ "epoch": 0.649772579597141,
+ "grad_norm": 0.7976633906364441,
+ "learning_rate": 0.00021619244217455558,
+ "loss": 0.6958013153076172,
+ "mean_token_accuracy": 0.8039823538064956,
+ "num_tokens": 607197.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.7755034640431404,
+ "epoch": 0.7797270955165692,
+ "grad_norm": 0.7086226940155029,
+ "learning_rate": 0.00025960457915739807,
+ "loss": 0.6792864990234375,
+ "mean_token_accuracy": 0.8074422472715378,
+ "num_tokens": 726579.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.7452501457929611,
+ "epoch": 0.9096816114359974,
+ "grad_norm": 0.6065998673439026,
+ "learning_rate": 0.00030301671614024056,
+ "loss": 0.6627902984619141,
+ "mean_token_accuracy": 0.8118429997563362,
+ "num_tokens": 848944.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.7490686719807295,
+ "eval_loss": 0.718986451625824,
+ "eval_mean_token_accuracy": 0.7978653156986604,
+ "eval_num_tokens": 931087.0,
+ "eval_runtime": 93.2605,
+ "eval_samples_per_second": 17.767,
+ "eval_steps_per_second": 2.23,
+ "step": 385
+ },
+ {
+ "entropy": 0.7291471499893534,
+ "epoch": 1.0389863547758285,
+ "grad_norm": 0.7773680090904236,
+ "learning_rate": 0.0003342599904174574,
+ "loss": 0.6407540893554687,
+ "mean_token_accuracy": 0.816381629387937,
+ "num_tokens": 968580.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.6813811221718789,
+ "epoch": 1.1689408706952567,
+ "grad_norm": 0.7393642067909241,
+ "learning_rate": 0.0003339921524880377,
+ "loss": 0.6079624938964844,
+ "mean_token_accuracy": 0.825104131102562,
+ "num_tokens": 1090002.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.682051141411066,
+ "epoch": 1.2988953866146848,
+ "grad_norm": 0.6833657026290894,
+ "learning_rate": 0.0003333814684455912,
+ "loss": 0.6068630218505859,
+ "mean_token_accuracy": 0.8236467111110687,
+ "num_tokens": 1212161.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.6863492175936698,
+ "epoch": 1.428849902534113,
+ "grad_norm": 0.6060423851013184,
+ "learning_rate": 0.0003324291930928916,
+ "loss": 0.6073496627807617,
+ "mean_token_accuracy": 0.8244831630587578,
+ "num_tokens": 1332852.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.6790362980961799,
+ "epoch": 1.5588044184535412,
+ "grad_norm": 0.7274985313415527,
+ "learning_rate": 0.00033113728311731064,
+ "loss": 0.5929213714599609,
+ "mean_token_accuracy": 0.8260384351015091,
+ "num_tokens": 1453569.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.6581329819560051,
+ "epoch": 1.6887589343729694,
+ "grad_norm": 0.6212530136108398,
+ "learning_rate": 0.00032950839307031545,
+ "loss": 0.5812822341918945,
+ "mean_token_accuracy": 0.8288319337368012,
+ "num_tokens": 1577152.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.6893622297048568,
+ "epoch": 1.8187134502923976,
+ "grad_norm": 0.5525115132331848,
+ "learning_rate": 0.0003275458699130301,
+ "loss": 0.6035601425170899,
+ "mean_token_accuracy": 0.8239563027024269,
+ "num_tokens": 1691926.0,
+ "step": 700
+ },
+ {
+ "entropy": 0.6713312044739723,
+ "epoch": 1.9486679662118258,
+ "grad_norm": 0.746279776096344,
+ "learning_rate": 0.0003252537461390686,
+ "loss": 0.59674560546875,
+ "mean_token_accuracy": 0.8267503470182419,
+ "num_tokens": 1816059.0,
+ "step": 750
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6752213362890941,
+ "eval_loss": 0.6740940809249878,
+ "eval_mean_token_accuracy": 0.8084349325643136,
+ "eval_num_tokens": 1862174.0,
+ "eval_runtime": 94.8197,
+ "eval_samples_per_second": 17.475,
+ "eval_steps_per_second": 2.194,
+ "step": 770
+ },
+ {
+ "entropy": 0.6279655678487902,
+ "epoch": 2.077972709551657,
+ "grad_norm": 0.7196288108825684,
+ "learning_rate": 0.00032263673148877055,
+ "loss": 0.5418233871459961,
+ "mean_token_accuracy": 0.8383135430177852,
+ "num_tokens": 1931345.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.5900149047374725,
+ "epoch": 2.207927225471085,
+ "grad_norm": 0.6000070571899414,
+ "learning_rate": 0.00031970020327186393,
+ "loss": 0.5065702056884765,
+ "mean_token_accuracy": 0.8456701630353928,
+ "num_tokens": 2057087.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.5950760287046433,
+ "epoch": 2.3378817413905133,
+ "grad_norm": 0.5828987956047058,
+ "learning_rate": 0.0003164501953184404,
+ "loss": 0.5115758514404297,
+ "mean_token_accuracy": 0.8449143546819687,
+ "num_tokens": 2179801.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.5921229894459248,
+ "epoch": 2.4678362573099415,
+ "grad_norm": 0.6523755788803101,
+ "learning_rate": 0.00031289338558094566,
+ "loss": 0.5099533462524414,
+ "mean_token_accuracy": 0.8458523535728455,
+ "num_tokens": 2300775.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.5937536461651325,
+ "epoch": 2.5977907732293697,
+ "grad_norm": 0.6264522075653076,
+ "learning_rate": 0.0003090370824126593,
+ "loss": 0.5148015594482422,
+ "mean_token_accuracy": 0.8433699604868888,
+ "num_tokens": 2425909.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.594799654185772,
+ "epoch": 2.727745289148798,
+ "grad_norm": 0.867798388004303,
+ "learning_rate": 0.000304889209550859,
+ "loss": 0.5112411499023437,
+ "mean_token_accuracy": 0.844317550957203,
+ "num_tokens": 2542506.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.5839655491709709,
+ "epoch": 2.857699805068226,
+ "grad_norm": 0.5512472987174988,
+ "learning_rate": 0.0003004582898355252,
+ "loss": 0.5124132537841797,
+ "mean_token_accuracy": 0.8445275443792343,
+ "num_tokens": 2666261.0,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6069950518012047,
+ "epoch": 2.9876543209876543,
+ "grad_norm": 0.6030759215354919,
+ "learning_rate": 0.00029575342769703865,
+ "loss": 0.521760368347168,
+ "mean_token_accuracy": 0.8433170795440674,
+ "num_tokens": 2781874.0,
+ "step": 1150
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.6072480179942571,
+ "eval_loss": 0.666907787322998,
+ "eval_mean_token_accuracy": 0.8143598649364251,
+ "eval_num_tokens": 2793261.0,
+ "eval_runtime": 95.3076,
+ "eval_samples_per_second": 17.386,
+ "eval_steps_per_second": 2.182,
+ "step": 1155
+ },
+ {
+ "entropy": 0.5205143204885512,
+ "epoch": 3.116959064327485,
+ "grad_norm": 0.663673996925354,
+ "learning_rate": 0.00029078429044885647,
+ "loss": 0.42832801818847654,
+ "mean_token_accuracy": 0.8651104230976584,
+ "num_tokens": 2896983.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.5157249197363853,
+ "epoch": 3.246913580246914,
+ "grad_norm": 0.8357914686203003,
+ "learning_rate": 0.00028556108842360404,
+ "loss": 0.4264420700073242,
+ "mean_token_accuracy": 0.8651721149682998,
+ "num_tokens": 3019005.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.5092187704145908,
+ "epoch": 3.3768680961663415,
+ "grad_norm": 0.7227477431297302,
+ "learning_rate": 0.00028009455399339797,
+ "loss": 0.4299094772338867,
+ "mean_token_accuracy": 0.8649037969112396,
+ "num_tokens": 3143415.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5237507157027721,
+ "epoch": 3.50682261208577,
+ "grad_norm": 0.6605198383331299,
+ "learning_rate": 0.00027439591951750953,
+ "loss": 0.44015567779541015,
+ "mean_token_accuracy": 0.8605978584289551,
+ "num_tokens": 3263845.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.5314217877388,
+ "epoch": 3.636777128005198,
+ "grad_norm": 0.9264621734619141,
+ "learning_rate": 0.00026847689426267907,
+ "loss": 0.4449501037597656,
+ "mean_token_accuracy": 0.862554576098919,
+ "num_tokens": 3383117.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.5156177791953087,
+ "epoch": 3.7667316439246266,
+ "grad_norm": 0.674964964389801,
+ "learning_rate": 0.0002623496403435053,
+ "loss": 0.4325400161743164,
+ "mean_token_accuracy": 0.8629470297694206,
+ "num_tokens": 3509724.0,
+ "step": 1450
+ },
+ {
+ "entropy": 0.5290651166439057,
+ "epoch": 3.8966861598440543,
+ "grad_norm": 0.8422386050224304,
+ "learning_rate": 0.00025602674773234535,
+ "loss": 0.44204349517822267,
+ "mean_token_accuracy": 0.8624594563245773,
+ "num_tokens": 3631110.0,
+ "step": 1500
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5855578513672719,
+ "eval_loss": 0.6782004833221436,
+ "eval_mean_token_accuracy": 0.8134025845390099,
+ "eval_num_tokens": 3724348.0,
+ "eval_runtime": 95.7097,
+ "eval_samples_per_second": 17.313,
+ "eval_steps_per_second": 2.173,
+ "step": 1540
+ },
+ {
+ "entropy": 0.513688079526077,
+ "epoch": 4.025990903183885,
+ "grad_norm": 0.71764075756073,
+ "learning_rate": 0.0002495212083900749,
+ "loss": 0.4214688491821289,
+ "mean_token_accuracy": 0.8665371561170223,
+ "num_tokens": 3749929.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.4366507241129875,
+ "epoch": 4.155945419103314,
+ "grad_norm": 0.8995222449302673,
+ "learning_rate": 0.00024284638957086167,
+ "loss": 0.3381219482421875,
+ "mean_token_accuracy": 0.8888532662391663,
+ "num_tokens": 3871403.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.4323717120289803,
+ "epoch": 4.2858999350227425,
+ "grad_norm": 0.7425215840339661,
+ "learning_rate": 0.00023601600635580594,
+ "loss": 0.3422117233276367,
+ "mean_token_accuracy": 0.8885110357403755,
+ "num_tokens": 3990410.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.44047351568937304,
+ "epoch": 4.41585445094217,
+ "grad_norm": 0.7615042328834534,
+ "learning_rate": 0.0002290440934718835,
+ "loss": 0.352070198059082,
+ "mean_token_accuracy": 0.8850394701957702,
+ "num_tokens": 4107377.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.4474208961427212,
+ "epoch": 4.545808966861598,
+ "grad_norm": 0.7616956830024719,
+ "learning_rate": 0.00022194497645409647,
+ "loss": 0.35494945526123045,
+ "mean_token_accuracy": 0.8840466964244843,
+ "num_tokens": 4225238.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.4457523064315319,
+ "epoch": 4.675763482781027,
+ "grad_norm": 0.7743670344352722,
+ "learning_rate": 0.00021473324221008624,
+ "loss": 0.35500614166259764,
+ "mean_token_accuracy": 0.8841654139757157,
+ "num_tokens": 4348229.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4438612677156925,
+ "epoch": 4.805717998700455,
+ "grad_norm": 1.1020246744155884,
+ "learning_rate": 0.0002074237090476915,
+ "loss": 0.3541957092285156,
+ "mean_token_accuracy": 0.8849165737628937,
+ "num_tokens": 4471337.0,
+ "step": 1850
+ },
+ {
+ "entropy": 0.44938020199537276,
+ "epoch": 4.935672514619883,
+ "grad_norm": 0.8200531601905823,
+ "learning_rate": 0.00020003139622703643,
+ "loss": 0.35912376403808594,
+ "mean_token_accuracy": 0.8818741375207901,
+ "num_tokens": 4596516.0,
+ "step": 1900
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.5258857041883928,
+ "eval_loss": 0.7124063372612,
+ "eval_mean_token_accuracy": 0.8141827643490754,
+ "eval_num_tokens": 4655435.0,
+ "eval_runtime": 95.068,
+ "eval_samples_per_second": 17.43,
+ "eval_steps_per_second": 2.188,
+ "step": 1925
+ },
+ {
+ "entropy": 0.3827864169774942,
+ "epoch": 5.064977257959714,
+ "grad_norm": 0.7284737825393677,
+ "learning_rate": 0.00019257149309971248,
+ "loss": 0.2970101737976074,
+ "mean_token_accuracy": 0.9038334503844755,
+ "num_tokens": 4720113.0,
+ "step": 1950
+ },
+ {
+ "entropy": 0.3430165499448776,
+ "epoch": 5.1949317738791425,
+ "grad_norm": 0.8974202275276184,
+ "learning_rate": 0.00018505932789846454,
+ "loss": 0.2502327537536621,
+ "mean_token_accuracy": 0.9151435106992721,
+ "num_tokens": 4842234.0,
+ "step": 2000
+ },
+ {
+ "entropy": 0.3617735780030489,
+ "epoch": 5.32488628979857,
+ "grad_norm": 0.7926186323165894,
+ "learning_rate": 0.00017751033624151158,
+ "loss": 0.25911506652832034,
+ "mean_token_accuracy": 0.9122897359728813,
+ "num_tokens": 4959991.0,
+ "step": 2050
+ },
+ {
+ "entropy": 0.3655681725591421,
+ "epoch": 5.454840805717999,
+ "grad_norm": 0.9221174716949463,
+ "learning_rate": 0.00016994002941621663,
+ "loss": 0.2651521873474121,
+ "mean_token_accuracy": 0.9111330035328865,
+ "num_tokens": 5082208.0,
+ "step": 2100
+ },
+ {
+ "entropy": 0.3543636628985405,
+ "epoch": 5.584795321637427,
+ "grad_norm": 0.7796305418014526,
+ "learning_rate": 0.00016236396250727552,
+ "loss": 0.2593145751953125,
+ "mean_token_accuracy": 0.912629722058773,
+ "num_tokens": 5199699.0,
+ "step": 2150
+ },
+ {
+ "entropy": 0.3563989091664553,
+ "epoch": 5.714749837556855,
+ "grad_norm": 0.8301284313201904,
+ "learning_rate": 0.00015479770243491373,
+ "loss": 0.26270965576171873,
+ "mean_token_accuracy": 0.9109204006195069,
+ "num_tokens": 5320501.0,
+ "step": 2200
+ },
+ {
+ "entropy": 0.3632403892278671,
+ "epoch": 5.844704353476283,
+ "grad_norm": 1.0204681158065796,
+ "learning_rate": 0.00014725679596876322,
+ "loss": 0.2649832725524902,
+ "mean_token_accuracy": 0.9091781166195869,
+ "num_tokens": 5442291.0,
+ "step": 2250
+ },
+ {
+ "entropy": 0.3748371870070696,
+ "epoch": 5.974658869395712,
+ "grad_norm": 0.9021655917167664,
+ "learning_rate": 0.00013975673778314469,
+ "loss": 0.2670526313781738,
+ "mean_token_accuracy": 0.9082769411802292,
+ "num_tokens": 5560531.0,
+ "step": 2300
+ },
+ {
+ "epoch": 6.0,
+ "eval_entropy": 0.4775148034095764,
+ "eval_loss": 0.7729372978210449,
+ "eval_mean_token_accuracy": 0.8107542619109154,
+ "eval_num_tokens": 5586522.0,
+ "eval_runtime": 95.0564,
+ "eval_samples_per_second": 17.432,
+ "eval_steps_per_second": 2.188,
+ "step": 2310
+ },
+ {
+ "entropy": 0.29565749216319326,
+ "epoch": 6.1039636127355426,
+ "grad_norm": 1.0129460096359253,
+ "learning_rate": 0.00013231293861939197,
+ "loss": 0.18774839401245116,
+ "mean_token_accuracy": 0.9347463528714587,
+ "num_tokens": 5682578.0,
+ "step": 2350
+ },
+ {
+ "entropy": 0.2680290696769953,
+ "epoch": 6.23391812865497,
+ "grad_norm": 0.7047284841537476,
+ "learning_rate": 0.00012494069362063769,
+ "loss": 0.1707102584838867,
+ "mean_token_accuracy": 0.9414009821414947,
+ "num_tokens": 5803409.0,
+ "step": 2400
+ },
+ {
+ "entropy": 0.2615845823287964,
+ "epoch": 6.363872644574399,
+ "grad_norm": 0.8583500981330872,
+ "learning_rate": 0.00011765515090412491,
+ "loss": 0.1720401382446289,
+ "mean_token_accuracy": 0.941322917342186,
+ "num_tokens": 5927439.0,
+ "step": 2450
+ },
+ {
+ "entropy": 0.26200685687363146,
+ "epoch": 6.493827160493828,
+ "grad_norm": 0.9596366286277771,
+ "learning_rate": 0.00011047128043561942,
+ "loss": 0.16701702117919923,
+ "mean_token_accuracy": 0.9415252614021301,
+ "num_tokens": 6052633.0,
+ "step": 2500
+ },
+ {
+ "entropy": 0.2842141987383366,
+ "epoch": 6.623781676413255,
+ "grad_norm": 1.1430736780166626,
+ "learning_rate": 0.00010340384326987922,
+ "loss": 0.18171667098999023,
+ "mean_token_accuracy": 0.935710808634758,
+ "num_tokens": 6167503.0,
+ "step": 2550
+ },
+ {
+ "entropy": 0.27876456022262575,
+ "epoch": 6.753736192332683,
+ "grad_norm": 1.0698399543762207,
+ "learning_rate": 9.646736122038319e-05,
+ "loss": 0.17995393753051758,
+ "mean_token_accuracy": 0.9379886907339096,
+ "num_tokens": 6284023.0,
+ "step": 2600
+ },
+ {
+ "entropy": 0.2657020591944456,
+ "epoch": 6.883690708252112,
+ "grad_norm": 0.8577425479888916,
+ "learning_rate": 8.96760870206418e-05,
+ "loss": 0.17407890319824218,
+ "mean_token_accuracy": 0.9403641560673713,
+ "num_tokens": 6408327.0,
+ "step": 2650
+ },
+ {
+ "epoch": 7.0,
+ "eval_entropy": 0.4080844047264411,
+ "eval_loss": 0.8632635474205017,
+ "eval_mean_token_accuracy": 0.8096621922002389,
+ "eval_num_tokens": 6517609.0,
+ "eval_runtime": 95.784,
+ "eval_samples_per_second": 17.299,
+ "eval_steps_per_second": 2.172,
+ "step": 2695
+ },
+ {
+ "entropy": 0.25731984889087967,
+ "epoch": 7.012995451591943,
+ "grad_norm": 0.5453973412513733,
+ "learning_rate": 8.304397503839827e-05,
+ "loss": 0.16415910720825194,
+ "mean_token_accuracy": 0.9442648204726789,
+ "num_tokens": 6530430.0,
+ "step": 2700
+ },
+ {
+ "entropy": 0.20041885927319528,
+ "epoch": 7.142949967511371,
+ "grad_norm": 0.817810595035553,
+ "learning_rate": 7.658465260289818e-05,
+ "loss": 0.11065036773681641,
+ "mean_token_accuracy": 0.9624890893697738,
+ "num_tokens": 6648623.0,
+ "step": 2750
+ },
+ {
+ "entropy": 0.19894523780792953,
+ "epoch": 7.272904483430799,
+ "grad_norm": 0.7567524313926697,
+ "learning_rate": 7.031139200413975e-05,
+ "loss": 0.10799540519714355,
+ "mean_token_accuracy": 0.962652114033699,
+ "num_tokens": 6768614.0,
+ "step": 2800
+ },
+ {
+ "entropy": 0.19293844722211362,
+ "epoch": 7.402858999350228,
+ "grad_norm": 0.8500803709030151,
+ "learning_rate": 6.423708322164252e-05,
+ "loss": 0.10800057411193847,
+ "mean_token_accuracy": 0.9634559273719787,
+ "num_tokens": 6892420.0,
+ "step": 2850
+ },
+ {
+ "entropy": 0.2034384235367179,
+ "epoch": 7.532813515269655,
+ "grad_norm": 0.9063189029693604,
+ "learning_rate": 5.837420743876688e-05,
+ "loss": 0.11020708084106445,
+ "mean_token_accuracy": 0.9601405537128449,
+ "num_tokens": 7009857.0,
+ "step": 2900
+ },
+ {
+ "entropy": 0.19702354729175567,
+ "epoch": 7.662768031189084,
+ "grad_norm": 0.8015642166137695,
+ "learning_rate": 5.27348113970081e-05,
+ "loss": 0.10956458091735839,
+ "mean_token_accuracy": 0.9614202988147735,
+ "num_tokens": 7132184.0,
+ "step": 2950
+ },
+ {
+ "entropy": 0.20174269448965787,
+ "epoch": 7.792722547108512,
+ "grad_norm": 0.8129525184631348,
+ "learning_rate": 4.733048264295942e-05,
+ "loss": 0.11115622520446777,
+ "mean_token_accuracy": 0.9607632473111153,
+ "num_tokens": 7250668.0,
+ "step": 3000
+ },
+ {
+ "entropy": 0.20127187445759773,
+ "epoch": 7.92267706302794,
+ "grad_norm": 0.916645884513855,
+ "learning_rate": 4.2172325718805845e-05,
+ "loss": 0.10734798431396485,
+ "mean_token_accuracy": 0.9610647630691528,
+ "num_tokens": 7373911.0,
+ "step": 3050
+ },
+ {
+ "epoch": 8.0,
+ "eval_entropy": 0.35548742294598085,
+ "eval_loss": 0.9974557757377625,
+ "eval_mean_token_accuracy": 0.8071597482149417,
+ "eval_num_tokens": 7448696.0,
+ "eval_runtime": 93.987,
+ "eval_samples_per_second": 17.63,
+ "eval_steps_per_second": 2.213,
+ "step": 3080
+ },
+ {
+ "entropy": 0.18474318536382225,
+ "epoch": 8.05198180636777,
+ "grad_norm": 0.5207017660140991,
+ "learning_rate": 3.727093934527048e-05,
+ "loss": 0.09326278686523437,
+ "mean_token_accuracy": 0.967371942409918,
+ "num_tokens": 7497025.0,
+ "step": 3100
+ },
+ {
+ "entropy": 0.165668417327106,
+ "epoch": 8.1819363222872,
+ "grad_norm": 0.6080297827720642,
+ "learning_rate": 3.263639464389803e-05,
+ "loss": 0.0780036211013794,
+ "mean_token_accuracy": 0.9718483003973961,
+ "num_tokens": 7615828.0,
+ "step": 3150
+ },
+ {
+ "entropy": 0.16921501949429513,
+ "epoch": 8.311890838206628,
+ "grad_norm": 0.48100438714027405,
+ "learning_rate": 2.8278214443421888e-05,
+ "loss": 0.07943209648132324,
+ "mean_token_accuracy": 0.9710959446430206,
+ "num_tokens": 7734498.0,
+ "step": 3200
+ },
+ {
+ "entropy": 0.1627163116633892,
+ "epoch": 8.441845354126055,
+ "grad_norm": 0.5401694774627686,
+ "learning_rate": 2.4205353712736153e-05,
+ "loss": 0.0771342420578003,
+ "mean_token_accuracy": 0.9723158326745033,
+ "num_tokens": 7855780.0,
+ "step": 3250
+ },
+ {
+ "entropy": 0.1762950336188078,
+ "epoch": 8.571799870045485,
+ "grad_norm": 0.48023927211761475,
+ "learning_rate": 2.0426181160676853e-05,
+ "loss": 0.0832386302947998,
+ "mean_token_accuracy": 0.9702270576357841,
+ "num_tokens": 7968890.0,
+ "step": 3300
+ },
+ {
+ "entropy": 0.1601188000291586,
+ "epoch": 8.701754385964913,
+ "grad_norm": 0.46383845806121826,
+ "learning_rate": 1.6948462040421234e-05,
+ "loss": 0.07640301704406738,
+ "mean_token_accuracy": 0.9734132272005082,
+ "num_tokens": 8094439.0,
+ "step": 3350
+ },
+ {
+ "entropy": 0.159210016541183,
+ "epoch": 8.83170890188434,
+ "grad_norm": 0.7575987577438354,
+ "learning_rate": 1.3779342193837082e-05,
+ "loss": 0.07497062206268311,
+ "mean_token_accuracy": 0.9726977890729904,
+ "num_tokens": 8219578.0,
+ "step": 3400
+ },
+ {
+ "entropy": 0.15734629351645707,
+ "epoch": 8.961663417803768,
+ "grad_norm": 0.4772493839263916,
+ "learning_rate": 1.0925333368567306e-05,
+ "loss": 0.076096453666687,
+ "mean_token_accuracy": 0.9726072958111763,
+ "num_tokens": 8345441.0,
+ "step": 3450
+ },
+ {
+ "epoch": 9.0,
+ "eval_entropy": 0.3270084271207452,
+ "eval_loss": 1.1192649602890015,
+ "eval_mean_token_accuracy": 0.8066606197792751,
+ "eval_num_tokens": 8379783.0,
+ "eval_runtime": 95.3174,
+ "eval_samples_per_second": 17.384,
+ "eval_steps_per_second": 2.182,
+ "step": 3465
+ },
+ {
+ "entropy": 0.14795680003399825,
+ "epoch": 9.0909681611436,
+ "grad_norm": 0.418293833732605,
+ "learning_rate": 8.392299838019066e-06,
+ "loss": 0.06981793403625489,
+ "mean_token_accuracy": 0.976432801491052,
+ "num_tokens": 8466213.0,
+ "step": 3500
+ },
+ {
+ "entropy": 0.14579848881810903,
+ "epoch": 9.220922677063028,
+ "grad_norm": 0.5972597002983093,
+ "learning_rate": 6.1854463517507314e-06,
+ "loss": 0.06543442726135254,
+ "mean_token_accuracy": 0.9764857029914856,
+ "num_tokens": 8591232.0,
+ "step": 3550
+ },
+ {
+ "entropy": 0.15288641860708593,
+ "epoch": 9.350877192982455,
+ "grad_norm": 0.5795395374298096,
+ "learning_rate": 4.3093074410147585e-06,
+ "loss": 0.0663717269897461,
+ "mean_token_accuracy": 0.9753753918409348,
+ "num_tokens": 8714639.0,
+ "step": 3600
+ },
+ {
+ "entropy": 0.15911089312285184,
+ "epoch": 9.480831708901885,
+ "grad_norm": 0.358585000038147,
+ "learning_rate": 2.7677381014317833e-06,
+ "loss": 0.06874057292938232,
+ "mean_token_accuracy": 0.974975820183754,
+ "num_tokens": 8833066.0,
+ "step": 3650
+ },
+ {
+ "entropy": 0.15434207927435636,
+ "epoch": 9.610786224821313,
+ "grad_norm": 0.733188271522522,
+ "learning_rate": 1.5639058719397926e-06,
+ "loss": 0.07093849658966064,
+ "mean_token_accuracy": 0.9747754210233688,
+ "num_tokens": 8949213.0,
+ "step": 3700
+ },
+ {
+ "entropy": 0.15123332396149636,
+ "epoch": 9.74074074074074,
+ "grad_norm": 0.40724924206733704,
+ "learning_rate": 7.002843262949894e-07,
+ "loss": 0.06883333206176757,
+ "mean_token_accuracy": 0.9754865205287934,
+ "num_tokens": 9069096.0,
+ "step": 3750
+ },
+ {
+ "entropy": 0.1510397135093808,
+ "epoch": 9.870695256660168,
+ "grad_norm": 0.4793565571308136,
+ "learning_rate": 1.7864799049715843e-07,
+ "loss": 0.06909588813781738,
+ "mean_token_accuracy": 0.9750119921565056,
+ "num_tokens": 9188485.0,
+ "step": 3800
+ },
+ {
+ "entropy": 0.1522918649804053,
+ "epoch": 10.0,
+ "grad_norm": 0.9604567885398865,
+ "learning_rate": 6.869658312007251e-11,
+ "loss": 0.06686938762664794,
+ "mean_token_accuracy": 0.975529134273529,
+ "num_tokens": 9310870.0,
+ "step": 3850
+ },
+ {
+ "epoch": 10.0,
+ "eval_entropy": 0.31789582760001606,
+ "eval_loss": 1.177817702293396,
+ "eval_mean_token_accuracy": 0.8061441314908174,
+ "eval_num_tokens": 9310870.0,
+ "eval_runtime": 95.9273,
+ "eval_samples_per_second": 17.274,
+ "eval_steps_per_second": 2.168,
+ "step": 3850
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3850,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": true
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 3.9069477834095616e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-770/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-770/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-770/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-770/adapter_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-770/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..13015f013b98b31173d1f215c91d6b5ff9f679cc
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-770/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 32,
+ "lora_bias": false,
+ "lora_dropout": 0.08930249338906042,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 16,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "k_proj",
+ "down_proj",
+ "q_proj",
+ "v_proj",
+ "up_proj",
+ "gate_proj",
+ "o_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-770/chat_template.jinja b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-770/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-770/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-770/tokenizer_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-770/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-770/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-770/trainer_state.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-770/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..03ad2400fffadeb77e32db942ac0326db2f5c5bc
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-770/trainer_state.json
@@ -0,0 +1,206 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 2.0,
+ "eval_steps": 500,
+ "global_step": 770,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.8437769564986228,
+ "epoch": 0.1299545159194282,
+ "grad_norm": 1.6047701835632324,
+ "learning_rate": 4.2543894243185634e-05,
+ "loss": 1.6604945373535156,
+ "mean_token_accuracy": 0.6506646998226643,
+ "num_tokens": 127115.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.9294079938530921,
+ "epoch": 0.2599090318388564,
+ "grad_norm": 1.1633784770965576,
+ "learning_rate": 8.595603122602811e-05,
+ "loss": 0.8363616180419922,
+ "mean_token_accuracy": 0.7776120388507843,
+ "num_tokens": 251796.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.8483590838313103,
+ "epoch": 0.3898635477582846,
+ "grad_norm": 0.8963253498077393,
+ "learning_rate": 0.00012936816820887062,
+ "loss": 0.7476110076904297,
+ "mean_token_accuracy": 0.7931549972295762,
+ "num_tokens": 366505.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.7840312227606774,
+ "epoch": 0.5198180636777128,
+ "grad_norm": 0.7876676917076111,
+ "learning_rate": 0.00017278030519171308,
+ "loss": 0.6980364990234375,
+ "mean_token_accuracy": 0.8027918112277984,
+ "num_tokens": 493150.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.7967149233818054,
+ "epoch": 0.649772579597141,
+ "grad_norm": 0.7976633906364441,
+ "learning_rate": 0.00021619244217455558,
+ "loss": 0.6958013153076172,
+ "mean_token_accuracy": 0.8039823538064956,
+ "num_tokens": 607197.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.7755034640431404,
+ "epoch": 0.7797270955165692,
+ "grad_norm": 0.7086226940155029,
+ "learning_rate": 0.00025960457915739807,
+ "loss": 0.6792864990234375,
+ "mean_token_accuracy": 0.8074422472715378,
+ "num_tokens": 726579.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.7452501457929611,
+ "epoch": 0.9096816114359974,
+ "grad_norm": 0.6065998673439026,
+ "learning_rate": 0.00030301671614024056,
+ "loss": 0.6627902984619141,
+ "mean_token_accuracy": 0.8118429997563362,
+ "num_tokens": 848944.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.7490686719807295,
+ "eval_loss": 0.718986451625824,
+ "eval_mean_token_accuracy": 0.7978653156986604,
+ "eval_num_tokens": 931087.0,
+ "eval_runtime": 93.2605,
+ "eval_samples_per_second": 17.767,
+ "eval_steps_per_second": 2.23,
+ "step": 385
+ },
+ {
+ "entropy": 0.7291471499893534,
+ "epoch": 1.0389863547758285,
+ "grad_norm": 0.7773680090904236,
+ "learning_rate": 0.0003342599904174574,
+ "loss": 0.6407540893554687,
+ "mean_token_accuracy": 0.816381629387937,
+ "num_tokens": 968580.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.6813811221718789,
+ "epoch": 1.1689408706952567,
+ "grad_norm": 0.7393642067909241,
+ "learning_rate": 0.0003339921524880377,
+ "loss": 0.6079624938964844,
+ "mean_token_accuracy": 0.825104131102562,
+ "num_tokens": 1090002.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.682051141411066,
+ "epoch": 1.2988953866146848,
+ "grad_norm": 0.6833657026290894,
+ "learning_rate": 0.0003333814684455912,
+ "loss": 0.6068630218505859,
+ "mean_token_accuracy": 0.8236467111110687,
+ "num_tokens": 1212161.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.6863492175936698,
+ "epoch": 1.428849902534113,
+ "grad_norm": 0.6060423851013184,
+ "learning_rate": 0.0003324291930928916,
+ "loss": 0.6073496627807617,
+ "mean_token_accuracy": 0.8244831630587578,
+ "num_tokens": 1332852.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.6790362980961799,
+ "epoch": 1.5588044184535412,
+ "grad_norm": 0.7274985313415527,
+ "learning_rate": 0.00033113728311731064,
+ "loss": 0.5929213714599609,
+ "mean_token_accuracy": 0.8260384351015091,
+ "num_tokens": 1453569.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.6581329819560051,
+ "epoch": 1.6887589343729694,
+ "grad_norm": 0.6212530136108398,
+ "learning_rate": 0.00032950839307031545,
+ "loss": 0.5812822341918945,
+ "mean_token_accuracy": 0.8288319337368012,
+ "num_tokens": 1577152.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.6893622297048568,
+ "epoch": 1.8187134502923976,
+ "grad_norm": 0.5525115132331848,
+ "learning_rate": 0.0003275458699130301,
+ "loss": 0.6035601425170899,
+ "mean_token_accuracy": 0.8239563027024269,
+ "num_tokens": 1691926.0,
+ "step": 700
+ },
+ {
+ "entropy": 0.6713312044739723,
+ "epoch": 1.9486679662118258,
+ "grad_norm": 0.746279776096344,
+ "learning_rate": 0.0003252537461390686,
+ "loss": 0.59674560546875,
+ "mean_token_accuracy": 0.8267503470182419,
+ "num_tokens": 1816059.0,
+ "step": 750
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6752213362890941,
+ "eval_loss": 0.6740940809249878,
+ "eval_mean_token_accuracy": 0.8084349325643136,
+ "eval_num_tokens": 1862174.0,
+ "eval_runtime": 94.8197,
+ "eval_samples_per_second": 17.475,
+ "eval_steps_per_second": 2.194,
+ "step": 770
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3850,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 7.826295845594112e+16,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..2fe40bb9bc751e70fba503c202f8e390373fc9bf
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/README.md
@@ -0,0 +1,58 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: transformers
+model_name: Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2
+tags:
+- generated_from_trainer
+- trl
+- sft
+licence: license
+---
+
+# Model Card for Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2
+
+This model is a fine-tuned version of [Qwen/Qwen3.5-4B-Base](https://huggingface.co/Qwen/Qwen3.5-4B-Base).
+It has been trained using [TRL](https://github.com/huggingface/trl).
+
+## Quick start
+
+```python
+from transformers import pipeline
+
+question = "If you had a time machine, but could only go to the past or the future once and never return, which would you choose and why?"
+generator = pipeline("text-generation", model="None", device="cuda")
+output = generator([{"role": "user", "content": question}], max_new_tokens=128, return_full_text=False)[0]
+print(output["generated_text"])
+```
+
+## Training procedure
+
+[
](https://wandb.ai/katriin-kukk/Cross_lingual_morphological_generalization/runs/eoqvqgs6)
+
+
+
+This model was trained with SFT.
+
+### Framework versions
+
+- TRL: 0.29.0
+- Transformers: 5.5.4
+- Pytorch: 2.10.0
+- Datasets: 4.6.1
+- Tokenizers: 0.22.2
+
+## Citations
+
+
+
+Cite TRL as:
+
+```bibtex
+@software{vonwerra2020trl,
+ title = {{TRL: Transformers Reinforcement Learning}},
+ author = {von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and Beeching, Edward and Thrush, Tristan and Lambert, Nathan and Huang, Shengyi and Rasul, Kashif and Gallouédec, Quentin},
+ license = {Apache-2.0},
+ url = {https://github.com/huggingface/trl},
+ year = {2020}
+}
+```
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1155/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1155/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1155/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1155/adapter_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1155/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..26742a540eca0befcc61f09a270940ded2abaca9
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1155/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 128,
+ "lora_bias": false,
+ "lora_dropout": 0.07386918480460901,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "k_proj",
+ "down_proj",
+ "q_proj",
+ "v_proj",
+ "up_proj",
+ "gate_proj",
+ "o_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1155/chat_template.jinja b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1155/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1155/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1155/tokenizer_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1155/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1155/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1155/trainer_state.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1155/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..4cd57faab60bc691a13688d73c6f485ab898cb12
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1155/trainer_state.json
@@ -0,0 +1,297 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 3.0,
+ "eval_steps": 500,
+ "global_step": 1155,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5888759356737137,
+ "epoch": 0.1299545159194282,
+ "grad_norm": 1.1323111057281494,
+ "learning_rate": 3.473456711878286e-05,
+ "loss": 1.4242234802246094,
+ "mean_token_accuracy": 0.6861670060455799,
+ "num_tokens": 127115.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.8638093116879463,
+ "epoch": 0.2599090318388564,
+ "grad_norm": 0.7135503888130188,
+ "learning_rate": 7.017800295427557e-05,
+ "loss": 0.7731611633300781,
+ "mean_token_accuracy": 0.7900564879179001,
+ "num_tokens": 251796.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.8165761995315551,
+ "epoch": 0.3898635477582846,
+ "grad_norm": 0.5835468173027039,
+ "learning_rate": 0.0001056214387897683,
+ "loss": 0.7186955261230469,
+ "mean_token_accuracy": 0.7992442399263382,
+ "num_tokens": 366505.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.7586414662003517,
+ "epoch": 0.5198180636777128,
+ "grad_norm": 0.5721960663795471,
+ "learning_rate": 0.00014106487462526099,
+ "loss": 0.6786507415771484,
+ "mean_token_accuracy": 0.8069057819247246,
+ "num_tokens": 493150.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.7753578001260757,
+ "epoch": 0.649772579597141,
+ "grad_norm": 0.6396453380584717,
+ "learning_rate": 0.00017650831046075373,
+ "loss": 0.6813118743896485,
+ "mean_token_accuracy": 0.806209616959095,
+ "num_tokens": 607197.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.7591985522210598,
+ "epoch": 0.7797270955165692,
+ "grad_norm": 0.577964186668396,
+ "learning_rate": 0.00021195174629624644,
+ "loss": 0.67177490234375,
+ "mean_token_accuracy": 0.8090321889519692,
+ "num_tokens": 726579.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.731777046918869,
+ "epoch": 0.9096816114359974,
+ "grad_norm": 0.597808301448822,
+ "learning_rate": 0.00024739518213173913,
+ "loss": 0.6601863861083984,
+ "mean_token_accuracy": 0.8120808112621307,
+ "num_tokens": 848944.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.7157932419616443,
+ "eval_loss": 0.7318932414054871,
+ "eval_mean_token_accuracy": 0.7979064600972029,
+ "eval_num_tokens": 931087.0,
+ "eval_runtime": 94.7419,
+ "eval_samples_per_second": 17.49,
+ "eval_steps_per_second": 2.195,
+ "step": 385
+ },
+ {
+ "entropy": 0.7265488104005555,
+ "epoch": 1.0389863547758285,
+ "grad_norm": 0.7224313020706177,
+ "learning_rate": 0.00027290346308950203,
+ "loss": 0.6452214050292969,
+ "mean_token_accuracy": 0.8156503406002293,
+ "num_tokens": 968580.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.6715995308756828,
+ "epoch": 1.1689408706952567,
+ "grad_norm": 0.6290106177330017,
+ "learning_rate": 0.000272684789300892,
+ "loss": 0.6042086791992187,
+ "mean_token_accuracy": 0.8257312250137329,
+ "num_tokens": 1090002.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.6747516925632954,
+ "epoch": 1.2988953866146848,
+ "grad_norm": 0.66269850730896,
+ "learning_rate": 0.00027218620198917387,
+ "loss": 0.6056245803833008,
+ "mean_token_accuracy": 0.8240418726205826,
+ "num_tokens": 1212161.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.6872365647554397,
+ "epoch": 1.428849902534113,
+ "grad_norm": 0.4894042909145355,
+ "learning_rate": 0.00027140872562641226,
+ "loss": 0.615659523010254,
+ "mean_token_accuracy": 0.8227636247873307,
+ "num_tokens": 1332852.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.6712110449373722,
+ "epoch": 1.5588044184535412,
+ "grad_norm": 0.9492104649543762,
+ "learning_rate": 0.00027035395773182937,
+ "loss": 0.595902099609375,
+ "mean_token_accuracy": 0.8245656326413154,
+ "num_tokens": 1453569.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.6576551312208175,
+ "epoch": 1.6887589343729694,
+ "grad_norm": 0.6480705142021179,
+ "learning_rate": 0.00026902406558930333,
+ "loss": 0.5864276885986328,
+ "mean_token_accuracy": 0.8273968940973282,
+ "num_tokens": 1577152.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.6850327280163765,
+ "epoch": 1.8187134502923976,
+ "grad_norm": 0.516665518283844,
+ "learning_rate": 0.0002674217817941425,
+ "loss": 0.6083594512939453,
+ "mean_token_accuracy": 0.8232863625884056,
+ "num_tokens": 1691926.0,
+ "step": 700
+ },
+ {
+ "entropy": 0.6665814685821533,
+ "epoch": 1.9486679662118258,
+ "grad_norm": 0.6762926578521729,
+ "learning_rate": 0.00026555039863828637,
+ "loss": 0.6011666488647461,
+ "mean_token_accuracy": 0.8256330332159996,
+ "num_tokens": 1816059.0,
+ "step": 750
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6573245783264821,
+ "eval_loss": 0.6852125525474548,
+ "eval_mean_token_accuracy": 0.808370346060166,
+ "eval_num_tokens": 1862174.0,
+ "eval_runtime": 94.2081,
+ "eval_samples_per_second": 17.589,
+ "eval_steps_per_second": 2.208,
+ "step": 770
+ },
+ {
+ "entropy": 0.612116006301276,
+ "epoch": 2.077972709551657,
+ "grad_norm": 0.5827146172523499,
+ "learning_rate": 0.0002634137613454702,
+ "loss": 0.5394676208496094,
+ "mean_token_accuracy": 0.8394461966040146,
+ "num_tokens": 1931345.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.5671873143315316,
+ "epoch": 2.207927225471085,
+ "grad_norm": 0.523046612739563,
+ "learning_rate": 0.00026101626017025327,
+ "loss": 0.4936346435546875,
+ "mean_token_accuracy": 0.8486950224637986,
+ "num_tokens": 2057087.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.5720223182439804,
+ "epoch": 2.3378817413905133,
+ "grad_norm": 0.47807472944259644,
+ "learning_rate": 0.0002583628213771449,
+ "loss": 0.5041716003417969,
+ "mean_token_accuracy": 0.8464827784895896,
+ "num_tokens": 2179801.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.5660848580300808,
+ "epoch": 2.4678362573099415,
+ "grad_norm": 0.5875942707061768,
+ "learning_rate": 0.0002554588971183651,
+ "loss": 0.4948244094848633,
+ "mean_token_accuracy": 0.8494577088952064,
+ "num_tokens": 2300775.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.5680096058547497,
+ "epoch": 2.5977907732293697,
+ "grad_norm": 0.5312080383300781,
+ "learning_rate": 0.0002523104542310364,
+ "loss": 0.5015496826171875,
+ "mean_token_accuracy": 0.847268882393837,
+ "num_tokens": 2425909.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.5735848733782768,
+ "epoch": 2.727745289148798,
+ "grad_norm": 0.5523076057434082,
+ "learning_rate": 0.0002489239619768288,
+ "loss": 0.4966643524169922,
+ "mean_token_accuracy": 0.8478639185428619,
+ "num_tokens": 2542506.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.5666417135298252,
+ "epoch": 2.857699805068226,
+ "grad_norm": 0.4719190299510956,
+ "learning_rate": 0.00024530637874924604,
+ "loss": 0.5055577850341797,
+ "mean_token_accuracy": 0.8468799209594726,
+ "num_tokens": 2666261.0,
+ "step": 1100
+ },
+ {
+ "entropy": 0.5885502319037914,
+ "epoch": 2.9876543209876543,
+ "grad_norm": 0.5237556099891663,
+ "learning_rate": 0.00024146513777587047,
+ "loss": 0.5136178588867187,
+ "mean_token_accuracy": 0.845189718902111,
+ "num_tokens": 2781874.0,
+ "step": 1150
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.6404194602599511,
+ "eval_loss": 0.6698319315910339,
+ "eval_mean_token_accuracy": 0.8120474551732724,
+ "eval_num_tokens": 2793261.0,
+ "eval_runtime": 94.1854,
+ "eval_samples_per_second": 17.593,
+ "eval_steps_per_second": 2.208,
+ "step": 1155
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3850,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.224542794634834e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1540/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1540/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1540/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1540/adapter_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1540/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..26742a540eca0befcc61f09a270940ded2abaca9
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1540/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 128,
+ "lora_bias": false,
+ "lora_dropout": 0.07386918480460901,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "k_proj",
+ "down_proj",
+ "q_proj",
+ "v_proj",
+ "up_proj",
+ "gate_proj",
+ "o_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1540/chat_template.jinja b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1540/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1540/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1540/tokenizer_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1540/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1540/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1540/trainer_state.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1540/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..ae393dbec0a80faa356a941280a711997de435b4
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1540/trainer_state.json
@@ -0,0 +1,378 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 4.0,
+ "eval_steps": 500,
+ "global_step": 1540,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5888759356737137,
+ "epoch": 0.1299545159194282,
+ "grad_norm": 1.1323111057281494,
+ "learning_rate": 3.473456711878286e-05,
+ "loss": 1.4242234802246094,
+ "mean_token_accuracy": 0.6861670060455799,
+ "num_tokens": 127115.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.8638093116879463,
+ "epoch": 0.2599090318388564,
+ "grad_norm": 0.7135503888130188,
+ "learning_rate": 7.017800295427557e-05,
+ "loss": 0.7731611633300781,
+ "mean_token_accuracy": 0.7900564879179001,
+ "num_tokens": 251796.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.8165761995315551,
+ "epoch": 0.3898635477582846,
+ "grad_norm": 0.5835468173027039,
+ "learning_rate": 0.0001056214387897683,
+ "loss": 0.7186955261230469,
+ "mean_token_accuracy": 0.7992442399263382,
+ "num_tokens": 366505.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.7586414662003517,
+ "epoch": 0.5198180636777128,
+ "grad_norm": 0.5721960663795471,
+ "learning_rate": 0.00014106487462526099,
+ "loss": 0.6786507415771484,
+ "mean_token_accuracy": 0.8069057819247246,
+ "num_tokens": 493150.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.7753578001260757,
+ "epoch": 0.649772579597141,
+ "grad_norm": 0.6396453380584717,
+ "learning_rate": 0.00017650831046075373,
+ "loss": 0.6813118743896485,
+ "mean_token_accuracy": 0.806209616959095,
+ "num_tokens": 607197.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.7591985522210598,
+ "epoch": 0.7797270955165692,
+ "grad_norm": 0.577964186668396,
+ "learning_rate": 0.00021195174629624644,
+ "loss": 0.67177490234375,
+ "mean_token_accuracy": 0.8090321889519692,
+ "num_tokens": 726579.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.731777046918869,
+ "epoch": 0.9096816114359974,
+ "grad_norm": 0.597808301448822,
+ "learning_rate": 0.00024739518213173913,
+ "loss": 0.6601863861083984,
+ "mean_token_accuracy": 0.8120808112621307,
+ "num_tokens": 848944.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.7157932419616443,
+ "eval_loss": 0.7318932414054871,
+ "eval_mean_token_accuracy": 0.7979064600972029,
+ "eval_num_tokens": 931087.0,
+ "eval_runtime": 94.7419,
+ "eval_samples_per_second": 17.49,
+ "eval_steps_per_second": 2.195,
+ "step": 385
+ },
+ {
+ "entropy": 0.7265488104005555,
+ "epoch": 1.0389863547758285,
+ "grad_norm": 0.7224313020706177,
+ "learning_rate": 0.00027290346308950203,
+ "loss": 0.6452214050292969,
+ "mean_token_accuracy": 0.8156503406002293,
+ "num_tokens": 968580.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.6715995308756828,
+ "epoch": 1.1689408706952567,
+ "grad_norm": 0.6290106177330017,
+ "learning_rate": 0.000272684789300892,
+ "loss": 0.6042086791992187,
+ "mean_token_accuracy": 0.8257312250137329,
+ "num_tokens": 1090002.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.6747516925632954,
+ "epoch": 1.2988953866146848,
+ "grad_norm": 0.66269850730896,
+ "learning_rate": 0.00027218620198917387,
+ "loss": 0.6056245803833008,
+ "mean_token_accuracy": 0.8240418726205826,
+ "num_tokens": 1212161.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.6872365647554397,
+ "epoch": 1.428849902534113,
+ "grad_norm": 0.4894042909145355,
+ "learning_rate": 0.00027140872562641226,
+ "loss": 0.615659523010254,
+ "mean_token_accuracy": 0.8227636247873307,
+ "num_tokens": 1332852.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.6712110449373722,
+ "epoch": 1.5588044184535412,
+ "grad_norm": 0.9492104649543762,
+ "learning_rate": 0.00027035395773182937,
+ "loss": 0.595902099609375,
+ "mean_token_accuracy": 0.8245656326413154,
+ "num_tokens": 1453569.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.6576551312208175,
+ "epoch": 1.6887589343729694,
+ "grad_norm": 0.6480705142021179,
+ "learning_rate": 0.00026902406558930333,
+ "loss": 0.5864276885986328,
+ "mean_token_accuracy": 0.8273968940973282,
+ "num_tokens": 1577152.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.6850327280163765,
+ "epoch": 1.8187134502923976,
+ "grad_norm": 0.516665518283844,
+ "learning_rate": 0.0002674217817941425,
+ "loss": 0.6083594512939453,
+ "mean_token_accuracy": 0.8232863625884056,
+ "num_tokens": 1691926.0,
+ "step": 700
+ },
+ {
+ "entropy": 0.6665814685821533,
+ "epoch": 1.9486679662118258,
+ "grad_norm": 0.6762926578521729,
+ "learning_rate": 0.00026555039863828637,
+ "loss": 0.6011666488647461,
+ "mean_token_accuracy": 0.8256330332159996,
+ "num_tokens": 1816059.0,
+ "step": 750
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6573245783264821,
+ "eval_loss": 0.6852125525474548,
+ "eval_mean_token_accuracy": 0.808370346060166,
+ "eval_num_tokens": 1862174.0,
+ "eval_runtime": 94.2081,
+ "eval_samples_per_second": 17.589,
+ "eval_steps_per_second": 2.208,
+ "step": 770
+ },
+ {
+ "entropy": 0.612116006301276,
+ "epoch": 2.077972709551657,
+ "grad_norm": 0.5827146172523499,
+ "learning_rate": 0.0002634137613454702,
+ "loss": 0.5394676208496094,
+ "mean_token_accuracy": 0.8394461966040146,
+ "num_tokens": 1931345.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.5671873143315316,
+ "epoch": 2.207927225471085,
+ "grad_norm": 0.523046612739563,
+ "learning_rate": 0.00026101626017025327,
+ "loss": 0.4936346435546875,
+ "mean_token_accuracy": 0.8486950224637986,
+ "num_tokens": 2057087.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.5720223182439804,
+ "epoch": 2.3378817413905133,
+ "grad_norm": 0.47807472944259644,
+ "learning_rate": 0.0002583628213771449,
+ "loss": 0.5041716003417969,
+ "mean_token_accuracy": 0.8464827784895896,
+ "num_tokens": 2179801.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.5660848580300808,
+ "epoch": 2.4678362573099415,
+ "grad_norm": 0.5875942707061768,
+ "learning_rate": 0.0002554588971183651,
+ "loss": 0.4948244094848633,
+ "mean_token_accuracy": 0.8494577088952064,
+ "num_tokens": 2300775.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.5680096058547497,
+ "epoch": 2.5977907732293697,
+ "grad_norm": 0.5312080383300781,
+ "learning_rate": 0.0002523104542310364,
+ "loss": 0.5015496826171875,
+ "mean_token_accuracy": 0.847268882393837,
+ "num_tokens": 2425909.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.5735848733782768,
+ "epoch": 2.727745289148798,
+ "grad_norm": 0.5523076057434082,
+ "learning_rate": 0.0002489239619768288,
+ "loss": 0.4966643524169922,
+ "mean_token_accuracy": 0.8478639185428619,
+ "num_tokens": 2542506.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.5666417135298252,
+ "epoch": 2.857699805068226,
+ "grad_norm": 0.4719190299510956,
+ "learning_rate": 0.00024530637874924604,
+ "loss": 0.5055577850341797,
+ "mean_token_accuracy": 0.8468799209594726,
+ "num_tokens": 2666261.0,
+ "step": 1100
+ },
+ {
+ "entropy": 0.5885502319037914,
+ "epoch": 2.9876543209876543,
+ "grad_norm": 0.5237556099891663,
+ "learning_rate": 0.00024146513777587047,
+ "loss": 0.5136178588867187,
+ "mean_token_accuracy": 0.845189718902111,
+ "num_tokens": 2781874.0,
+ "step": 1150
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.6404194602599511,
+ "eval_loss": 0.6698319315910339,
+ "eval_mean_token_accuracy": 0.8120474551732724,
+ "eval_num_tokens": 2793261.0,
+ "eval_runtime": 94.1854,
+ "eval_samples_per_second": 17.593,
+ "eval_steps_per_second": 2.208,
+ "step": 1155
+ },
+ {
+ "entropy": 0.48768596628203464,
+ "epoch": 3.116959064327485,
+ "grad_norm": 0.45523950457572937,
+ "learning_rate": 0.0002374081318449413,
+ "loss": 0.4034775924682617,
+ "mean_token_accuracy": 0.8715206759059848,
+ "num_tokens": 2896983.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.4783519075810909,
+ "epoch": 3.246913580246914,
+ "grad_norm": 0.5953325629234314,
+ "learning_rate": 0.0002331436970876517,
+ "loss": 0.39595497131347657,
+ "mean_token_accuracy": 0.8736060819029808,
+ "num_tokens": 3019005.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.4747548791766167,
+ "epoch": 3.3768680961663415,
+ "grad_norm": 0.5665897130966187,
+ "learning_rate": 0.00022868059584948654,
+ "loss": 0.4032122039794922,
+ "mean_token_accuracy": 0.8709388446807861,
+ "num_tokens": 3143415.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.4843991830945015,
+ "epoch": 3.50682261208577,
+ "grad_norm": 0.6089479923248291,
+ "learning_rate": 0.00022402799868579694,
+ "loss": 0.4088896179199219,
+ "mean_token_accuracy": 0.8691128557920456,
+ "num_tokens": 3263845.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.49369323417544364,
+ "epoch": 3.636777128005198,
+ "grad_norm": 0.7913607954978943,
+ "learning_rate": 0.00021919546551860548,
+ "loss": 0.4160056304931641,
+ "mean_token_accuracy": 0.8698329451680183,
+ "num_tokens": 3383117.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.47668287783861163,
+ "epoch": 3.7667316439246266,
+ "grad_norm": 0.5151591897010803,
+ "learning_rate": 0.00021419292599336078,
+ "loss": 0.4031526184082031,
+ "mean_token_accuracy": 0.8698683533072472,
+ "num_tokens": 3509724.0,
+ "step": 1450
+ },
+ {
+ "entropy": 0.4885013411939144,
+ "epoch": 3.8966861598440543,
+ "grad_norm": 0.6061839461326599,
+ "learning_rate": 0.0002090306590760024,
+ "loss": 0.4131336212158203,
+ "mean_token_accuracy": 0.8689957037568092,
+ "num_tokens": 3631110.0,
+ "step": 1500
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5588342215006168,
+ "eval_loss": 0.6875180006027222,
+ "eval_mean_token_accuracy": 0.8144597127460517,
+ "eval_num_tokens": 3724348.0,
+ "eval_runtime": 94.8129,
+ "eval_samples_per_second": 17.477,
+ "eval_steps_per_second": 2.194,
+ "step": 1540
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3850,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.6306733516670566e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1925/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1925/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1925/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1925/adapter_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1925/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..26742a540eca0befcc61f09a270940ded2abaca9
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1925/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 128,
+ "lora_bias": false,
+ "lora_dropout": 0.07386918480460901,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "k_proj",
+ "down_proj",
+ "q_proj",
+ "v_proj",
+ "up_proj",
+ "gate_proj",
+ "o_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1925/chat_template.jinja b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1925/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1925/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1925/tokenizer_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1925/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1925/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1925/trainer_state.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1925/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..47b11840ae6453b4559d145a370a59b96425630c
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1925/trainer_state.json
@@ -0,0 +1,469 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 5.0,
+ "eval_steps": 500,
+ "global_step": 1925,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5888759356737137,
+ "epoch": 0.1299545159194282,
+ "grad_norm": 1.1323111057281494,
+ "learning_rate": 3.473456711878286e-05,
+ "loss": 1.4242234802246094,
+ "mean_token_accuracy": 0.6861670060455799,
+ "num_tokens": 127115.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.8638093116879463,
+ "epoch": 0.2599090318388564,
+ "grad_norm": 0.7135503888130188,
+ "learning_rate": 7.017800295427557e-05,
+ "loss": 0.7731611633300781,
+ "mean_token_accuracy": 0.7900564879179001,
+ "num_tokens": 251796.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.8165761995315551,
+ "epoch": 0.3898635477582846,
+ "grad_norm": 0.5835468173027039,
+ "learning_rate": 0.0001056214387897683,
+ "loss": 0.7186955261230469,
+ "mean_token_accuracy": 0.7992442399263382,
+ "num_tokens": 366505.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.7586414662003517,
+ "epoch": 0.5198180636777128,
+ "grad_norm": 0.5721960663795471,
+ "learning_rate": 0.00014106487462526099,
+ "loss": 0.6786507415771484,
+ "mean_token_accuracy": 0.8069057819247246,
+ "num_tokens": 493150.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.7753578001260757,
+ "epoch": 0.649772579597141,
+ "grad_norm": 0.6396453380584717,
+ "learning_rate": 0.00017650831046075373,
+ "loss": 0.6813118743896485,
+ "mean_token_accuracy": 0.806209616959095,
+ "num_tokens": 607197.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.7591985522210598,
+ "epoch": 0.7797270955165692,
+ "grad_norm": 0.577964186668396,
+ "learning_rate": 0.00021195174629624644,
+ "loss": 0.67177490234375,
+ "mean_token_accuracy": 0.8090321889519692,
+ "num_tokens": 726579.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.731777046918869,
+ "epoch": 0.9096816114359974,
+ "grad_norm": 0.597808301448822,
+ "learning_rate": 0.00024739518213173913,
+ "loss": 0.6601863861083984,
+ "mean_token_accuracy": 0.8120808112621307,
+ "num_tokens": 848944.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.7157932419616443,
+ "eval_loss": 0.7318932414054871,
+ "eval_mean_token_accuracy": 0.7979064600972029,
+ "eval_num_tokens": 931087.0,
+ "eval_runtime": 94.7419,
+ "eval_samples_per_second": 17.49,
+ "eval_steps_per_second": 2.195,
+ "step": 385
+ },
+ {
+ "entropy": 0.7265488104005555,
+ "epoch": 1.0389863547758285,
+ "grad_norm": 0.7224313020706177,
+ "learning_rate": 0.00027290346308950203,
+ "loss": 0.6452214050292969,
+ "mean_token_accuracy": 0.8156503406002293,
+ "num_tokens": 968580.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.6715995308756828,
+ "epoch": 1.1689408706952567,
+ "grad_norm": 0.6290106177330017,
+ "learning_rate": 0.000272684789300892,
+ "loss": 0.6042086791992187,
+ "mean_token_accuracy": 0.8257312250137329,
+ "num_tokens": 1090002.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.6747516925632954,
+ "epoch": 1.2988953866146848,
+ "grad_norm": 0.66269850730896,
+ "learning_rate": 0.00027218620198917387,
+ "loss": 0.6056245803833008,
+ "mean_token_accuracy": 0.8240418726205826,
+ "num_tokens": 1212161.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.6872365647554397,
+ "epoch": 1.428849902534113,
+ "grad_norm": 0.4894042909145355,
+ "learning_rate": 0.00027140872562641226,
+ "loss": 0.615659523010254,
+ "mean_token_accuracy": 0.8227636247873307,
+ "num_tokens": 1332852.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.6712110449373722,
+ "epoch": 1.5588044184535412,
+ "grad_norm": 0.9492104649543762,
+ "learning_rate": 0.00027035395773182937,
+ "loss": 0.595902099609375,
+ "mean_token_accuracy": 0.8245656326413154,
+ "num_tokens": 1453569.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.6576551312208175,
+ "epoch": 1.6887589343729694,
+ "grad_norm": 0.6480705142021179,
+ "learning_rate": 0.00026902406558930333,
+ "loss": 0.5864276885986328,
+ "mean_token_accuracy": 0.8273968940973282,
+ "num_tokens": 1577152.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.6850327280163765,
+ "epoch": 1.8187134502923976,
+ "grad_norm": 0.516665518283844,
+ "learning_rate": 0.0002674217817941425,
+ "loss": 0.6083594512939453,
+ "mean_token_accuracy": 0.8232863625884056,
+ "num_tokens": 1691926.0,
+ "step": 700
+ },
+ {
+ "entropy": 0.6665814685821533,
+ "epoch": 1.9486679662118258,
+ "grad_norm": 0.6762926578521729,
+ "learning_rate": 0.00026555039863828637,
+ "loss": 0.6011666488647461,
+ "mean_token_accuracy": 0.8256330332159996,
+ "num_tokens": 1816059.0,
+ "step": 750
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6573245783264821,
+ "eval_loss": 0.6852125525474548,
+ "eval_mean_token_accuracy": 0.808370346060166,
+ "eval_num_tokens": 1862174.0,
+ "eval_runtime": 94.2081,
+ "eval_samples_per_second": 17.589,
+ "eval_steps_per_second": 2.208,
+ "step": 770
+ },
+ {
+ "entropy": 0.612116006301276,
+ "epoch": 2.077972709551657,
+ "grad_norm": 0.5827146172523499,
+ "learning_rate": 0.0002634137613454702,
+ "loss": 0.5394676208496094,
+ "mean_token_accuracy": 0.8394461966040146,
+ "num_tokens": 1931345.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.5671873143315316,
+ "epoch": 2.207927225471085,
+ "grad_norm": 0.523046612739563,
+ "learning_rate": 0.00026101626017025327,
+ "loss": 0.4936346435546875,
+ "mean_token_accuracy": 0.8486950224637986,
+ "num_tokens": 2057087.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.5720223182439804,
+ "epoch": 2.3378817413905133,
+ "grad_norm": 0.47807472944259644,
+ "learning_rate": 0.0002583628213771449,
+ "loss": 0.5041716003417969,
+ "mean_token_accuracy": 0.8464827784895896,
+ "num_tokens": 2179801.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.5660848580300808,
+ "epoch": 2.4678362573099415,
+ "grad_norm": 0.5875942707061768,
+ "learning_rate": 0.0002554588971183651,
+ "loss": 0.4948244094848633,
+ "mean_token_accuracy": 0.8494577088952064,
+ "num_tokens": 2300775.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.5680096058547497,
+ "epoch": 2.5977907732293697,
+ "grad_norm": 0.5312080383300781,
+ "learning_rate": 0.0002523104542310364,
+ "loss": 0.5015496826171875,
+ "mean_token_accuracy": 0.847268882393837,
+ "num_tokens": 2425909.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.5735848733782768,
+ "epoch": 2.727745289148798,
+ "grad_norm": 0.5523076057434082,
+ "learning_rate": 0.0002489239619768288,
+ "loss": 0.4966643524169922,
+ "mean_token_accuracy": 0.8478639185428619,
+ "num_tokens": 2542506.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.5666417135298252,
+ "epoch": 2.857699805068226,
+ "grad_norm": 0.4719190299510956,
+ "learning_rate": 0.00024530637874924604,
+ "loss": 0.5055577850341797,
+ "mean_token_accuracy": 0.8468799209594726,
+ "num_tokens": 2666261.0,
+ "step": 1100
+ },
+ {
+ "entropy": 0.5885502319037914,
+ "epoch": 2.9876543209876543,
+ "grad_norm": 0.5237556099891663,
+ "learning_rate": 0.00024146513777587047,
+ "loss": 0.5136178588867187,
+ "mean_token_accuracy": 0.845189718902111,
+ "num_tokens": 2781874.0,
+ "step": 1150
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.6404194602599511,
+ "eval_loss": 0.6698319315910339,
+ "eval_mean_token_accuracy": 0.8120474551732724,
+ "eval_num_tokens": 2793261.0,
+ "eval_runtime": 94.1854,
+ "eval_samples_per_second": 17.593,
+ "eval_steps_per_second": 2.208,
+ "step": 1155
+ },
+ {
+ "entropy": 0.48768596628203464,
+ "epoch": 3.116959064327485,
+ "grad_norm": 0.45523950457572937,
+ "learning_rate": 0.0002374081318449413,
+ "loss": 0.4034775924682617,
+ "mean_token_accuracy": 0.8715206759059848,
+ "num_tokens": 2896983.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.4783519075810909,
+ "epoch": 3.246913580246914,
+ "grad_norm": 0.5953325629234314,
+ "learning_rate": 0.0002331436970876517,
+ "loss": 0.39595497131347657,
+ "mean_token_accuracy": 0.8736060819029808,
+ "num_tokens": 3019005.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.4747548791766167,
+ "epoch": 3.3768680961663415,
+ "grad_norm": 0.5665897130966187,
+ "learning_rate": 0.00022868059584948654,
+ "loss": 0.4032122039794922,
+ "mean_token_accuracy": 0.8709388446807861,
+ "num_tokens": 3143415.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.4843991830945015,
+ "epoch": 3.50682261208577,
+ "grad_norm": 0.6089479923248291,
+ "learning_rate": 0.00022402799868579694,
+ "loss": 0.4088896179199219,
+ "mean_token_accuracy": 0.8691128557920456,
+ "num_tokens": 3263845.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.49369323417544364,
+ "epoch": 3.636777128005198,
+ "grad_norm": 0.7913607954978943,
+ "learning_rate": 0.00021919546551860548,
+ "loss": 0.4160056304931641,
+ "mean_token_accuracy": 0.8698329451680183,
+ "num_tokens": 3383117.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.47668287783861163,
+ "epoch": 3.7667316439246266,
+ "grad_norm": 0.5151591897010803,
+ "learning_rate": 0.00021419292599336078,
+ "loss": 0.4031526184082031,
+ "mean_token_accuracy": 0.8698683533072472,
+ "num_tokens": 3509724.0,
+ "step": 1450
+ },
+ {
+ "entropy": 0.4885013411939144,
+ "epoch": 3.8966861598440543,
+ "grad_norm": 0.6061839461326599,
+ "learning_rate": 0.0002090306590760024,
+ "loss": 0.4131336212158203,
+ "mean_token_accuracy": 0.8689957037568092,
+ "num_tokens": 3631110.0,
+ "step": 1500
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5588342215006168,
+ "eval_loss": 0.6875180006027222,
+ "eval_mean_token_accuracy": 0.8144597127460517,
+ "eval_num_tokens": 3724348.0,
+ "eval_runtime": 94.8129,
+ "eval_samples_per_second": 17.477,
+ "eval_steps_per_second": 2.194,
+ "step": 1540
+ },
+ {
+ "entropy": 0.4742255074594488,
+ "epoch": 4.025990903183885,
+ "grad_norm": 0.692275881767273,
+ "learning_rate": 0.0002037192719322604,
+ "loss": 0.39202346801757815,
+ "mean_token_accuracy": 0.8734938379508167,
+ "num_tokens": 3749929.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.3813336396217346,
+ "epoch": 4.155945419103314,
+ "grad_norm": 0.5534346699714661,
+ "learning_rate": 0.00019826967813258578,
+ "loss": 0.29004505157470706,
+ "mean_token_accuracy": 0.9025143319368363,
+ "num_tokens": 3871403.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.38276420064270494,
+ "epoch": 4.2858999350227425,
+ "grad_norm": 0.567842960357666,
+ "learning_rate": 0.00019269307522749648,
+ "loss": 0.2957651138305664,
+ "mean_token_accuracy": 0.9018667274713517,
+ "num_tokens": 3990410.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.38759261518716814,
+ "epoch": 4.41585445094217,
+ "grad_norm": 0.5272114276885986,
+ "learning_rate": 0.00018700092173941517,
+ "loss": 0.3017059135437012,
+ "mean_token_accuracy": 0.8991367623209954,
+ "num_tokens": 4107377.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.39710743524134157,
+ "epoch": 4.545808966861598,
+ "grad_norm": 0.5341858267784119,
+ "learning_rate": 0.00018120491361827535,
+ "loss": 0.30882442474365235,
+ "mean_token_accuracy": 0.8962600353360176,
+ "num_tokens": 4225238.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.39052032932639125,
+ "epoch": 4.675763482781027,
+ "grad_norm": 0.5387603044509888,
+ "learning_rate": 0.00017531696020927306,
+ "loss": 0.3076332473754883,
+ "mean_token_accuracy": 0.8972256830334664,
+ "num_tokens": 4348229.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.38603785961866377,
+ "epoch": 4.805717998700455,
+ "grad_norm": 0.6838284134864807,
+ "learning_rate": 0.00016934915978214512,
+ "loss": 0.3052788162231445,
+ "mean_token_accuracy": 0.8977296784520149,
+ "num_tokens": 4471337.0,
+ "step": 1850
+ },
+ {
+ "entropy": 0.387482148706913,
+ "epoch": 4.935672514619883,
+ "grad_norm": 0.6066106557846069,
+ "learning_rate": 0.00016331377467225454,
+ "loss": 0.3069518852233887,
+ "mean_token_accuracy": 0.8959452581405639,
+ "num_tokens": 4596516.0,
+ "step": 1900
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.47827344932235205,
+ "eval_loss": 0.734696626663208,
+ "eval_mean_token_accuracy": 0.8166056708074533,
+ "eval_num_tokens": 4655435.0,
+ "eval_runtime": 94.3316,
+ "eval_samples_per_second": 17.566,
+ "eval_steps_per_second": 2.205,
+ "step": 1925
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3850,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.037033240680018e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2310/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2310/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2310/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2310/adapter_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2310/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..26742a540eca0befcc61f09a270940ded2abaca9
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2310/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 128,
+ "lora_bias": false,
+ "lora_dropout": 0.07386918480460901,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "k_proj",
+ "down_proj",
+ "q_proj",
+ "v_proj",
+ "up_proj",
+ "gate_proj",
+ "o_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2310/chat_template.jinja b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2310/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2310/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2310/tokenizer_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2310/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2310/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2310/trainer_state.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2310/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..bbda34e8df89bdf53485a00001d762c92ec37af2
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2310/trainer_state.json
@@ -0,0 +1,560 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 6.0,
+ "eval_steps": 500,
+ "global_step": 2310,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5888759356737137,
+ "epoch": 0.1299545159194282,
+ "grad_norm": 1.1323111057281494,
+ "learning_rate": 3.473456711878286e-05,
+ "loss": 1.4242234802246094,
+ "mean_token_accuracy": 0.6861670060455799,
+ "num_tokens": 127115.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.8638093116879463,
+ "epoch": 0.2599090318388564,
+ "grad_norm": 0.7135503888130188,
+ "learning_rate": 7.017800295427557e-05,
+ "loss": 0.7731611633300781,
+ "mean_token_accuracy": 0.7900564879179001,
+ "num_tokens": 251796.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.8165761995315551,
+ "epoch": 0.3898635477582846,
+ "grad_norm": 0.5835468173027039,
+ "learning_rate": 0.0001056214387897683,
+ "loss": 0.7186955261230469,
+ "mean_token_accuracy": 0.7992442399263382,
+ "num_tokens": 366505.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.7586414662003517,
+ "epoch": 0.5198180636777128,
+ "grad_norm": 0.5721960663795471,
+ "learning_rate": 0.00014106487462526099,
+ "loss": 0.6786507415771484,
+ "mean_token_accuracy": 0.8069057819247246,
+ "num_tokens": 493150.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.7753578001260757,
+ "epoch": 0.649772579597141,
+ "grad_norm": 0.6396453380584717,
+ "learning_rate": 0.00017650831046075373,
+ "loss": 0.6813118743896485,
+ "mean_token_accuracy": 0.806209616959095,
+ "num_tokens": 607197.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.7591985522210598,
+ "epoch": 0.7797270955165692,
+ "grad_norm": 0.577964186668396,
+ "learning_rate": 0.00021195174629624644,
+ "loss": 0.67177490234375,
+ "mean_token_accuracy": 0.8090321889519692,
+ "num_tokens": 726579.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.731777046918869,
+ "epoch": 0.9096816114359974,
+ "grad_norm": 0.597808301448822,
+ "learning_rate": 0.00024739518213173913,
+ "loss": 0.6601863861083984,
+ "mean_token_accuracy": 0.8120808112621307,
+ "num_tokens": 848944.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.7157932419616443,
+ "eval_loss": 0.7318932414054871,
+ "eval_mean_token_accuracy": 0.7979064600972029,
+ "eval_num_tokens": 931087.0,
+ "eval_runtime": 94.7419,
+ "eval_samples_per_second": 17.49,
+ "eval_steps_per_second": 2.195,
+ "step": 385
+ },
+ {
+ "entropy": 0.7265488104005555,
+ "epoch": 1.0389863547758285,
+ "grad_norm": 0.7224313020706177,
+ "learning_rate": 0.00027290346308950203,
+ "loss": 0.6452214050292969,
+ "mean_token_accuracy": 0.8156503406002293,
+ "num_tokens": 968580.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.6715995308756828,
+ "epoch": 1.1689408706952567,
+ "grad_norm": 0.6290106177330017,
+ "learning_rate": 0.000272684789300892,
+ "loss": 0.6042086791992187,
+ "mean_token_accuracy": 0.8257312250137329,
+ "num_tokens": 1090002.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.6747516925632954,
+ "epoch": 1.2988953866146848,
+ "grad_norm": 0.66269850730896,
+ "learning_rate": 0.00027218620198917387,
+ "loss": 0.6056245803833008,
+ "mean_token_accuracy": 0.8240418726205826,
+ "num_tokens": 1212161.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.6872365647554397,
+ "epoch": 1.428849902534113,
+ "grad_norm": 0.4894042909145355,
+ "learning_rate": 0.00027140872562641226,
+ "loss": 0.615659523010254,
+ "mean_token_accuracy": 0.8227636247873307,
+ "num_tokens": 1332852.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.6712110449373722,
+ "epoch": 1.5588044184535412,
+ "grad_norm": 0.9492104649543762,
+ "learning_rate": 0.00027035395773182937,
+ "loss": 0.595902099609375,
+ "mean_token_accuracy": 0.8245656326413154,
+ "num_tokens": 1453569.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.6576551312208175,
+ "epoch": 1.6887589343729694,
+ "grad_norm": 0.6480705142021179,
+ "learning_rate": 0.00026902406558930333,
+ "loss": 0.5864276885986328,
+ "mean_token_accuracy": 0.8273968940973282,
+ "num_tokens": 1577152.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.6850327280163765,
+ "epoch": 1.8187134502923976,
+ "grad_norm": 0.516665518283844,
+ "learning_rate": 0.0002674217817941425,
+ "loss": 0.6083594512939453,
+ "mean_token_accuracy": 0.8232863625884056,
+ "num_tokens": 1691926.0,
+ "step": 700
+ },
+ {
+ "entropy": 0.6665814685821533,
+ "epoch": 1.9486679662118258,
+ "grad_norm": 0.6762926578521729,
+ "learning_rate": 0.00026555039863828637,
+ "loss": 0.6011666488647461,
+ "mean_token_accuracy": 0.8256330332159996,
+ "num_tokens": 1816059.0,
+ "step": 750
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6573245783264821,
+ "eval_loss": 0.6852125525474548,
+ "eval_mean_token_accuracy": 0.808370346060166,
+ "eval_num_tokens": 1862174.0,
+ "eval_runtime": 94.2081,
+ "eval_samples_per_second": 17.589,
+ "eval_steps_per_second": 2.208,
+ "step": 770
+ },
+ {
+ "entropy": 0.612116006301276,
+ "epoch": 2.077972709551657,
+ "grad_norm": 0.5827146172523499,
+ "learning_rate": 0.0002634137613454702,
+ "loss": 0.5394676208496094,
+ "mean_token_accuracy": 0.8394461966040146,
+ "num_tokens": 1931345.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.5671873143315316,
+ "epoch": 2.207927225471085,
+ "grad_norm": 0.523046612739563,
+ "learning_rate": 0.00026101626017025327,
+ "loss": 0.4936346435546875,
+ "mean_token_accuracy": 0.8486950224637986,
+ "num_tokens": 2057087.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.5720223182439804,
+ "epoch": 2.3378817413905133,
+ "grad_norm": 0.47807472944259644,
+ "learning_rate": 0.0002583628213771449,
+ "loss": 0.5041716003417969,
+ "mean_token_accuracy": 0.8464827784895896,
+ "num_tokens": 2179801.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.5660848580300808,
+ "epoch": 2.4678362573099415,
+ "grad_norm": 0.5875942707061768,
+ "learning_rate": 0.0002554588971183651,
+ "loss": 0.4948244094848633,
+ "mean_token_accuracy": 0.8494577088952064,
+ "num_tokens": 2300775.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.5680096058547497,
+ "epoch": 2.5977907732293697,
+ "grad_norm": 0.5312080383300781,
+ "learning_rate": 0.0002523104542310364,
+ "loss": 0.5015496826171875,
+ "mean_token_accuracy": 0.847268882393837,
+ "num_tokens": 2425909.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.5735848733782768,
+ "epoch": 2.727745289148798,
+ "grad_norm": 0.5523076057434082,
+ "learning_rate": 0.0002489239619768288,
+ "loss": 0.4966643524169922,
+ "mean_token_accuracy": 0.8478639185428619,
+ "num_tokens": 2542506.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.5666417135298252,
+ "epoch": 2.857699805068226,
+ "grad_norm": 0.4719190299510956,
+ "learning_rate": 0.00024530637874924604,
+ "loss": 0.5055577850341797,
+ "mean_token_accuracy": 0.8468799209594726,
+ "num_tokens": 2666261.0,
+ "step": 1100
+ },
+ {
+ "entropy": 0.5885502319037914,
+ "epoch": 2.9876543209876543,
+ "grad_norm": 0.5237556099891663,
+ "learning_rate": 0.00024146513777587047,
+ "loss": 0.5136178588867187,
+ "mean_token_accuracy": 0.845189718902111,
+ "num_tokens": 2781874.0,
+ "step": 1150
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.6404194602599511,
+ "eval_loss": 0.6698319315910339,
+ "eval_mean_token_accuracy": 0.8120474551732724,
+ "eval_num_tokens": 2793261.0,
+ "eval_runtime": 94.1854,
+ "eval_samples_per_second": 17.593,
+ "eval_steps_per_second": 2.208,
+ "step": 1155
+ },
+ {
+ "entropy": 0.48768596628203464,
+ "epoch": 3.116959064327485,
+ "grad_norm": 0.45523950457572937,
+ "learning_rate": 0.0002374081318449413,
+ "loss": 0.4034775924682617,
+ "mean_token_accuracy": 0.8715206759059848,
+ "num_tokens": 2896983.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.4783519075810909,
+ "epoch": 3.246913580246914,
+ "grad_norm": 0.5953325629234314,
+ "learning_rate": 0.0002331436970876517,
+ "loss": 0.39595497131347657,
+ "mean_token_accuracy": 0.8736060819029808,
+ "num_tokens": 3019005.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.4747548791766167,
+ "epoch": 3.3768680961663415,
+ "grad_norm": 0.5665897130966187,
+ "learning_rate": 0.00022868059584948654,
+ "loss": 0.4032122039794922,
+ "mean_token_accuracy": 0.8709388446807861,
+ "num_tokens": 3143415.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.4843991830945015,
+ "epoch": 3.50682261208577,
+ "grad_norm": 0.6089479923248291,
+ "learning_rate": 0.00022402799868579694,
+ "loss": 0.4088896179199219,
+ "mean_token_accuracy": 0.8691128557920456,
+ "num_tokens": 3263845.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.49369323417544364,
+ "epoch": 3.636777128005198,
+ "grad_norm": 0.7913607954978943,
+ "learning_rate": 0.00021919546551860548,
+ "loss": 0.4160056304931641,
+ "mean_token_accuracy": 0.8698329451680183,
+ "num_tokens": 3383117.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.47668287783861163,
+ "epoch": 3.7667316439246266,
+ "grad_norm": 0.5151591897010803,
+ "learning_rate": 0.00021419292599336078,
+ "loss": 0.4031526184082031,
+ "mean_token_accuracy": 0.8698683533072472,
+ "num_tokens": 3509724.0,
+ "step": 1450
+ },
+ {
+ "entropy": 0.4885013411939144,
+ "epoch": 3.8966861598440543,
+ "grad_norm": 0.6061839461326599,
+ "learning_rate": 0.0002090306590760024,
+ "loss": 0.4131336212158203,
+ "mean_token_accuracy": 0.8689957037568092,
+ "num_tokens": 3631110.0,
+ "step": 1500
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5588342215006168,
+ "eval_loss": 0.6875180006027222,
+ "eval_mean_token_accuracy": 0.8144597127460517,
+ "eval_num_tokens": 3724348.0,
+ "eval_runtime": 94.8129,
+ "eval_samples_per_second": 17.477,
+ "eval_steps_per_second": 2.194,
+ "step": 1540
+ },
+ {
+ "entropy": 0.4742255074594488,
+ "epoch": 4.025990903183885,
+ "grad_norm": 0.692275881767273,
+ "learning_rate": 0.0002037192719322604,
+ "loss": 0.39202346801757815,
+ "mean_token_accuracy": 0.8734938379508167,
+ "num_tokens": 3749929.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.3813336396217346,
+ "epoch": 4.155945419103314,
+ "grad_norm": 0.5534346699714661,
+ "learning_rate": 0.00019826967813258578,
+ "loss": 0.29004505157470706,
+ "mean_token_accuracy": 0.9025143319368363,
+ "num_tokens": 3871403.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.38276420064270494,
+ "epoch": 4.2858999350227425,
+ "grad_norm": 0.567842960357666,
+ "learning_rate": 0.00019269307522749648,
+ "loss": 0.2957651138305664,
+ "mean_token_accuracy": 0.9018667274713517,
+ "num_tokens": 3990410.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.38759261518716814,
+ "epoch": 4.41585445094217,
+ "grad_norm": 0.5272114276885986,
+ "learning_rate": 0.00018700092173941517,
+ "loss": 0.3017059135437012,
+ "mean_token_accuracy": 0.8991367623209954,
+ "num_tokens": 4107377.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.39710743524134157,
+ "epoch": 4.545808966861598,
+ "grad_norm": 0.5341858267784119,
+ "learning_rate": 0.00018120491361827535,
+ "loss": 0.30882442474365235,
+ "mean_token_accuracy": 0.8962600353360176,
+ "num_tokens": 4225238.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.39052032932639125,
+ "epoch": 4.675763482781027,
+ "grad_norm": 0.5387603044509888,
+ "learning_rate": 0.00017531696020927306,
+ "loss": 0.3076332473754883,
+ "mean_token_accuracy": 0.8972256830334664,
+ "num_tokens": 4348229.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.38603785961866377,
+ "epoch": 4.805717998700455,
+ "grad_norm": 0.6838284134864807,
+ "learning_rate": 0.00016934915978214512,
+ "loss": 0.3052788162231445,
+ "mean_token_accuracy": 0.8977296784520149,
+ "num_tokens": 4471337.0,
+ "step": 1850
+ },
+ {
+ "entropy": 0.387482148706913,
+ "epoch": 4.935672514619883,
+ "grad_norm": 0.6066106557846069,
+ "learning_rate": 0.00016331377467225454,
+ "loss": 0.3069518852233887,
+ "mean_token_accuracy": 0.8959452581405639,
+ "num_tokens": 4596516.0,
+ "step": 1900
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.47827344932235205,
+ "eval_loss": 0.734696626663208,
+ "eval_mean_token_accuracy": 0.8166056708074533,
+ "eval_num_tokens": 4655435.0,
+ "eval_runtime": 94.3316,
+ "eval_samples_per_second": 17.566,
+ "eval_steps_per_second": 2.205,
+ "step": 1925
+ },
+ {
+ "entropy": 0.3189891557298114,
+ "epoch": 5.064977257959714,
+ "grad_norm": 0.4411616623401642,
+ "learning_rate": 0.00015722320608456218,
+ "loss": 0.24322872161865233,
+ "mean_token_accuracy": 0.9192550799355435,
+ "num_tokens": 4720113.0,
+ "step": 1950
+ },
+ {
+ "entropy": 0.2661040405929089,
+ "epoch": 5.1949317738791425,
+ "grad_norm": 0.5544848442077637,
+ "learning_rate": 0.00015108996861225613,
+ "loss": 0.19045450210571288,
+ "mean_token_accuracy": 0.9353114965558053,
+ "num_tokens": 4842234.0,
+ "step": 2000
+ },
+ {
+ "entropy": 0.2827081228792667,
+ "epoch": 5.32488628979857,
+ "grad_norm": 0.5472713708877563,
+ "learning_rate": 0.0001449266645223969,
+ "loss": 0.197703857421875,
+ "mean_token_accuracy": 0.9314222651720047,
+ "num_tokens": 4959991.0,
+ "step": 2050
+ },
+ {
+ "entropy": 0.28358375541865827,
+ "epoch": 5.454840805717999,
+ "grad_norm": 0.6511954069137573,
+ "learning_rate": 0.00013874595786141455,
+ "loss": 0.20100860595703124,
+ "mean_token_accuracy": 0.930547822713852,
+ "num_tokens": 5082208.0,
+ "step": 2100
+ },
+ {
+ "entropy": 0.2757134060561657,
+ "epoch": 5.584795321637427,
+ "grad_norm": 0.5925538539886475,
+ "learning_rate": 0.0001325605484336644,
+ "loss": 0.19688623428344726,
+ "mean_token_accuracy": 0.9317859390377998,
+ "num_tokens": 5199699.0,
+ "step": 2150
+ },
+ {
+ "entropy": 0.2787610155344009,
+ "epoch": 5.714749837556855,
+ "grad_norm": 0.7486807703971863,
+ "learning_rate": 0.00012638314570651013,
+ "loss": 0.198111629486084,
+ "mean_token_accuracy": 0.9316030797362328,
+ "num_tokens": 5320501.0,
+ "step": 2200
+ },
+ {
+ "entropy": 0.2822082404047251,
+ "epoch": 5.844704353476283,
+ "grad_norm": 0.5981082320213318,
+ "learning_rate": 0.00012022644269555052,
+ "loss": 0.19833782196044922,
+ "mean_token_accuracy": 0.9301980781555176,
+ "num_tokens": 5442291.0,
+ "step": 2250
+ },
+ {
+ "entropy": 0.29100623600184916,
+ "epoch": 5.974658869395712,
+ "grad_norm": 0.6062291860580444,
+ "learning_rate": 0.00011410308988365154,
+ "loss": 0.20115615844726562,
+ "mean_token_accuracy": 0.9285238465666771,
+ "num_tokens": 5560531.0,
+ "step": 2300
+ },
+ {
+ "epoch": 6.0,
+ "eval_entropy": 0.419523782025163,
+ "eval_loss": 0.8228668570518494,
+ "eval_mean_token_accuracy": 0.8120436149720962,
+ "eval_num_tokens": 5586522.0,
+ "eval_runtime": 94.2282,
+ "eval_samples_per_second": 17.585,
+ "eval_steps_per_second": 2.207,
+ "step": 2310
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3850,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.4437075074865357e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2695/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2695/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2695/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2695/adapter_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2695/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..26742a540eca0befcc61f09a270940ded2abaca9
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2695/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 128,
+ "lora_bias": false,
+ "lora_dropout": 0.07386918480460901,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "k_proj",
+ "down_proj",
+ "q_proj",
+ "v_proj",
+ "up_proj",
+ "gate_proj",
+ "o_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2695/chat_template.jinja b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2695/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2695/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2695/tokenizer_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2695/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2695/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2695/trainer_state.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2695/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..7b4f04e3eade61a925600777e08b42c9257a4a4d
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2695/trainer_state.json
@@ -0,0 +1,641 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 7.0,
+ "eval_steps": 500,
+ "global_step": 2695,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5888759356737137,
+ "epoch": 0.1299545159194282,
+ "grad_norm": 1.1323111057281494,
+ "learning_rate": 3.473456711878286e-05,
+ "loss": 1.4242234802246094,
+ "mean_token_accuracy": 0.6861670060455799,
+ "num_tokens": 127115.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.8638093116879463,
+ "epoch": 0.2599090318388564,
+ "grad_norm": 0.7135503888130188,
+ "learning_rate": 7.017800295427557e-05,
+ "loss": 0.7731611633300781,
+ "mean_token_accuracy": 0.7900564879179001,
+ "num_tokens": 251796.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.8165761995315551,
+ "epoch": 0.3898635477582846,
+ "grad_norm": 0.5835468173027039,
+ "learning_rate": 0.0001056214387897683,
+ "loss": 0.7186955261230469,
+ "mean_token_accuracy": 0.7992442399263382,
+ "num_tokens": 366505.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.7586414662003517,
+ "epoch": 0.5198180636777128,
+ "grad_norm": 0.5721960663795471,
+ "learning_rate": 0.00014106487462526099,
+ "loss": 0.6786507415771484,
+ "mean_token_accuracy": 0.8069057819247246,
+ "num_tokens": 493150.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.7753578001260757,
+ "epoch": 0.649772579597141,
+ "grad_norm": 0.6396453380584717,
+ "learning_rate": 0.00017650831046075373,
+ "loss": 0.6813118743896485,
+ "mean_token_accuracy": 0.806209616959095,
+ "num_tokens": 607197.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.7591985522210598,
+ "epoch": 0.7797270955165692,
+ "grad_norm": 0.577964186668396,
+ "learning_rate": 0.00021195174629624644,
+ "loss": 0.67177490234375,
+ "mean_token_accuracy": 0.8090321889519692,
+ "num_tokens": 726579.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.731777046918869,
+ "epoch": 0.9096816114359974,
+ "grad_norm": 0.597808301448822,
+ "learning_rate": 0.00024739518213173913,
+ "loss": 0.6601863861083984,
+ "mean_token_accuracy": 0.8120808112621307,
+ "num_tokens": 848944.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.7157932419616443,
+ "eval_loss": 0.7318932414054871,
+ "eval_mean_token_accuracy": 0.7979064600972029,
+ "eval_num_tokens": 931087.0,
+ "eval_runtime": 94.7419,
+ "eval_samples_per_second": 17.49,
+ "eval_steps_per_second": 2.195,
+ "step": 385
+ },
+ {
+ "entropy": 0.7265488104005555,
+ "epoch": 1.0389863547758285,
+ "grad_norm": 0.7224313020706177,
+ "learning_rate": 0.00027290346308950203,
+ "loss": 0.6452214050292969,
+ "mean_token_accuracy": 0.8156503406002293,
+ "num_tokens": 968580.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.6715995308756828,
+ "epoch": 1.1689408706952567,
+ "grad_norm": 0.6290106177330017,
+ "learning_rate": 0.000272684789300892,
+ "loss": 0.6042086791992187,
+ "mean_token_accuracy": 0.8257312250137329,
+ "num_tokens": 1090002.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.6747516925632954,
+ "epoch": 1.2988953866146848,
+ "grad_norm": 0.66269850730896,
+ "learning_rate": 0.00027218620198917387,
+ "loss": 0.6056245803833008,
+ "mean_token_accuracy": 0.8240418726205826,
+ "num_tokens": 1212161.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.6872365647554397,
+ "epoch": 1.428849902534113,
+ "grad_norm": 0.4894042909145355,
+ "learning_rate": 0.00027140872562641226,
+ "loss": 0.615659523010254,
+ "mean_token_accuracy": 0.8227636247873307,
+ "num_tokens": 1332852.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.6712110449373722,
+ "epoch": 1.5588044184535412,
+ "grad_norm": 0.9492104649543762,
+ "learning_rate": 0.00027035395773182937,
+ "loss": 0.595902099609375,
+ "mean_token_accuracy": 0.8245656326413154,
+ "num_tokens": 1453569.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.6576551312208175,
+ "epoch": 1.6887589343729694,
+ "grad_norm": 0.6480705142021179,
+ "learning_rate": 0.00026902406558930333,
+ "loss": 0.5864276885986328,
+ "mean_token_accuracy": 0.8273968940973282,
+ "num_tokens": 1577152.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.6850327280163765,
+ "epoch": 1.8187134502923976,
+ "grad_norm": 0.516665518283844,
+ "learning_rate": 0.0002674217817941425,
+ "loss": 0.6083594512939453,
+ "mean_token_accuracy": 0.8232863625884056,
+ "num_tokens": 1691926.0,
+ "step": 700
+ },
+ {
+ "entropy": 0.6665814685821533,
+ "epoch": 1.9486679662118258,
+ "grad_norm": 0.6762926578521729,
+ "learning_rate": 0.00026555039863828637,
+ "loss": 0.6011666488647461,
+ "mean_token_accuracy": 0.8256330332159996,
+ "num_tokens": 1816059.0,
+ "step": 750
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6573245783264821,
+ "eval_loss": 0.6852125525474548,
+ "eval_mean_token_accuracy": 0.808370346060166,
+ "eval_num_tokens": 1862174.0,
+ "eval_runtime": 94.2081,
+ "eval_samples_per_second": 17.589,
+ "eval_steps_per_second": 2.208,
+ "step": 770
+ },
+ {
+ "entropy": 0.612116006301276,
+ "epoch": 2.077972709551657,
+ "grad_norm": 0.5827146172523499,
+ "learning_rate": 0.0002634137613454702,
+ "loss": 0.5394676208496094,
+ "mean_token_accuracy": 0.8394461966040146,
+ "num_tokens": 1931345.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.5671873143315316,
+ "epoch": 2.207927225471085,
+ "grad_norm": 0.523046612739563,
+ "learning_rate": 0.00026101626017025327,
+ "loss": 0.4936346435546875,
+ "mean_token_accuracy": 0.8486950224637986,
+ "num_tokens": 2057087.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.5720223182439804,
+ "epoch": 2.3378817413905133,
+ "grad_norm": 0.47807472944259644,
+ "learning_rate": 0.0002583628213771449,
+ "loss": 0.5041716003417969,
+ "mean_token_accuracy": 0.8464827784895896,
+ "num_tokens": 2179801.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.5660848580300808,
+ "epoch": 2.4678362573099415,
+ "grad_norm": 0.5875942707061768,
+ "learning_rate": 0.0002554588971183651,
+ "loss": 0.4948244094848633,
+ "mean_token_accuracy": 0.8494577088952064,
+ "num_tokens": 2300775.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.5680096058547497,
+ "epoch": 2.5977907732293697,
+ "grad_norm": 0.5312080383300781,
+ "learning_rate": 0.0002523104542310364,
+ "loss": 0.5015496826171875,
+ "mean_token_accuracy": 0.847268882393837,
+ "num_tokens": 2425909.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.5735848733782768,
+ "epoch": 2.727745289148798,
+ "grad_norm": 0.5523076057434082,
+ "learning_rate": 0.0002489239619768288,
+ "loss": 0.4966643524169922,
+ "mean_token_accuracy": 0.8478639185428619,
+ "num_tokens": 2542506.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.5666417135298252,
+ "epoch": 2.857699805068226,
+ "grad_norm": 0.4719190299510956,
+ "learning_rate": 0.00024530637874924604,
+ "loss": 0.5055577850341797,
+ "mean_token_accuracy": 0.8468799209594726,
+ "num_tokens": 2666261.0,
+ "step": 1100
+ },
+ {
+ "entropy": 0.5885502319037914,
+ "epoch": 2.9876543209876543,
+ "grad_norm": 0.5237556099891663,
+ "learning_rate": 0.00024146513777587047,
+ "loss": 0.5136178588867187,
+ "mean_token_accuracy": 0.845189718902111,
+ "num_tokens": 2781874.0,
+ "step": 1150
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.6404194602599511,
+ "eval_loss": 0.6698319315910339,
+ "eval_mean_token_accuracy": 0.8120474551732724,
+ "eval_num_tokens": 2793261.0,
+ "eval_runtime": 94.1854,
+ "eval_samples_per_second": 17.593,
+ "eval_steps_per_second": 2.208,
+ "step": 1155
+ },
+ {
+ "entropy": 0.48768596628203464,
+ "epoch": 3.116959064327485,
+ "grad_norm": 0.45523950457572937,
+ "learning_rate": 0.0002374081318449413,
+ "loss": 0.4034775924682617,
+ "mean_token_accuracy": 0.8715206759059848,
+ "num_tokens": 2896983.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.4783519075810909,
+ "epoch": 3.246913580246914,
+ "grad_norm": 0.5953325629234314,
+ "learning_rate": 0.0002331436970876517,
+ "loss": 0.39595497131347657,
+ "mean_token_accuracy": 0.8736060819029808,
+ "num_tokens": 3019005.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.4747548791766167,
+ "epoch": 3.3768680961663415,
+ "grad_norm": 0.5665897130966187,
+ "learning_rate": 0.00022868059584948654,
+ "loss": 0.4032122039794922,
+ "mean_token_accuracy": 0.8709388446807861,
+ "num_tokens": 3143415.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.4843991830945015,
+ "epoch": 3.50682261208577,
+ "grad_norm": 0.6089479923248291,
+ "learning_rate": 0.00022402799868579694,
+ "loss": 0.4088896179199219,
+ "mean_token_accuracy": 0.8691128557920456,
+ "num_tokens": 3263845.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.49369323417544364,
+ "epoch": 3.636777128005198,
+ "grad_norm": 0.7913607954978943,
+ "learning_rate": 0.00021919546551860548,
+ "loss": 0.4160056304931641,
+ "mean_token_accuracy": 0.8698329451680183,
+ "num_tokens": 3383117.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.47668287783861163,
+ "epoch": 3.7667316439246266,
+ "grad_norm": 0.5151591897010803,
+ "learning_rate": 0.00021419292599336078,
+ "loss": 0.4031526184082031,
+ "mean_token_accuracy": 0.8698683533072472,
+ "num_tokens": 3509724.0,
+ "step": 1450
+ },
+ {
+ "entropy": 0.4885013411939144,
+ "epoch": 3.8966861598440543,
+ "grad_norm": 0.6061839461326599,
+ "learning_rate": 0.0002090306590760024,
+ "loss": 0.4131336212158203,
+ "mean_token_accuracy": 0.8689957037568092,
+ "num_tokens": 3631110.0,
+ "step": 1500
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5588342215006168,
+ "eval_loss": 0.6875180006027222,
+ "eval_mean_token_accuracy": 0.8144597127460517,
+ "eval_num_tokens": 3724348.0,
+ "eval_runtime": 94.8129,
+ "eval_samples_per_second": 17.477,
+ "eval_steps_per_second": 2.194,
+ "step": 1540
+ },
+ {
+ "entropy": 0.4742255074594488,
+ "epoch": 4.025990903183885,
+ "grad_norm": 0.692275881767273,
+ "learning_rate": 0.0002037192719322604,
+ "loss": 0.39202346801757815,
+ "mean_token_accuracy": 0.8734938379508167,
+ "num_tokens": 3749929.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.3813336396217346,
+ "epoch": 4.155945419103314,
+ "grad_norm": 0.5534346699714661,
+ "learning_rate": 0.00019826967813258578,
+ "loss": 0.29004505157470706,
+ "mean_token_accuracy": 0.9025143319368363,
+ "num_tokens": 3871403.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.38276420064270494,
+ "epoch": 4.2858999350227425,
+ "grad_norm": 0.567842960357666,
+ "learning_rate": 0.00019269307522749648,
+ "loss": 0.2957651138305664,
+ "mean_token_accuracy": 0.9018667274713517,
+ "num_tokens": 3990410.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.38759261518716814,
+ "epoch": 4.41585445094217,
+ "grad_norm": 0.5272114276885986,
+ "learning_rate": 0.00018700092173941517,
+ "loss": 0.3017059135437012,
+ "mean_token_accuracy": 0.8991367623209954,
+ "num_tokens": 4107377.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.39710743524134157,
+ "epoch": 4.545808966861598,
+ "grad_norm": 0.5341858267784119,
+ "learning_rate": 0.00018120491361827535,
+ "loss": 0.30882442474365235,
+ "mean_token_accuracy": 0.8962600353360176,
+ "num_tokens": 4225238.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.39052032932639125,
+ "epoch": 4.675763482781027,
+ "grad_norm": 0.5387603044509888,
+ "learning_rate": 0.00017531696020927306,
+ "loss": 0.3076332473754883,
+ "mean_token_accuracy": 0.8972256830334664,
+ "num_tokens": 4348229.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.38603785961866377,
+ "epoch": 4.805717998700455,
+ "grad_norm": 0.6838284134864807,
+ "learning_rate": 0.00016934915978214512,
+ "loss": 0.3052788162231445,
+ "mean_token_accuracy": 0.8977296784520149,
+ "num_tokens": 4471337.0,
+ "step": 1850
+ },
+ {
+ "entropy": 0.387482148706913,
+ "epoch": 4.935672514619883,
+ "grad_norm": 0.6066106557846069,
+ "learning_rate": 0.00016331377467225454,
+ "loss": 0.3069518852233887,
+ "mean_token_accuracy": 0.8959452581405639,
+ "num_tokens": 4596516.0,
+ "step": 1900
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.47827344932235205,
+ "eval_loss": 0.734696626663208,
+ "eval_mean_token_accuracy": 0.8166056708074533,
+ "eval_num_tokens": 4655435.0,
+ "eval_runtime": 94.3316,
+ "eval_samples_per_second": 17.566,
+ "eval_steps_per_second": 2.205,
+ "step": 1925
+ },
+ {
+ "entropy": 0.3189891557298114,
+ "epoch": 5.064977257959714,
+ "grad_norm": 0.4411616623401642,
+ "learning_rate": 0.00015722320608456218,
+ "loss": 0.24322872161865233,
+ "mean_token_accuracy": 0.9192550799355435,
+ "num_tokens": 4720113.0,
+ "step": 1950
+ },
+ {
+ "entropy": 0.2661040405929089,
+ "epoch": 5.1949317738791425,
+ "grad_norm": 0.5544848442077637,
+ "learning_rate": 0.00015108996861225613,
+ "loss": 0.19045450210571288,
+ "mean_token_accuracy": 0.9353114965558053,
+ "num_tokens": 4842234.0,
+ "step": 2000
+ },
+ {
+ "entropy": 0.2827081228792667,
+ "epoch": 5.32488628979857,
+ "grad_norm": 0.5472713708877563,
+ "learning_rate": 0.0001449266645223969,
+ "loss": 0.197703857421875,
+ "mean_token_accuracy": 0.9314222651720047,
+ "num_tokens": 4959991.0,
+ "step": 2050
+ },
+ {
+ "entropy": 0.28358375541865827,
+ "epoch": 5.454840805717999,
+ "grad_norm": 0.6511954069137573,
+ "learning_rate": 0.00013874595786141455,
+ "loss": 0.20100860595703124,
+ "mean_token_accuracy": 0.930547822713852,
+ "num_tokens": 5082208.0,
+ "step": 2100
+ },
+ {
+ "entropy": 0.2757134060561657,
+ "epoch": 5.584795321637427,
+ "grad_norm": 0.5925538539886475,
+ "learning_rate": 0.0001325605484336644,
+ "loss": 0.19688623428344726,
+ "mean_token_accuracy": 0.9317859390377998,
+ "num_tokens": 5199699.0,
+ "step": 2150
+ },
+ {
+ "entropy": 0.2787610155344009,
+ "epoch": 5.714749837556855,
+ "grad_norm": 0.7486807703971863,
+ "learning_rate": 0.00012638314570651013,
+ "loss": 0.198111629486084,
+ "mean_token_accuracy": 0.9316030797362328,
+ "num_tokens": 5320501.0,
+ "step": 2200
+ },
+ {
+ "entropy": 0.2822082404047251,
+ "epoch": 5.844704353476283,
+ "grad_norm": 0.5981082320213318,
+ "learning_rate": 0.00012022644269555052,
+ "loss": 0.19833782196044922,
+ "mean_token_accuracy": 0.9301980781555176,
+ "num_tokens": 5442291.0,
+ "step": 2250
+ },
+ {
+ "entropy": 0.29100623600184916,
+ "epoch": 5.974658869395712,
+ "grad_norm": 0.6062291860580444,
+ "learning_rate": 0.00011410308988365154,
+ "loss": 0.20115615844726562,
+ "mean_token_accuracy": 0.9285238465666771,
+ "num_tokens": 5560531.0,
+ "step": 2300
+ },
+ {
+ "epoch": 6.0,
+ "eval_entropy": 0.419523782025163,
+ "eval_loss": 0.8228668570518494,
+ "eval_mean_token_accuracy": 0.8120436149720962,
+ "eval_num_tokens": 5586522.0,
+ "eval_runtime": 94.2282,
+ "eval_samples_per_second": 17.585,
+ "eval_steps_per_second": 2.207,
+ "step": 2310
+ },
+ {
+ "entropy": 0.22129053263059215,
+ "epoch": 6.1039636127355426,
+ "grad_norm": 0.5541817545890808,
+ "learning_rate": 0.000108025669227372,
+ "loss": 0.1330717086791992,
+ "mean_token_accuracy": 0.9534786897688056,
+ "num_tokens": 5682578.0,
+ "step": 2350
+ },
+ {
+ "entropy": 0.19409072685986758,
+ "epoch": 6.23391812865497,
+ "grad_norm": 0.31028518080711365,
+ "learning_rate": 0.00010200666830419353,
+ "loss": 0.11706629753112793,
+ "mean_token_accuracy": 0.9599499633908272,
+ "num_tokens": 5803409.0,
+ "step": 2400
+ },
+ {
+ "entropy": 0.18704994775354863,
+ "epoch": 6.363872644574399,
+ "grad_norm": 0.426945298910141,
+ "learning_rate": 9.605845465367642e-05,
+ "loss": 0.11503399848937988,
+ "mean_token_accuracy": 0.9598734942078591,
+ "num_tokens": 5927439.0,
+ "step": 2450
+ },
+ {
+ "entropy": 0.1894980014115572,
+ "epoch": 6.493827160493828,
+ "grad_norm": 0.5031415224075317,
+ "learning_rate": 9.019325036526285e-05,
+ "loss": 0.11405850410461425,
+ "mean_token_accuracy": 0.959916945695877,
+ "num_tokens": 6052633.0,
+ "step": 2500
+ },
+ {
+ "entropy": 0.2080751285329461,
+ "epoch": 6.623781676413255,
+ "grad_norm": 0.6324969530105591,
+ "learning_rate": 8.442310696494393e-05,
+ "loss": 0.1220788288116455,
+ "mean_token_accuracy": 0.9559539663791656,
+ "num_tokens": 6167503.0,
+ "step": 2550
+ },
+ {
+ "entropy": 0.20628964144736528,
+ "epoch": 6.753736192332683,
+ "grad_norm": 0.6685518622398376,
+ "learning_rate": 7.875988065239161e-05,
+ "loss": 0.12388327598571777,
+ "mean_token_accuracy": 0.956074104309082,
+ "num_tokens": 6284023.0,
+ "step": 2600
+ },
+ {
+ "entropy": 0.19158254452049733,
+ "epoch": 6.883690708252112,
+ "grad_norm": 0.5295413136482239,
+ "learning_rate": 7.321520793943799e-05,
+ "loss": 0.11677920341491699,
+ "mean_token_accuracy": 0.9589221009612083,
+ "num_tokens": 6408327.0,
+ "step": 2650
+ },
+ {
+ "epoch": 7.0,
+ "eval_entropy": 0.3437433656161794,
+ "eval_loss": 0.9526592493057251,
+ "eval_mean_token_accuracy": 0.8121905275262319,
+ "eval_num_tokens": 6517609.0,
+ "eval_runtime": 94.2057,
+ "eval_samples_per_second": 17.589,
+ "eval_steps_per_second": 2.208,
+ "step": 2695
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3850,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.8507895005249536e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3080/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3080/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3080/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3080/adapter_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3080/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..26742a540eca0befcc61f09a270940ded2abaca9
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3080/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 128,
+ "lora_bias": false,
+ "lora_dropout": 0.07386918480460901,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "k_proj",
+ "down_proj",
+ "q_proj",
+ "v_proj",
+ "up_proj",
+ "gate_proj",
+ "o_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3080/chat_template.jinja b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3080/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3080/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3080/tokenizer_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3080/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3080/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3080/trainer_state.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3080/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..22f695b27cf7667c29c335b3a2d9db7b0173261a
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3080/trainer_state.json
@@ -0,0 +1,732 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 8.0,
+ "eval_steps": 500,
+ "global_step": 3080,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5888759356737137,
+ "epoch": 0.1299545159194282,
+ "grad_norm": 1.1323111057281494,
+ "learning_rate": 3.473456711878286e-05,
+ "loss": 1.4242234802246094,
+ "mean_token_accuracy": 0.6861670060455799,
+ "num_tokens": 127115.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.8638093116879463,
+ "epoch": 0.2599090318388564,
+ "grad_norm": 0.7135503888130188,
+ "learning_rate": 7.017800295427557e-05,
+ "loss": 0.7731611633300781,
+ "mean_token_accuracy": 0.7900564879179001,
+ "num_tokens": 251796.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.8165761995315551,
+ "epoch": 0.3898635477582846,
+ "grad_norm": 0.5835468173027039,
+ "learning_rate": 0.0001056214387897683,
+ "loss": 0.7186955261230469,
+ "mean_token_accuracy": 0.7992442399263382,
+ "num_tokens": 366505.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.7586414662003517,
+ "epoch": 0.5198180636777128,
+ "grad_norm": 0.5721960663795471,
+ "learning_rate": 0.00014106487462526099,
+ "loss": 0.6786507415771484,
+ "mean_token_accuracy": 0.8069057819247246,
+ "num_tokens": 493150.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.7753578001260757,
+ "epoch": 0.649772579597141,
+ "grad_norm": 0.6396453380584717,
+ "learning_rate": 0.00017650831046075373,
+ "loss": 0.6813118743896485,
+ "mean_token_accuracy": 0.806209616959095,
+ "num_tokens": 607197.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.7591985522210598,
+ "epoch": 0.7797270955165692,
+ "grad_norm": 0.577964186668396,
+ "learning_rate": 0.00021195174629624644,
+ "loss": 0.67177490234375,
+ "mean_token_accuracy": 0.8090321889519692,
+ "num_tokens": 726579.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.731777046918869,
+ "epoch": 0.9096816114359974,
+ "grad_norm": 0.597808301448822,
+ "learning_rate": 0.00024739518213173913,
+ "loss": 0.6601863861083984,
+ "mean_token_accuracy": 0.8120808112621307,
+ "num_tokens": 848944.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.7157932419616443,
+ "eval_loss": 0.7318932414054871,
+ "eval_mean_token_accuracy": 0.7979064600972029,
+ "eval_num_tokens": 931087.0,
+ "eval_runtime": 94.7419,
+ "eval_samples_per_second": 17.49,
+ "eval_steps_per_second": 2.195,
+ "step": 385
+ },
+ {
+ "entropy": 0.7265488104005555,
+ "epoch": 1.0389863547758285,
+ "grad_norm": 0.7224313020706177,
+ "learning_rate": 0.00027290346308950203,
+ "loss": 0.6452214050292969,
+ "mean_token_accuracy": 0.8156503406002293,
+ "num_tokens": 968580.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.6715995308756828,
+ "epoch": 1.1689408706952567,
+ "grad_norm": 0.6290106177330017,
+ "learning_rate": 0.000272684789300892,
+ "loss": 0.6042086791992187,
+ "mean_token_accuracy": 0.8257312250137329,
+ "num_tokens": 1090002.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.6747516925632954,
+ "epoch": 1.2988953866146848,
+ "grad_norm": 0.66269850730896,
+ "learning_rate": 0.00027218620198917387,
+ "loss": 0.6056245803833008,
+ "mean_token_accuracy": 0.8240418726205826,
+ "num_tokens": 1212161.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.6872365647554397,
+ "epoch": 1.428849902534113,
+ "grad_norm": 0.4894042909145355,
+ "learning_rate": 0.00027140872562641226,
+ "loss": 0.615659523010254,
+ "mean_token_accuracy": 0.8227636247873307,
+ "num_tokens": 1332852.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.6712110449373722,
+ "epoch": 1.5588044184535412,
+ "grad_norm": 0.9492104649543762,
+ "learning_rate": 0.00027035395773182937,
+ "loss": 0.595902099609375,
+ "mean_token_accuracy": 0.8245656326413154,
+ "num_tokens": 1453569.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.6576551312208175,
+ "epoch": 1.6887589343729694,
+ "grad_norm": 0.6480705142021179,
+ "learning_rate": 0.00026902406558930333,
+ "loss": 0.5864276885986328,
+ "mean_token_accuracy": 0.8273968940973282,
+ "num_tokens": 1577152.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.6850327280163765,
+ "epoch": 1.8187134502923976,
+ "grad_norm": 0.516665518283844,
+ "learning_rate": 0.0002674217817941425,
+ "loss": 0.6083594512939453,
+ "mean_token_accuracy": 0.8232863625884056,
+ "num_tokens": 1691926.0,
+ "step": 700
+ },
+ {
+ "entropy": 0.6665814685821533,
+ "epoch": 1.9486679662118258,
+ "grad_norm": 0.6762926578521729,
+ "learning_rate": 0.00026555039863828637,
+ "loss": 0.6011666488647461,
+ "mean_token_accuracy": 0.8256330332159996,
+ "num_tokens": 1816059.0,
+ "step": 750
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6573245783264821,
+ "eval_loss": 0.6852125525474548,
+ "eval_mean_token_accuracy": 0.808370346060166,
+ "eval_num_tokens": 1862174.0,
+ "eval_runtime": 94.2081,
+ "eval_samples_per_second": 17.589,
+ "eval_steps_per_second": 2.208,
+ "step": 770
+ },
+ {
+ "entropy": 0.612116006301276,
+ "epoch": 2.077972709551657,
+ "grad_norm": 0.5827146172523499,
+ "learning_rate": 0.0002634137613454702,
+ "loss": 0.5394676208496094,
+ "mean_token_accuracy": 0.8394461966040146,
+ "num_tokens": 1931345.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.5671873143315316,
+ "epoch": 2.207927225471085,
+ "grad_norm": 0.523046612739563,
+ "learning_rate": 0.00026101626017025327,
+ "loss": 0.4936346435546875,
+ "mean_token_accuracy": 0.8486950224637986,
+ "num_tokens": 2057087.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.5720223182439804,
+ "epoch": 2.3378817413905133,
+ "grad_norm": 0.47807472944259644,
+ "learning_rate": 0.0002583628213771449,
+ "loss": 0.5041716003417969,
+ "mean_token_accuracy": 0.8464827784895896,
+ "num_tokens": 2179801.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.5660848580300808,
+ "epoch": 2.4678362573099415,
+ "grad_norm": 0.5875942707061768,
+ "learning_rate": 0.0002554588971183651,
+ "loss": 0.4948244094848633,
+ "mean_token_accuracy": 0.8494577088952064,
+ "num_tokens": 2300775.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.5680096058547497,
+ "epoch": 2.5977907732293697,
+ "grad_norm": 0.5312080383300781,
+ "learning_rate": 0.0002523104542310364,
+ "loss": 0.5015496826171875,
+ "mean_token_accuracy": 0.847268882393837,
+ "num_tokens": 2425909.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.5735848733782768,
+ "epoch": 2.727745289148798,
+ "grad_norm": 0.5523076057434082,
+ "learning_rate": 0.0002489239619768288,
+ "loss": 0.4966643524169922,
+ "mean_token_accuracy": 0.8478639185428619,
+ "num_tokens": 2542506.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.5666417135298252,
+ "epoch": 2.857699805068226,
+ "grad_norm": 0.4719190299510956,
+ "learning_rate": 0.00024530637874924604,
+ "loss": 0.5055577850341797,
+ "mean_token_accuracy": 0.8468799209594726,
+ "num_tokens": 2666261.0,
+ "step": 1100
+ },
+ {
+ "entropy": 0.5885502319037914,
+ "epoch": 2.9876543209876543,
+ "grad_norm": 0.5237556099891663,
+ "learning_rate": 0.00024146513777587047,
+ "loss": 0.5136178588867187,
+ "mean_token_accuracy": 0.845189718902111,
+ "num_tokens": 2781874.0,
+ "step": 1150
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.6404194602599511,
+ "eval_loss": 0.6698319315910339,
+ "eval_mean_token_accuracy": 0.8120474551732724,
+ "eval_num_tokens": 2793261.0,
+ "eval_runtime": 94.1854,
+ "eval_samples_per_second": 17.593,
+ "eval_steps_per_second": 2.208,
+ "step": 1155
+ },
+ {
+ "entropy": 0.48768596628203464,
+ "epoch": 3.116959064327485,
+ "grad_norm": 0.45523950457572937,
+ "learning_rate": 0.0002374081318449413,
+ "loss": 0.4034775924682617,
+ "mean_token_accuracy": 0.8715206759059848,
+ "num_tokens": 2896983.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.4783519075810909,
+ "epoch": 3.246913580246914,
+ "grad_norm": 0.5953325629234314,
+ "learning_rate": 0.0002331436970876517,
+ "loss": 0.39595497131347657,
+ "mean_token_accuracy": 0.8736060819029808,
+ "num_tokens": 3019005.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.4747548791766167,
+ "epoch": 3.3768680961663415,
+ "grad_norm": 0.5665897130966187,
+ "learning_rate": 0.00022868059584948654,
+ "loss": 0.4032122039794922,
+ "mean_token_accuracy": 0.8709388446807861,
+ "num_tokens": 3143415.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.4843991830945015,
+ "epoch": 3.50682261208577,
+ "grad_norm": 0.6089479923248291,
+ "learning_rate": 0.00022402799868579694,
+ "loss": 0.4088896179199219,
+ "mean_token_accuracy": 0.8691128557920456,
+ "num_tokens": 3263845.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.49369323417544364,
+ "epoch": 3.636777128005198,
+ "grad_norm": 0.7913607954978943,
+ "learning_rate": 0.00021919546551860548,
+ "loss": 0.4160056304931641,
+ "mean_token_accuracy": 0.8698329451680183,
+ "num_tokens": 3383117.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.47668287783861163,
+ "epoch": 3.7667316439246266,
+ "grad_norm": 0.5151591897010803,
+ "learning_rate": 0.00021419292599336078,
+ "loss": 0.4031526184082031,
+ "mean_token_accuracy": 0.8698683533072472,
+ "num_tokens": 3509724.0,
+ "step": 1450
+ },
+ {
+ "entropy": 0.4885013411939144,
+ "epoch": 3.8966861598440543,
+ "grad_norm": 0.6061839461326599,
+ "learning_rate": 0.0002090306590760024,
+ "loss": 0.4131336212158203,
+ "mean_token_accuracy": 0.8689957037568092,
+ "num_tokens": 3631110.0,
+ "step": 1500
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5588342215006168,
+ "eval_loss": 0.6875180006027222,
+ "eval_mean_token_accuracy": 0.8144597127460517,
+ "eval_num_tokens": 3724348.0,
+ "eval_runtime": 94.8129,
+ "eval_samples_per_second": 17.477,
+ "eval_steps_per_second": 2.194,
+ "step": 1540
+ },
+ {
+ "entropy": 0.4742255074594488,
+ "epoch": 4.025990903183885,
+ "grad_norm": 0.692275881767273,
+ "learning_rate": 0.0002037192719322604,
+ "loss": 0.39202346801757815,
+ "mean_token_accuracy": 0.8734938379508167,
+ "num_tokens": 3749929.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.3813336396217346,
+ "epoch": 4.155945419103314,
+ "grad_norm": 0.5534346699714661,
+ "learning_rate": 0.00019826967813258578,
+ "loss": 0.29004505157470706,
+ "mean_token_accuracy": 0.9025143319368363,
+ "num_tokens": 3871403.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.38276420064270494,
+ "epoch": 4.2858999350227425,
+ "grad_norm": 0.567842960357666,
+ "learning_rate": 0.00019269307522749648,
+ "loss": 0.2957651138305664,
+ "mean_token_accuracy": 0.9018667274713517,
+ "num_tokens": 3990410.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.38759261518716814,
+ "epoch": 4.41585445094217,
+ "grad_norm": 0.5272114276885986,
+ "learning_rate": 0.00018700092173941517,
+ "loss": 0.3017059135437012,
+ "mean_token_accuracy": 0.8991367623209954,
+ "num_tokens": 4107377.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.39710743524134157,
+ "epoch": 4.545808966861598,
+ "grad_norm": 0.5341858267784119,
+ "learning_rate": 0.00018120491361827535,
+ "loss": 0.30882442474365235,
+ "mean_token_accuracy": 0.8962600353360176,
+ "num_tokens": 4225238.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.39052032932639125,
+ "epoch": 4.675763482781027,
+ "grad_norm": 0.5387603044509888,
+ "learning_rate": 0.00017531696020927306,
+ "loss": 0.3076332473754883,
+ "mean_token_accuracy": 0.8972256830334664,
+ "num_tokens": 4348229.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.38603785961866377,
+ "epoch": 4.805717998700455,
+ "grad_norm": 0.6838284134864807,
+ "learning_rate": 0.00016934915978214512,
+ "loss": 0.3052788162231445,
+ "mean_token_accuracy": 0.8977296784520149,
+ "num_tokens": 4471337.0,
+ "step": 1850
+ },
+ {
+ "entropy": 0.387482148706913,
+ "epoch": 4.935672514619883,
+ "grad_norm": 0.6066106557846069,
+ "learning_rate": 0.00016331377467225454,
+ "loss": 0.3069518852233887,
+ "mean_token_accuracy": 0.8959452581405639,
+ "num_tokens": 4596516.0,
+ "step": 1900
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.47827344932235205,
+ "eval_loss": 0.734696626663208,
+ "eval_mean_token_accuracy": 0.8166056708074533,
+ "eval_num_tokens": 4655435.0,
+ "eval_runtime": 94.3316,
+ "eval_samples_per_second": 17.566,
+ "eval_steps_per_second": 2.205,
+ "step": 1925
+ },
+ {
+ "entropy": 0.3189891557298114,
+ "epoch": 5.064977257959714,
+ "grad_norm": 0.4411616623401642,
+ "learning_rate": 0.00015722320608456218,
+ "loss": 0.24322872161865233,
+ "mean_token_accuracy": 0.9192550799355435,
+ "num_tokens": 4720113.0,
+ "step": 1950
+ },
+ {
+ "entropy": 0.2661040405929089,
+ "epoch": 5.1949317738791425,
+ "grad_norm": 0.5544848442077637,
+ "learning_rate": 0.00015108996861225613,
+ "loss": 0.19045450210571288,
+ "mean_token_accuracy": 0.9353114965558053,
+ "num_tokens": 4842234.0,
+ "step": 2000
+ },
+ {
+ "entropy": 0.2827081228792667,
+ "epoch": 5.32488628979857,
+ "grad_norm": 0.5472713708877563,
+ "learning_rate": 0.0001449266645223969,
+ "loss": 0.197703857421875,
+ "mean_token_accuracy": 0.9314222651720047,
+ "num_tokens": 4959991.0,
+ "step": 2050
+ },
+ {
+ "entropy": 0.28358375541865827,
+ "epoch": 5.454840805717999,
+ "grad_norm": 0.6511954069137573,
+ "learning_rate": 0.00013874595786141455,
+ "loss": 0.20100860595703124,
+ "mean_token_accuracy": 0.930547822713852,
+ "num_tokens": 5082208.0,
+ "step": 2100
+ },
+ {
+ "entropy": 0.2757134060561657,
+ "epoch": 5.584795321637427,
+ "grad_norm": 0.5925538539886475,
+ "learning_rate": 0.0001325605484336644,
+ "loss": 0.19688623428344726,
+ "mean_token_accuracy": 0.9317859390377998,
+ "num_tokens": 5199699.0,
+ "step": 2150
+ },
+ {
+ "entropy": 0.2787610155344009,
+ "epoch": 5.714749837556855,
+ "grad_norm": 0.7486807703971863,
+ "learning_rate": 0.00012638314570651013,
+ "loss": 0.198111629486084,
+ "mean_token_accuracy": 0.9316030797362328,
+ "num_tokens": 5320501.0,
+ "step": 2200
+ },
+ {
+ "entropy": 0.2822082404047251,
+ "epoch": 5.844704353476283,
+ "grad_norm": 0.5981082320213318,
+ "learning_rate": 0.00012022644269555052,
+ "loss": 0.19833782196044922,
+ "mean_token_accuracy": 0.9301980781555176,
+ "num_tokens": 5442291.0,
+ "step": 2250
+ },
+ {
+ "entropy": 0.29100623600184916,
+ "epoch": 5.974658869395712,
+ "grad_norm": 0.6062291860580444,
+ "learning_rate": 0.00011410308988365154,
+ "loss": 0.20115615844726562,
+ "mean_token_accuracy": 0.9285238465666771,
+ "num_tokens": 5560531.0,
+ "step": 2300
+ },
+ {
+ "epoch": 6.0,
+ "eval_entropy": 0.419523782025163,
+ "eval_loss": 0.8228668570518494,
+ "eval_mean_token_accuracy": 0.8120436149720962,
+ "eval_num_tokens": 5586522.0,
+ "eval_runtime": 94.2282,
+ "eval_samples_per_second": 17.585,
+ "eval_steps_per_second": 2.207,
+ "step": 2310
+ },
+ {
+ "entropy": 0.22129053263059215,
+ "epoch": 6.1039636127355426,
+ "grad_norm": 0.5541817545890808,
+ "learning_rate": 0.000108025669227372,
+ "loss": 0.1330717086791992,
+ "mean_token_accuracy": 0.9534786897688056,
+ "num_tokens": 5682578.0,
+ "step": 2350
+ },
+ {
+ "entropy": 0.19409072685986758,
+ "epoch": 6.23391812865497,
+ "grad_norm": 0.31028518080711365,
+ "learning_rate": 0.00010200666830419353,
+ "loss": 0.11706629753112793,
+ "mean_token_accuracy": 0.9599499633908272,
+ "num_tokens": 5803409.0,
+ "step": 2400
+ },
+ {
+ "entropy": 0.18704994775354863,
+ "epoch": 6.363872644574399,
+ "grad_norm": 0.426945298910141,
+ "learning_rate": 9.605845465367642e-05,
+ "loss": 0.11503399848937988,
+ "mean_token_accuracy": 0.9598734942078591,
+ "num_tokens": 5927439.0,
+ "step": 2450
+ },
+ {
+ "entropy": 0.1894980014115572,
+ "epoch": 6.493827160493828,
+ "grad_norm": 0.5031415224075317,
+ "learning_rate": 9.019325036526285e-05,
+ "loss": 0.11405850410461425,
+ "mean_token_accuracy": 0.959916945695877,
+ "num_tokens": 6052633.0,
+ "step": 2500
+ },
+ {
+ "entropy": 0.2080751285329461,
+ "epoch": 6.623781676413255,
+ "grad_norm": 0.6324969530105591,
+ "learning_rate": 8.442310696494393e-05,
+ "loss": 0.1220788288116455,
+ "mean_token_accuracy": 0.9559539663791656,
+ "num_tokens": 6167503.0,
+ "step": 2550
+ },
+ {
+ "entropy": 0.20628964144736528,
+ "epoch": 6.753736192332683,
+ "grad_norm": 0.6685518622398376,
+ "learning_rate": 7.875988065239161e-05,
+ "loss": 0.12388327598571777,
+ "mean_token_accuracy": 0.956074104309082,
+ "num_tokens": 6284023.0,
+ "step": 2600
+ },
+ {
+ "entropy": 0.19158254452049733,
+ "epoch": 6.883690708252112,
+ "grad_norm": 0.5295413136482239,
+ "learning_rate": 7.321520793943799e-05,
+ "loss": 0.11677920341491699,
+ "mean_token_accuracy": 0.9589221009612083,
+ "num_tokens": 6408327.0,
+ "step": 2650
+ },
+ {
+ "epoch": 7.0,
+ "eval_entropy": 0.3437433656161794,
+ "eval_loss": 0.9526592493057251,
+ "eval_mean_token_accuracy": 0.8121905275262319,
+ "eval_num_tokens": 6517609.0,
+ "eval_runtime": 94.2057,
+ "eval_samples_per_second": 17.589,
+ "eval_steps_per_second": 2.208,
+ "step": 2695
+ },
+ {
+ "entropy": 0.18569647883949567,
+ "epoch": 7.012995451591943,
+ "grad_norm": 0.4049379825592041,
+ "learning_rate": 6.78004817399571e-05,
+ "loss": 0.10891223907470703,
+ "mean_token_accuracy": 0.9622611065006735,
+ "num_tokens": 6530430.0,
+ "step": 2700
+ },
+ {
+ "entropy": 0.14998033180832862,
+ "epoch": 7.142949967511371,
+ "grad_norm": 0.39533936977386475,
+ "learning_rate": 6.252682796028047e-05,
+ "loss": 0.07854948997497559,
+ "mean_token_accuracy": 0.9729882365465164,
+ "num_tokens": 6648623.0,
+ "step": 2750
+ },
+ {
+ "entropy": 0.1513198823109269,
+ "epoch": 7.272904483430799,
+ "grad_norm": 0.40231335163116455,
+ "learning_rate": 5.740508263824541e-05,
+ "loss": 0.07896852493286133,
+ "mean_token_accuracy": 0.9724258282780647,
+ "num_tokens": 6768614.0,
+ "step": 2800
+ },
+ {
+ "entropy": 0.1426772189885378,
+ "epoch": 7.402858999350228,
+ "grad_norm": 0.39278867840766907,
+ "learning_rate": 5.244576967785141e-05,
+ "loss": 0.07671234607696534,
+ "mean_token_accuracy": 0.9739191561937333,
+ "num_tokens": 6892420.0,
+ "step": 2850
+ },
+ {
+ "entropy": 0.15510446103289724,
+ "epoch": 7.532813515269655,
+ "grad_norm": 0.45904669165611267,
+ "learning_rate": 4.7659079225272575e-05,
+ "loss": 0.08021929740905762,
+ "mean_token_accuracy": 0.9704085075855255,
+ "num_tokens": 7009857.0,
+ "step": 2900
+ },
+ {
+ "entropy": 0.14695645896717907,
+ "epoch": 7.662768031189084,
+ "grad_norm": 0.2910745441913605,
+ "learning_rate": 4.305484673065923e-05,
+ "loss": 0.07789832592010498,
+ "mean_token_accuracy": 0.9713015568256378,
+ "num_tokens": 7132184.0,
+ "step": 2950
+ },
+ {
+ "entropy": 0.15163813527673484,
+ "epoch": 7.792722547108512,
+ "grad_norm": 0.26081109046936035,
+ "learning_rate": 3.8642532738751e-05,
+ "loss": 0.07821487426757813,
+ "mean_token_accuracy": 0.9711482360959053,
+ "num_tokens": 7250668.0,
+ "step": 3000
+ },
+ {
+ "entropy": 0.15378343284130097,
+ "epoch": 7.92267706302794,
+ "grad_norm": 0.35973286628723145,
+ "learning_rate": 3.443120344982667e-05,
+ "loss": 0.07847892284393311,
+ "mean_token_accuracy": 0.9707216301560402,
+ "num_tokens": 7373911.0,
+ "step": 3050
+ },
+ {
+ "epoch": 8.0,
+ "eval_entropy": 0.30457046585014236,
+ "eval_loss": 1.0573656558990479,
+ "eval_mean_token_accuracy": 0.8123422947067481,
+ "eval_num_tokens": 7448696.0,
+ "eval_runtime": 94.7627,
+ "eval_samples_per_second": 17.486,
+ "eval_steps_per_second": 2.195,
+ "step": 3080
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3850,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 3.257258894436741e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3465/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3465/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3465/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3465/adapter_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3465/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..26742a540eca0befcc61f09a270940ded2abaca9
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3465/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 128,
+ "lora_bias": false,
+ "lora_dropout": 0.07386918480460901,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "k_proj",
+ "down_proj",
+ "q_proj",
+ "v_proj",
+ "up_proj",
+ "gate_proj",
+ "o_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3465/chat_template.jinja b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3465/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3465/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3465/tokenizer_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3465/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3465/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3465/trainer_state.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3465/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..de2d84cc52d9eabc6c0e63ec1249964cdeddc010
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3465/trainer_state.json
@@ -0,0 +1,823 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 9.0,
+ "eval_steps": 500,
+ "global_step": 3465,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5888759356737137,
+ "epoch": 0.1299545159194282,
+ "grad_norm": 1.1323111057281494,
+ "learning_rate": 3.473456711878286e-05,
+ "loss": 1.4242234802246094,
+ "mean_token_accuracy": 0.6861670060455799,
+ "num_tokens": 127115.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.8638093116879463,
+ "epoch": 0.2599090318388564,
+ "grad_norm": 0.7135503888130188,
+ "learning_rate": 7.017800295427557e-05,
+ "loss": 0.7731611633300781,
+ "mean_token_accuracy": 0.7900564879179001,
+ "num_tokens": 251796.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.8165761995315551,
+ "epoch": 0.3898635477582846,
+ "grad_norm": 0.5835468173027039,
+ "learning_rate": 0.0001056214387897683,
+ "loss": 0.7186955261230469,
+ "mean_token_accuracy": 0.7992442399263382,
+ "num_tokens": 366505.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.7586414662003517,
+ "epoch": 0.5198180636777128,
+ "grad_norm": 0.5721960663795471,
+ "learning_rate": 0.00014106487462526099,
+ "loss": 0.6786507415771484,
+ "mean_token_accuracy": 0.8069057819247246,
+ "num_tokens": 493150.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.7753578001260757,
+ "epoch": 0.649772579597141,
+ "grad_norm": 0.6396453380584717,
+ "learning_rate": 0.00017650831046075373,
+ "loss": 0.6813118743896485,
+ "mean_token_accuracy": 0.806209616959095,
+ "num_tokens": 607197.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.7591985522210598,
+ "epoch": 0.7797270955165692,
+ "grad_norm": 0.577964186668396,
+ "learning_rate": 0.00021195174629624644,
+ "loss": 0.67177490234375,
+ "mean_token_accuracy": 0.8090321889519692,
+ "num_tokens": 726579.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.731777046918869,
+ "epoch": 0.9096816114359974,
+ "grad_norm": 0.597808301448822,
+ "learning_rate": 0.00024739518213173913,
+ "loss": 0.6601863861083984,
+ "mean_token_accuracy": 0.8120808112621307,
+ "num_tokens": 848944.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.7157932419616443,
+ "eval_loss": 0.7318932414054871,
+ "eval_mean_token_accuracy": 0.7979064600972029,
+ "eval_num_tokens": 931087.0,
+ "eval_runtime": 94.7419,
+ "eval_samples_per_second": 17.49,
+ "eval_steps_per_second": 2.195,
+ "step": 385
+ },
+ {
+ "entropy": 0.7265488104005555,
+ "epoch": 1.0389863547758285,
+ "grad_norm": 0.7224313020706177,
+ "learning_rate": 0.00027290346308950203,
+ "loss": 0.6452214050292969,
+ "mean_token_accuracy": 0.8156503406002293,
+ "num_tokens": 968580.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.6715995308756828,
+ "epoch": 1.1689408706952567,
+ "grad_norm": 0.6290106177330017,
+ "learning_rate": 0.000272684789300892,
+ "loss": 0.6042086791992187,
+ "mean_token_accuracy": 0.8257312250137329,
+ "num_tokens": 1090002.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.6747516925632954,
+ "epoch": 1.2988953866146848,
+ "grad_norm": 0.66269850730896,
+ "learning_rate": 0.00027218620198917387,
+ "loss": 0.6056245803833008,
+ "mean_token_accuracy": 0.8240418726205826,
+ "num_tokens": 1212161.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.6872365647554397,
+ "epoch": 1.428849902534113,
+ "grad_norm": 0.4894042909145355,
+ "learning_rate": 0.00027140872562641226,
+ "loss": 0.615659523010254,
+ "mean_token_accuracy": 0.8227636247873307,
+ "num_tokens": 1332852.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.6712110449373722,
+ "epoch": 1.5588044184535412,
+ "grad_norm": 0.9492104649543762,
+ "learning_rate": 0.00027035395773182937,
+ "loss": 0.595902099609375,
+ "mean_token_accuracy": 0.8245656326413154,
+ "num_tokens": 1453569.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.6576551312208175,
+ "epoch": 1.6887589343729694,
+ "grad_norm": 0.6480705142021179,
+ "learning_rate": 0.00026902406558930333,
+ "loss": 0.5864276885986328,
+ "mean_token_accuracy": 0.8273968940973282,
+ "num_tokens": 1577152.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.6850327280163765,
+ "epoch": 1.8187134502923976,
+ "grad_norm": 0.516665518283844,
+ "learning_rate": 0.0002674217817941425,
+ "loss": 0.6083594512939453,
+ "mean_token_accuracy": 0.8232863625884056,
+ "num_tokens": 1691926.0,
+ "step": 700
+ },
+ {
+ "entropy": 0.6665814685821533,
+ "epoch": 1.9486679662118258,
+ "grad_norm": 0.6762926578521729,
+ "learning_rate": 0.00026555039863828637,
+ "loss": 0.6011666488647461,
+ "mean_token_accuracy": 0.8256330332159996,
+ "num_tokens": 1816059.0,
+ "step": 750
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6573245783264821,
+ "eval_loss": 0.6852125525474548,
+ "eval_mean_token_accuracy": 0.808370346060166,
+ "eval_num_tokens": 1862174.0,
+ "eval_runtime": 94.2081,
+ "eval_samples_per_second": 17.589,
+ "eval_steps_per_second": 2.208,
+ "step": 770
+ },
+ {
+ "entropy": 0.612116006301276,
+ "epoch": 2.077972709551657,
+ "grad_norm": 0.5827146172523499,
+ "learning_rate": 0.0002634137613454702,
+ "loss": 0.5394676208496094,
+ "mean_token_accuracy": 0.8394461966040146,
+ "num_tokens": 1931345.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.5671873143315316,
+ "epoch": 2.207927225471085,
+ "grad_norm": 0.523046612739563,
+ "learning_rate": 0.00026101626017025327,
+ "loss": 0.4936346435546875,
+ "mean_token_accuracy": 0.8486950224637986,
+ "num_tokens": 2057087.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.5720223182439804,
+ "epoch": 2.3378817413905133,
+ "grad_norm": 0.47807472944259644,
+ "learning_rate": 0.0002583628213771449,
+ "loss": 0.5041716003417969,
+ "mean_token_accuracy": 0.8464827784895896,
+ "num_tokens": 2179801.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.5660848580300808,
+ "epoch": 2.4678362573099415,
+ "grad_norm": 0.5875942707061768,
+ "learning_rate": 0.0002554588971183651,
+ "loss": 0.4948244094848633,
+ "mean_token_accuracy": 0.8494577088952064,
+ "num_tokens": 2300775.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.5680096058547497,
+ "epoch": 2.5977907732293697,
+ "grad_norm": 0.5312080383300781,
+ "learning_rate": 0.0002523104542310364,
+ "loss": 0.5015496826171875,
+ "mean_token_accuracy": 0.847268882393837,
+ "num_tokens": 2425909.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.5735848733782768,
+ "epoch": 2.727745289148798,
+ "grad_norm": 0.5523076057434082,
+ "learning_rate": 0.0002489239619768288,
+ "loss": 0.4966643524169922,
+ "mean_token_accuracy": 0.8478639185428619,
+ "num_tokens": 2542506.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.5666417135298252,
+ "epoch": 2.857699805068226,
+ "grad_norm": 0.4719190299510956,
+ "learning_rate": 0.00024530637874924604,
+ "loss": 0.5055577850341797,
+ "mean_token_accuracy": 0.8468799209594726,
+ "num_tokens": 2666261.0,
+ "step": 1100
+ },
+ {
+ "entropy": 0.5885502319037914,
+ "epoch": 2.9876543209876543,
+ "grad_norm": 0.5237556099891663,
+ "learning_rate": 0.00024146513777587047,
+ "loss": 0.5136178588867187,
+ "mean_token_accuracy": 0.845189718902111,
+ "num_tokens": 2781874.0,
+ "step": 1150
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.6404194602599511,
+ "eval_loss": 0.6698319315910339,
+ "eval_mean_token_accuracy": 0.8120474551732724,
+ "eval_num_tokens": 2793261.0,
+ "eval_runtime": 94.1854,
+ "eval_samples_per_second": 17.593,
+ "eval_steps_per_second": 2.208,
+ "step": 1155
+ },
+ {
+ "entropy": 0.48768596628203464,
+ "epoch": 3.116959064327485,
+ "grad_norm": 0.45523950457572937,
+ "learning_rate": 0.0002374081318449413,
+ "loss": 0.4034775924682617,
+ "mean_token_accuracy": 0.8715206759059848,
+ "num_tokens": 2896983.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.4783519075810909,
+ "epoch": 3.246913580246914,
+ "grad_norm": 0.5953325629234314,
+ "learning_rate": 0.0002331436970876517,
+ "loss": 0.39595497131347657,
+ "mean_token_accuracy": 0.8736060819029808,
+ "num_tokens": 3019005.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.4747548791766167,
+ "epoch": 3.3768680961663415,
+ "grad_norm": 0.5665897130966187,
+ "learning_rate": 0.00022868059584948654,
+ "loss": 0.4032122039794922,
+ "mean_token_accuracy": 0.8709388446807861,
+ "num_tokens": 3143415.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.4843991830945015,
+ "epoch": 3.50682261208577,
+ "grad_norm": 0.6089479923248291,
+ "learning_rate": 0.00022402799868579694,
+ "loss": 0.4088896179199219,
+ "mean_token_accuracy": 0.8691128557920456,
+ "num_tokens": 3263845.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.49369323417544364,
+ "epoch": 3.636777128005198,
+ "grad_norm": 0.7913607954978943,
+ "learning_rate": 0.00021919546551860548,
+ "loss": 0.4160056304931641,
+ "mean_token_accuracy": 0.8698329451680183,
+ "num_tokens": 3383117.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.47668287783861163,
+ "epoch": 3.7667316439246266,
+ "grad_norm": 0.5151591897010803,
+ "learning_rate": 0.00021419292599336078,
+ "loss": 0.4031526184082031,
+ "mean_token_accuracy": 0.8698683533072472,
+ "num_tokens": 3509724.0,
+ "step": 1450
+ },
+ {
+ "entropy": 0.4885013411939144,
+ "epoch": 3.8966861598440543,
+ "grad_norm": 0.6061839461326599,
+ "learning_rate": 0.0002090306590760024,
+ "loss": 0.4131336212158203,
+ "mean_token_accuracy": 0.8689957037568092,
+ "num_tokens": 3631110.0,
+ "step": 1500
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5588342215006168,
+ "eval_loss": 0.6875180006027222,
+ "eval_mean_token_accuracy": 0.8144597127460517,
+ "eval_num_tokens": 3724348.0,
+ "eval_runtime": 94.8129,
+ "eval_samples_per_second": 17.477,
+ "eval_steps_per_second": 2.194,
+ "step": 1540
+ },
+ {
+ "entropy": 0.4742255074594488,
+ "epoch": 4.025990903183885,
+ "grad_norm": 0.692275881767273,
+ "learning_rate": 0.0002037192719322604,
+ "loss": 0.39202346801757815,
+ "mean_token_accuracy": 0.8734938379508167,
+ "num_tokens": 3749929.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.3813336396217346,
+ "epoch": 4.155945419103314,
+ "grad_norm": 0.5534346699714661,
+ "learning_rate": 0.00019826967813258578,
+ "loss": 0.29004505157470706,
+ "mean_token_accuracy": 0.9025143319368363,
+ "num_tokens": 3871403.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.38276420064270494,
+ "epoch": 4.2858999350227425,
+ "grad_norm": 0.567842960357666,
+ "learning_rate": 0.00019269307522749648,
+ "loss": 0.2957651138305664,
+ "mean_token_accuracy": 0.9018667274713517,
+ "num_tokens": 3990410.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.38759261518716814,
+ "epoch": 4.41585445094217,
+ "grad_norm": 0.5272114276885986,
+ "learning_rate": 0.00018700092173941517,
+ "loss": 0.3017059135437012,
+ "mean_token_accuracy": 0.8991367623209954,
+ "num_tokens": 4107377.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.39710743524134157,
+ "epoch": 4.545808966861598,
+ "grad_norm": 0.5341858267784119,
+ "learning_rate": 0.00018120491361827535,
+ "loss": 0.30882442474365235,
+ "mean_token_accuracy": 0.8962600353360176,
+ "num_tokens": 4225238.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.39052032932639125,
+ "epoch": 4.675763482781027,
+ "grad_norm": 0.5387603044509888,
+ "learning_rate": 0.00017531696020927306,
+ "loss": 0.3076332473754883,
+ "mean_token_accuracy": 0.8972256830334664,
+ "num_tokens": 4348229.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.38603785961866377,
+ "epoch": 4.805717998700455,
+ "grad_norm": 0.6838284134864807,
+ "learning_rate": 0.00016934915978214512,
+ "loss": 0.3052788162231445,
+ "mean_token_accuracy": 0.8977296784520149,
+ "num_tokens": 4471337.0,
+ "step": 1850
+ },
+ {
+ "entropy": 0.387482148706913,
+ "epoch": 4.935672514619883,
+ "grad_norm": 0.6066106557846069,
+ "learning_rate": 0.00016331377467225454,
+ "loss": 0.3069518852233887,
+ "mean_token_accuracy": 0.8959452581405639,
+ "num_tokens": 4596516.0,
+ "step": 1900
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.47827344932235205,
+ "eval_loss": 0.734696626663208,
+ "eval_mean_token_accuracy": 0.8166056708074533,
+ "eval_num_tokens": 4655435.0,
+ "eval_runtime": 94.3316,
+ "eval_samples_per_second": 17.566,
+ "eval_steps_per_second": 2.205,
+ "step": 1925
+ },
+ {
+ "entropy": 0.3189891557298114,
+ "epoch": 5.064977257959714,
+ "grad_norm": 0.4411616623401642,
+ "learning_rate": 0.00015722320608456218,
+ "loss": 0.24322872161865233,
+ "mean_token_accuracy": 0.9192550799355435,
+ "num_tokens": 4720113.0,
+ "step": 1950
+ },
+ {
+ "entropy": 0.2661040405929089,
+ "epoch": 5.1949317738791425,
+ "grad_norm": 0.5544848442077637,
+ "learning_rate": 0.00015108996861225613,
+ "loss": 0.19045450210571288,
+ "mean_token_accuracy": 0.9353114965558053,
+ "num_tokens": 4842234.0,
+ "step": 2000
+ },
+ {
+ "entropy": 0.2827081228792667,
+ "epoch": 5.32488628979857,
+ "grad_norm": 0.5472713708877563,
+ "learning_rate": 0.0001449266645223969,
+ "loss": 0.197703857421875,
+ "mean_token_accuracy": 0.9314222651720047,
+ "num_tokens": 4959991.0,
+ "step": 2050
+ },
+ {
+ "entropy": 0.28358375541865827,
+ "epoch": 5.454840805717999,
+ "grad_norm": 0.6511954069137573,
+ "learning_rate": 0.00013874595786141455,
+ "loss": 0.20100860595703124,
+ "mean_token_accuracy": 0.930547822713852,
+ "num_tokens": 5082208.0,
+ "step": 2100
+ },
+ {
+ "entropy": 0.2757134060561657,
+ "epoch": 5.584795321637427,
+ "grad_norm": 0.5925538539886475,
+ "learning_rate": 0.0001325605484336644,
+ "loss": 0.19688623428344726,
+ "mean_token_accuracy": 0.9317859390377998,
+ "num_tokens": 5199699.0,
+ "step": 2150
+ },
+ {
+ "entropy": 0.2787610155344009,
+ "epoch": 5.714749837556855,
+ "grad_norm": 0.7486807703971863,
+ "learning_rate": 0.00012638314570651013,
+ "loss": 0.198111629486084,
+ "mean_token_accuracy": 0.9316030797362328,
+ "num_tokens": 5320501.0,
+ "step": 2200
+ },
+ {
+ "entropy": 0.2822082404047251,
+ "epoch": 5.844704353476283,
+ "grad_norm": 0.5981082320213318,
+ "learning_rate": 0.00012022644269555052,
+ "loss": 0.19833782196044922,
+ "mean_token_accuracy": 0.9301980781555176,
+ "num_tokens": 5442291.0,
+ "step": 2250
+ },
+ {
+ "entropy": 0.29100623600184916,
+ "epoch": 5.974658869395712,
+ "grad_norm": 0.6062291860580444,
+ "learning_rate": 0.00011410308988365154,
+ "loss": 0.20115615844726562,
+ "mean_token_accuracy": 0.9285238465666771,
+ "num_tokens": 5560531.0,
+ "step": 2300
+ },
+ {
+ "epoch": 6.0,
+ "eval_entropy": 0.419523782025163,
+ "eval_loss": 0.8228668570518494,
+ "eval_mean_token_accuracy": 0.8120436149720962,
+ "eval_num_tokens": 5586522.0,
+ "eval_runtime": 94.2282,
+ "eval_samples_per_second": 17.585,
+ "eval_steps_per_second": 2.207,
+ "step": 2310
+ },
+ {
+ "entropy": 0.22129053263059215,
+ "epoch": 6.1039636127355426,
+ "grad_norm": 0.5541817545890808,
+ "learning_rate": 0.000108025669227372,
+ "loss": 0.1330717086791992,
+ "mean_token_accuracy": 0.9534786897688056,
+ "num_tokens": 5682578.0,
+ "step": 2350
+ },
+ {
+ "entropy": 0.19409072685986758,
+ "epoch": 6.23391812865497,
+ "grad_norm": 0.31028518080711365,
+ "learning_rate": 0.00010200666830419353,
+ "loss": 0.11706629753112793,
+ "mean_token_accuracy": 0.9599499633908272,
+ "num_tokens": 5803409.0,
+ "step": 2400
+ },
+ {
+ "entropy": 0.18704994775354863,
+ "epoch": 6.363872644574399,
+ "grad_norm": 0.426945298910141,
+ "learning_rate": 9.605845465367642e-05,
+ "loss": 0.11503399848937988,
+ "mean_token_accuracy": 0.9598734942078591,
+ "num_tokens": 5927439.0,
+ "step": 2450
+ },
+ {
+ "entropy": 0.1894980014115572,
+ "epoch": 6.493827160493828,
+ "grad_norm": 0.5031415224075317,
+ "learning_rate": 9.019325036526285e-05,
+ "loss": 0.11405850410461425,
+ "mean_token_accuracy": 0.959916945695877,
+ "num_tokens": 6052633.0,
+ "step": 2500
+ },
+ {
+ "entropy": 0.2080751285329461,
+ "epoch": 6.623781676413255,
+ "grad_norm": 0.6324969530105591,
+ "learning_rate": 8.442310696494393e-05,
+ "loss": 0.1220788288116455,
+ "mean_token_accuracy": 0.9559539663791656,
+ "num_tokens": 6167503.0,
+ "step": 2550
+ },
+ {
+ "entropy": 0.20628964144736528,
+ "epoch": 6.753736192332683,
+ "grad_norm": 0.6685518622398376,
+ "learning_rate": 7.875988065239161e-05,
+ "loss": 0.12388327598571777,
+ "mean_token_accuracy": 0.956074104309082,
+ "num_tokens": 6284023.0,
+ "step": 2600
+ },
+ {
+ "entropy": 0.19158254452049733,
+ "epoch": 6.883690708252112,
+ "grad_norm": 0.5295413136482239,
+ "learning_rate": 7.321520793943799e-05,
+ "loss": 0.11677920341491699,
+ "mean_token_accuracy": 0.9589221009612083,
+ "num_tokens": 6408327.0,
+ "step": 2650
+ },
+ {
+ "epoch": 7.0,
+ "eval_entropy": 0.3437433656161794,
+ "eval_loss": 0.9526592493057251,
+ "eval_mean_token_accuracy": 0.8121905275262319,
+ "eval_num_tokens": 6517609.0,
+ "eval_runtime": 94.2057,
+ "eval_samples_per_second": 17.589,
+ "eval_steps_per_second": 2.208,
+ "step": 2695
+ },
+ {
+ "entropy": 0.18569647883949567,
+ "epoch": 7.012995451591943,
+ "grad_norm": 0.4049379825592041,
+ "learning_rate": 6.78004817399571e-05,
+ "loss": 0.10891223907470703,
+ "mean_token_accuracy": 0.9622611065006735,
+ "num_tokens": 6530430.0,
+ "step": 2700
+ },
+ {
+ "entropy": 0.14998033180832862,
+ "epoch": 7.142949967511371,
+ "grad_norm": 0.39533936977386475,
+ "learning_rate": 6.252682796028047e-05,
+ "loss": 0.07854948997497559,
+ "mean_token_accuracy": 0.9729882365465164,
+ "num_tokens": 6648623.0,
+ "step": 2750
+ },
+ {
+ "entropy": 0.1513198823109269,
+ "epoch": 7.272904483430799,
+ "grad_norm": 0.40231335163116455,
+ "learning_rate": 5.740508263824541e-05,
+ "loss": 0.07896852493286133,
+ "mean_token_accuracy": 0.9724258282780647,
+ "num_tokens": 6768614.0,
+ "step": 2800
+ },
+ {
+ "entropy": 0.1426772189885378,
+ "epoch": 7.402858999350228,
+ "grad_norm": 0.39278867840766907,
+ "learning_rate": 5.244576967785141e-05,
+ "loss": 0.07671234607696534,
+ "mean_token_accuracy": 0.9739191561937333,
+ "num_tokens": 6892420.0,
+ "step": 2850
+ },
+ {
+ "entropy": 0.15510446103289724,
+ "epoch": 7.532813515269655,
+ "grad_norm": 0.45904669165611267,
+ "learning_rate": 4.7659079225272575e-05,
+ "loss": 0.08021929740905762,
+ "mean_token_accuracy": 0.9704085075855255,
+ "num_tokens": 7009857.0,
+ "step": 2900
+ },
+ {
+ "entropy": 0.14695645896717907,
+ "epoch": 7.662768031189084,
+ "grad_norm": 0.2910745441913605,
+ "learning_rate": 4.305484673065923e-05,
+ "loss": 0.07789832592010498,
+ "mean_token_accuracy": 0.9713015568256378,
+ "num_tokens": 7132184.0,
+ "step": 2950
+ },
+ {
+ "entropy": 0.15163813527673484,
+ "epoch": 7.792722547108512,
+ "grad_norm": 0.26081109046936035,
+ "learning_rate": 3.8642532738751e-05,
+ "loss": 0.07821487426757813,
+ "mean_token_accuracy": 0.9711482360959053,
+ "num_tokens": 7250668.0,
+ "step": 3000
+ },
+ {
+ "entropy": 0.15378343284130097,
+ "epoch": 7.92267706302794,
+ "grad_norm": 0.35973286628723145,
+ "learning_rate": 3.443120344982667e-05,
+ "loss": 0.07847892284393311,
+ "mean_token_accuracy": 0.9707216301560402,
+ "num_tokens": 7373911.0,
+ "step": 3050
+ },
+ {
+ "epoch": 8.0,
+ "eval_entropy": 0.30457046585014236,
+ "eval_loss": 1.0573656558990479,
+ "eval_mean_token_accuracy": 0.8123422947067481,
+ "eval_num_tokens": 7448696.0,
+ "eval_runtime": 94.7627,
+ "eval_samples_per_second": 17.486,
+ "eval_steps_per_second": 2.195,
+ "step": 3080
+ },
+ {
+ "entropy": 0.14268321996957214,
+ "epoch": 8.05198180636777,
+ "grad_norm": 0.17200376093387604,
+ "learning_rate": 3.0429512090932775e-05,
+ "loss": 0.06907873630523681,
+ "mean_token_accuracy": 0.9740184904941961,
+ "num_tokens": 7497025.0,
+ "step": 3100
+ },
+ {
+ "entropy": 0.13242442483082414,
+ "epoch": 8.1819363222872,
+ "grad_norm": 0.24965617060661316,
+ "learning_rate": 2.6645681135669307e-05,
+ "loss": 0.06381748676300049,
+ "mean_token_accuracy": 0.9758135312795639,
+ "num_tokens": 7615828.0,
+ "step": 3150
+ },
+ {
+ "entropy": 0.13472190242260695,
+ "epoch": 8.311890838206628,
+ "grad_norm": 0.24549734592437744,
+ "learning_rate": 2.3087485409065462e-05,
+ "loss": 0.06418798446655273,
+ "mean_token_accuracy": 0.9750474771857262,
+ "num_tokens": 7734498.0,
+ "step": 3200
+ },
+ {
+ "entropy": 0.12980990454554558,
+ "epoch": 8.441845354126055,
+ "grad_norm": 0.16283877193927765,
+ "learning_rate": 1.9762236112261417e-05,
+ "loss": 0.06304348945617676,
+ "mean_token_accuracy": 0.9762365907430649,
+ "num_tokens": 7855780.0,
+ "step": 3250
+ },
+ {
+ "entropy": 0.14281578866764902,
+ "epoch": 8.571799870045485,
+ "grad_norm": 0.25190046429634094,
+ "learning_rate": 1.667676579982073e-05,
+ "loss": 0.06881052494049072,
+ "mean_token_accuracy": 0.9739083594083786,
+ "num_tokens": 7968890.0,
+ "step": 3300
+ },
+ {
+ "entropy": 0.1275078719109297,
+ "epoch": 8.701754385964913,
+ "grad_norm": 0.14581522345542908,
+ "learning_rate": 1.3837414340541901e-05,
+ "loss": 0.06218691825866699,
+ "mean_token_accuracy": 0.9765720379352569,
+ "num_tokens": 8094439.0,
+ "step": 3350
+ },
+ {
+ "entropy": 0.12800191594287752,
+ "epoch": 8.83170890188434,
+ "grad_norm": 0.15930895507335663,
+ "learning_rate": 1.1250015890615666e-05,
+ "loss": 0.061544852256774904,
+ "mean_token_accuracy": 0.9760145637392997,
+ "num_tokens": 8219578.0,
+ "step": 3400
+ },
+ {
+ "entropy": 0.12469653191044927,
+ "epoch": 8.961663417803768,
+ "grad_norm": 0.3019782304763794,
+ "learning_rate": 8.91988690589514e-06,
+ "loss": 0.06210709571838379,
+ "mean_token_accuracy": 0.9761517831683159,
+ "num_tokens": 8345441.0,
+ "step": 3450
+ },
+ {
+ "epoch": 9.0,
+ "eval_entropy": 0.2802187033140889,
+ "eval_loss": 1.1696693897247314,
+ "eval_mean_token_accuracy": 0.8132876937205975,
+ "eval_num_tokens": 8379783.0,
+ "eval_runtime": 94.6664,
+ "eval_samples_per_second": 17.504,
+ "eval_steps_per_second": 2.197,
+ "step": 3465
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3850,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 3.662904308863918e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-385/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-385/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-385/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-385/adapter_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-385/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..26742a540eca0befcc61f09a270940ded2abaca9
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-385/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 128,
+ "lora_bias": false,
+ "lora_dropout": 0.07386918480460901,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "k_proj",
+ "down_proj",
+ "q_proj",
+ "v_proj",
+ "up_proj",
+ "gate_proj",
+ "o_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-385/chat_template.jinja b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-385/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-385/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-385/tokenizer_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-385/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-385/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-385/trainer_state.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-385/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..bbcbfc89c256035c475f2d3ebe49f7f435b4a618
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-385/trainer_state.json
@@ -0,0 +1,115 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 1.0,
+ "eval_steps": 500,
+ "global_step": 385,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5888759356737137,
+ "epoch": 0.1299545159194282,
+ "grad_norm": 1.1323111057281494,
+ "learning_rate": 3.473456711878286e-05,
+ "loss": 1.4242234802246094,
+ "mean_token_accuracy": 0.6861670060455799,
+ "num_tokens": 127115.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.8638093116879463,
+ "epoch": 0.2599090318388564,
+ "grad_norm": 0.7135503888130188,
+ "learning_rate": 7.017800295427557e-05,
+ "loss": 0.7731611633300781,
+ "mean_token_accuracy": 0.7900564879179001,
+ "num_tokens": 251796.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.8165761995315551,
+ "epoch": 0.3898635477582846,
+ "grad_norm": 0.5835468173027039,
+ "learning_rate": 0.0001056214387897683,
+ "loss": 0.7186955261230469,
+ "mean_token_accuracy": 0.7992442399263382,
+ "num_tokens": 366505.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.7586414662003517,
+ "epoch": 0.5198180636777128,
+ "grad_norm": 0.5721960663795471,
+ "learning_rate": 0.00014106487462526099,
+ "loss": 0.6786507415771484,
+ "mean_token_accuracy": 0.8069057819247246,
+ "num_tokens": 493150.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.7753578001260757,
+ "epoch": 0.649772579597141,
+ "grad_norm": 0.6396453380584717,
+ "learning_rate": 0.00017650831046075373,
+ "loss": 0.6813118743896485,
+ "mean_token_accuracy": 0.806209616959095,
+ "num_tokens": 607197.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.7591985522210598,
+ "epoch": 0.7797270955165692,
+ "grad_norm": 0.577964186668396,
+ "learning_rate": 0.00021195174629624644,
+ "loss": 0.67177490234375,
+ "mean_token_accuracy": 0.8090321889519692,
+ "num_tokens": 726579.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.731777046918869,
+ "epoch": 0.9096816114359974,
+ "grad_norm": 0.597808301448822,
+ "learning_rate": 0.00024739518213173913,
+ "loss": 0.6601863861083984,
+ "mean_token_accuracy": 0.8120808112621307,
+ "num_tokens": 848944.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.7157932419616443,
+ "eval_loss": 0.7318932414054871,
+ "eval_mean_token_accuracy": 0.7979064600972029,
+ "eval_num_tokens": 931087.0,
+ "eval_runtime": 94.7419,
+ "eval_samples_per_second": 17.49,
+ "eval_steps_per_second": 2.195,
+ "step": 385
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3850,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 4.054851961940582e+16,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3850/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3850/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3850/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3850/adapter_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3850/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..26742a540eca0befcc61f09a270940ded2abaca9
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3850/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 128,
+ "lora_bias": false,
+ "lora_dropout": 0.07386918480460901,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "k_proj",
+ "down_proj",
+ "q_proj",
+ "v_proj",
+ "up_proj",
+ "gate_proj",
+ "o_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3850/chat_template.jinja b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3850/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3850/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3850/tokenizer_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3850/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3850/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3850/trainer_state.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3850/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..15b95fdbf252d74535881c5259ee763e91fe2647
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3850/trainer_state.json
@@ -0,0 +1,914 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 10.0,
+ "eval_steps": 500,
+ "global_step": 3850,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5888759356737137,
+ "epoch": 0.1299545159194282,
+ "grad_norm": 1.1323111057281494,
+ "learning_rate": 3.473456711878286e-05,
+ "loss": 1.4242234802246094,
+ "mean_token_accuracy": 0.6861670060455799,
+ "num_tokens": 127115.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.8638093116879463,
+ "epoch": 0.2599090318388564,
+ "grad_norm": 0.7135503888130188,
+ "learning_rate": 7.017800295427557e-05,
+ "loss": 0.7731611633300781,
+ "mean_token_accuracy": 0.7900564879179001,
+ "num_tokens": 251796.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.8165761995315551,
+ "epoch": 0.3898635477582846,
+ "grad_norm": 0.5835468173027039,
+ "learning_rate": 0.0001056214387897683,
+ "loss": 0.7186955261230469,
+ "mean_token_accuracy": 0.7992442399263382,
+ "num_tokens": 366505.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.7586414662003517,
+ "epoch": 0.5198180636777128,
+ "grad_norm": 0.5721960663795471,
+ "learning_rate": 0.00014106487462526099,
+ "loss": 0.6786507415771484,
+ "mean_token_accuracy": 0.8069057819247246,
+ "num_tokens": 493150.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.7753578001260757,
+ "epoch": 0.649772579597141,
+ "grad_norm": 0.6396453380584717,
+ "learning_rate": 0.00017650831046075373,
+ "loss": 0.6813118743896485,
+ "mean_token_accuracy": 0.806209616959095,
+ "num_tokens": 607197.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.7591985522210598,
+ "epoch": 0.7797270955165692,
+ "grad_norm": 0.577964186668396,
+ "learning_rate": 0.00021195174629624644,
+ "loss": 0.67177490234375,
+ "mean_token_accuracy": 0.8090321889519692,
+ "num_tokens": 726579.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.731777046918869,
+ "epoch": 0.9096816114359974,
+ "grad_norm": 0.597808301448822,
+ "learning_rate": 0.00024739518213173913,
+ "loss": 0.6601863861083984,
+ "mean_token_accuracy": 0.8120808112621307,
+ "num_tokens": 848944.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.7157932419616443,
+ "eval_loss": 0.7318932414054871,
+ "eval_mean_token_accuracy": 0.7979064600972029,
+ "eval_num_tokens": 931087.0,
+ "eval_runtime": 94.7419,
+ "eval_samples_per_second": 17.49,
+ "eval_steps_per_second": 2.195,
+ "step": 385
+ },
+ {
+ "entropy": 0.7265488104005555,
+ "epoch": 1.0389863547758285,
+ "grad_norm": 0.7224313020706177,
+ "learning_rate": 0.00027290346308950203,
+ "loss": 0.6452214050292969,
+ "mean_token_accuracy": 0.8156503406002293,
+ "num_tokens": 968580.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.6715995308756828,
+ "epoch": 1.1689408706952567,
+ "grad_norm": 0.6290106177330017,
+ "learning_rate": 0.000272684789300892,
+ "loss": 0.6042086791992187,
+ "mean_token_accuracy": 0.8257312250137329,
+ "num_tokens": 1090002.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.6747516925632954,
+ "epoch": 1.2988953866146848,
+ "grad_norm": 0.66269850730896,
+ "learning_rate": 0.00027218620198917387,
+ "loss": 0.6056245803833008,
+ "mean_token_accuracy": 0.8240418726205826,
+ "num_tokens": 1212161.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.6872365647554397,
+ "epoch": 1.428849902534113,
+ "grad_norm": 0.4894042909145355,
+ "learning_rate": 0.00027140872562641226,
+ "loss": 0.615659523010254,
+ "mean_token_accuracy": 0.8227636247873307,
+ "num_tokens": 1332852.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.6712110449373722,
+ "epoch": 1.5588044184535412,
+ "grad_norm": 0.9492104649543762,
+ "learning_rate": 0.00027035395773182937,
+ "loss": 0.595902099609375,
+ "mean_token_accuracy": 0.8245656326413154,
+ "num_tokens": 1453569.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.6576551312208175,
+ "epoch": 1.6887589343729694,
+ "grad_norm": 0.6480705142021179,
+ "learning_rate": 0.00026902406558930333,
+ "loss": 0.5864276885986328,
+ "mean_token_accuracy": 0.8273968940973282,
+ "num_tokens": 1577152.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.6850327280163765,
+ "epoch": 1.8187134502923976,
+ "grad_norm": 0.516665518283844,
+ "learning_rate": 0.0002674217817941425,
+ "loss": 0.6083594512939453,
+ "mean_token_accuracy": 0.8232863625884056,
+ "num_tokens": 1691926.0,
+ "step": 700
+ },
+ {
+ "entropy": 0.6665814685821533,
+ "epoch": 1.9486679662118258,
+ "grad_norm": 0.6762926578521729,
+ "learning_rate": 0.00026555039863828637,
+ "loss": 0.6011666488647461,
+ "mean_token_accuracy": 0.8256330332159996,
+ "num_tokens": 1816059.0,
+ "step": 750
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6573245783264821,
+ "eval_loss": 0.6852125525474548,
+ "eval_mean_token_accuracy": 0.808370346060166,
+ "eval_num_tokens": 1862174.0,
+ "eval_runtime": 94.2081,
+ "eval_samples_per_second": 17.589,
+ "eval_steps_per_second": 2.208,
+ "step": 770
+ },
+ {
+ "entropy": 0.612116006301276,
+ "epoch": 2.077972709551657,
+ "grad_norm": 0.5827146172523499,
+ "learning_rate": 0.0002634137613454702,
+ "loss": 0.5394676208496094,
+ "mean_token_accuracy": 0.8394461966040146,
+ "num_tokens": 1931345.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.5671873143315316,
+ "epoch": 2.207927225471085,
+ "grad_norm": 0.523046612739563,
+ "learning_rate": 0.00026101626017025327,
+ "loss": 0.4936346435546875,
+ "mean_token_accuracy": 0.8486950224637986,
+ "num_tokens": 2057087.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.5720223182439804,
+ "epoch": 2.3378817413905133,
+ "grad_norm": 0.47807472944259644,
+ "learning_rate": 0.0002583628213771449,
+ "loss": 0.5041716003417969,
+ "mean_token_accuracy": 0.8464827784895896,
+ "num_tokens": 2179801.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.5660848580300808,
+ "epoch": 2.4678362573099415,
+ "grad_norm": 0.5875942707061768,
+ "learning_rate": 0.0002554588971183651,
+ "loss": 0.4948244094848633,
+ "mean_token_accuracy": 0.8494577088952064,
+ "num_tokens": 2300775.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.5680096058547497,
+ "epoch": 2.5977907732293697,
+ "grad_norm": 0.5312080383300781,
+ "learning_rate": 0.0002523104542310364,
+ "loss": 0.5015496826171875,
+ "mean_token_accuracy": 0.847268882393837,
+ "num_tokens": 2425909.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.5735848733782768,
+ "epoch": 2.727745289148798,
+ "grad_norm": 0.5523076057434082,
+ "learning_rate": 0.0002489239619768288,
+ "loss": 0.4966643524169922,
+ "mean_token_accuracy": 0.8478639185428619,
+ "num_tokens": 2542506.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.5666417135298252,
+ "epoch": 2.857699805068226,
+ "grad_norm": 0.4719190299510956,
+ "learning_rate": 0.00024530637874924604,
+ "loss": 0.5055577850341797,
+ "mean_token_accuracy": 0.8468799209594726,
+ "num_tokens": 2666261.0,
+ "step": 1100
+ },
+ {
+ "entropy": 0.5885502319037914,
+ "epoch": 2.9876543209876543,
+ "grad_norm": 0.5237556099891663,
+ "learning_rate": 0.00024146513777587047,
+ "loss": 0.5136178588867187,
+ "mean_token_accuracy": 0.845189718902111,
+ "num_tokens": 2781874.0,
+ "step": 1150
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.6404194602599511,
+ "eval_loss": 0.6698319315910339,
+ "eval_mean_token_accuracy": 0.8120474551732724,
+ "eval_num_tokens": 2793261.0,
+ "eval_runtime": 94.1854,
+ "eval_samples_per_second": 17.593,
+ "eval_steps_per_second": 2.208,
+ "step": 1155
+ },
+ {
+ "entropy": 0.48768596628203464,
+ "epoch": 3.116959064327485,
+ "grad_norm": 0.45523950457572937,
+ "learning_rate": 0.0002374081318449413,
+ "loss": 0.4034775924682617,
+ "mean_token_accuracy": 0.8715206759059848,
+ "num_tokens": 2896983.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.4783519075810909,
+ "epoch": 3.246913580246914,
+ "grad_norm": 0.5953325629234314,
+ "learning_rate": 0.0002331436970876517,
+ "loss": 0.39595497131347657,
+ "mean_token_accuracy": 0.8736060819029808,
+ "num_tokens": 3019005.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.4747548791766167,
+ "epoch": 3.3768680961663415,
+ "grad_norm": 0.5665897130966187,
+ "learning_rate": 0.00022868059584948654,
+ "loss": 0.4032122039794922,
+ "mean_token_accuracy": 0.8709388446807861,
+ "num_tokens": 3143415.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.4843991830945015,
+ "epoch": 3.50682261208577,
+ "grad_norm": 0.6089479923248291,
+ "learning_rate": 0.00022402799868579694,
+ "loss": 0.4088896179199219,
+ "mean_token_accuracy": 0.8691128557920456,
+ "num_tokens": 3263845.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.49369323417544364,
+ "epoch": 3.636777128005198,
+ "grad_norm": 0.7913607954978943,
+ "learning_rate": 0.00021919546551860548,
+ "loss": 0.4160056304931641,
+ "mean_token_accuracy": 0.8698329451680183,
+ "num_tokens": 3383117.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.47668287783861163,
+ "epoch": 3.7667316439246266,
+ "grad_norm": 0.5151591897010803,
+ "learning_rate": 0.00021419292599336078,
+ "loss": 0.4031526184082031,
+ "mean_token_accuracy": 0.8698683533072472,
+ "num_tokens": 3509724.0,
+ "step": 1450
+ },
+ {
+ "entropy": 0.4885013411939144,
+ "epoch": 3.8966861598440543,
+ "grad_norm": 0.6061839461326599,
+ "learning_rate": 0.0002090306590760024,
+ "loss": 0.4131336212158203,
+ "mean_token_accuracy": 0.8689957037568092,
+ "num_tokens": 3631110.0,
+ "step": 1500
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5588342215006168,
+ "eval_loss": 0.6875180006027222,
+ "eval_mean_token_accuracy": 0.8144597127460517,
+ "eval_num_tokens": 3724348.0,
+ "eval_runtime": 94.8129,
+ "eval_samples_per_second": 17.477,
+ "eval_steps_per_second": 2.194,
+ "step": 1540
+ },
+ {
+ "entropy": 0.4742255074594488,
+ "epoch": 4.025990903183885,
+ "grad_norm": 0.692275881767273,
+ "learning_rate": 0.0002037192719322604,
+ "loss": 0.39202346801757815,
+ "mean_token_accuracy": 0.8734938379508167,
+ "num_tokens": 3749929.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.3813336396217346,
+ "epoch": 4.155945419103314,
+ "grad_norm": 0.5534346699714661,
+ "learning_rate": 0.00019826967813258578,
+ "loss": 0.29004505157470706,
+ "mean_token_accuracy": 0.9025143319368363,
+ "num_tokens": 3871403.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.38276420064270494,
+ "epoch": 4.2858999350227425,
+ "grad_norm": 0.567842960357666,
+ "learning_rate": 0.00019269307522749648,
+ "loss": 0.2957651138305664,
+ "mean_token_accuracy": 0.9018667274713517,
+ "num_tokens": 3990410.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.38759261518716814,
+ "epoch": 4.41585445094217,
+ "grad_norm": 0.5272114276885986,
+ "learning_rate": 0.00018700092173941517,
+ "loss": 0.3017059135437012,
+ "mean_token_accuracy": 0.8991367623209954,
+ "num_tokens": 4107377.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.39710743524134157,
+ "epoch": 4.545808966861598,
+ "grad_norm": 0.5341858267784119,
+ "learning_rate": 0.00018120491361827535,
+ "loss": 0.30882442474365235,
+ "mean_token_accuracy": 0.8962600353360176,
+ "num_tokens": 4225238.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.39052032932639125,
+ "epoch": 4.675763482781027,
+ "grad_norm": 0.5387603044509888,
+ "learning_rate": 0.00017531696020927306,
+ "loss": 0.3076332473754883,
+ "mean_token_accuracy": 0.8972256830334664,
+ "num_tokens": 4348229.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.38603785961866377,
+ "epoch": 4.805717998700455,
+ "grad_norm": 0.6838284134864807,
+ "learning_rate": 0.00016934915978214512,
+ "loss": 0.3052788162231445,
+ "mean_token_accuracy": 0.8977296784520149,
+ "num_tokens": 4471337.0,
+ "step": 1850
+ },
+ {
+ "entropy": 0.387482148706913,
+ "epoch": 4.935672514619883,
+ "grad_norm": 0.6066106557846069,
+ "learning_rate": 0.00016331377467225454,
+ "loss": 0.3069518852233887,
+ "mean_token_accuracy": 0.8959452581405639,
+ "num_tokens": 4596516.0,
+ "step": 1900
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.47827344932235205,
+ "eval_loss": 0.734696626663208,
+ "eval_mean_token_accuracy": 0.8166056708074533,
+ "eval_num_tokens": 4655435.0,
+ "eval_runtime": 94.3316,
+ "eval_samples_per_second": 17.566,
+ "eval_steps_per_second": 2.205,
+ "step": 1925
+ },
+ {
+ "entropy": 0.3189891557298114,
+ "epoch": 5.064977257959714,
+ "grad_norm": 0.4411616623401642,
+ "learning_rate": 0.00015722320608456218,
+ "loss": 0.24322872161865233,
+ "mean_token_accuracy": 0.9192550799355435,
+ "num_tokens": 4720113.0,
+ "step": 1950
+ },
+ {
+ "entropy": 0.2661040405929089,
+ "epoch": 5.1949317738791425,
+ "grad_norm": 0.5544848442077637,
+ "learning_rate": 0.00015108996861225613,
+ "loss": 0.19045450210571288,
+ "mean_token_accuracy": 0.9353114965558053,
+ "num_tokens": 4842234.0,
+ "step": 2000
+ },
+ {
+ "entropy": 0.2827081228792667,
+ "epoch": 5.32488628979857,
+ "grad_norm": 0.5472713708877563,
+ "learning_rate": 0.0001449266645223969,
+ "loss": 0.197703857421875,
+ "mean_token_accuracy": 0.9314222651720047,
+ "num_tokens": 4959991.0,
+ "step": 2050
+ },
+ {
+ "entropy": 0.28358375541865827,
+ "epoch": 5.454840805717999,
+ "grad_norm": 0.6511954069137573,
+ "learning_rate": 0.00013874595786141455,
+ "loss": 0.20100860595703124,
+ "mean_token_accuracy": 0.930547822713852,
+ "num_tokens": 5082208.0,
+ "step": 2100
+ },
+ {
+ "entropy": 0.2757134060561657,
+ "epoch": 5.584795321637427,
+ "grad_norm": 0.5925538539886475,
+ "learning_rate": 0.0001325605484336644,
+ "loss": 0.19688623428344726,
+ "mean_token_accuracy": 0.9317859390377998,
+ "num_tokens": 5199699.0,
+ "step": 2150
+ },
+ {
+ "entropy": 0.2787610155344009,
+ "epoch": 5.714749837556855,
+ "grad_norm": 0.7486807703971863,
+ "learning_rate": 0.00012638314570651013,
+ "loss": 0.198111629486084,
+ "mean_token_accuracy": 0.9316030797362328,
+ "num_tokens": 5320501.0,
+ "step": 2200
+ },
+ {
+ "entropy": 0.2822082404047251,
+ "epoch": 5.844704353476283,
+ "grad_norm": 0.5981082320213318,
+ "learning_rate": 0.00012022644269555052,
+ "loss": 0.19833782196044922,
+ "mean_token_accuracy": 0.9301980781555176,
+ "num_tokens": 5442291.0,
+ "step": 2250
+ },
+ {
+ "entropy": 0.29100623600184916,
+ "epoch": 5.974658869395712,
+ "grad_norm": 0.6062291860580444,
+ "learning_rate": 0.00011410308988365154,
+ "loss": 0.20115615844726562,
+ "mean_token_accuracy": 0.9285238465666771,
+ "num_tokens": 5560531.0,
+ "step": 2300
+ },
+ {
+ "epoch": 6.0,
+ "eval_entropy": 0.419523782025163,
+ "eval_loss": 0.8228668570518494,
+ "eval_mean_token_accuracy": 0.8120436149720962,
+ "eval_num_tokens": 5586522.0,
+ "eval_runtime": 94.2282,
+ "eval_samples_per_second": 17.585,
+ "eval_steps_per_second": 2.207,
+ "step": 2310
+ },
+ {
+ "entropy": 0.22129053263059215,
+ "epoch": 6.1039636127355426,
+ "grad_norm": 0.5541817545890808,
+ "learning_rate": 0.000108025669227372,
+ "loss": 0.1330717086791992,
+ "mean_token_accuracy": 0.9534786897688056,
+ "num_tokens": 5682578.0,
+ "step": 2350
+ },
+ {
+ "entropy": 0.19409072685986758,
+ "epoch": 6.23391812865497,
+ "grad_norm": 0.31028518080711365,
+ "learning_rate": 0.00010200666830419353,
+ "loss": 0.11706629753112793,
+ "mean_token_accuracy": 0.9599499633908272,
+ "num_tokens": 5803409.0,
+ "step": 2400
+ },
+ {
+ "entropy": 0.18704994775354863,
+ "epoch": 6.363872644574399,
+ "grad_norm": 0.426945298910141,
+ "learning_rate": 9.605845465367642e-05,
+ "loss": 0.11503399848937988,
+ "mean_token_accuracy": 0.9598734942078591,
+ "num_tokens": 5927439.0,
+ "step": 2450
+ },
+ {
+ "entropy": 0.1894980014115572,
+ "epoch": 6.493827160493828,
+ "grad_norm": 0.5031415224075317,
+ "learning_rate": 9.019325036526285e-05,
+ "loss": 0.11405850410461425,
+ "mean_token_accuracy": 0.959916945695877,
+ "num_tokens": 6052633.0,
+ "step": 2500
+ },
+ {
+ "entropy": 0.2080751285329461,
+ "epoch": 6.623781676413255,
+ "grad_norm": 0.6324969530105591,
+ "learning_rate": 8.442310696494393e-05,
+ "loss": 0.1220788288116455,
+ "mean_token_accuracy": 0.9559539663791656,
+ "num_tokens": 6167503.0,
+ "step": 2550
+ },
+ {
+ "entropy": 0.20628964144736528,
+ "epoch": 6.753736192332683,
+ "grad_norm": 0.6685518622398376,
+ "learning_rate": 7.875988065239161e-05,
+ "loss": 0.12388327598571777,
+ "mean_token_accuracy": 0.956074104309082,
+ "num_tokens": 6284023.0,
+ "step": 2600
+ },
+ {
+ "entropy": 0.19158254452049733,
+ "epoch": 6.883690708252112,
+ "grad_norm": 0.5295413136482239,
+ "learning_rate": 7.321520793943799e-05,
+ "loss": 0.11677920341491699,
+ "mean_token_accuracy": 0.9589221009612083,
+ "num_tokens": 6408327.0,
+ "step": 2650
+ },
+ {
+ "epoch": 7.0,
+ "eval_entropy": 0.3437433656161794,
+ "eval_loss": 0.9526592493057251,
+ "eval_mean_token_accuracy": 0.8121905275262319,
+ "eval_num_tokens": 6517609.0,
+ "eval_runtime": 94.2057,
+ "eval_samples_per_second": 17.589,
+ "eval_steps_per_second": 2.208,
+ "step": 2695
+ },
+ {
+ "entropy": 0.18569647883949567,
+ "epoch": 7.012995451591943,
+ "grad_norm": 0.4049379825592041,
+ "learning_rate": 6.78004817399571e-05,
+ "loss": 0.10891223907470703,
+ "mean_token_accuracy": 0.9622611065006735,
+ "num_tokens": 6530430.0,
+ "step": 2700
+ },
+ {
+ "entropy": 0.14998033180832862,
+ "epoch": 7.142949967511371,
+ "grad_norm": 0.39533936977386475,
+ "learning_rate": 6.252682796028047e-05,
+ "loss": 0.07854948997497559,
+ "mean_token_accuracy": 0.9729882365465164,
+ "num_tokens": 6648623.0,
+ "step": 2750
+ },
+ {
+ "entropy": 0.1513198823109269,
+ "epoch": 7.272904483430799,
+ "grad_norm": 0.40231335163116455,
+ "learning_rate": 5.740508263824541e-05,
+ "loss": 0.07896852493286133,
+ "mean_token_accuracy": 0.9724258282780647,
+ "num_tokens": 6768614.0,
+ "step": 2800
+ },
+ {
+ "entropy": 0.1426772189885378,
+ "epoch": 7.402858999350228,
+ "grad_norm": 0.39278867840766907,
+ "learning_rate": 5.244576967785141e-05,
+ "loss": 0.07671234607696534,
+ "mean_token_accuracy": 0.9739191561937333,
+ "num_tokens": 6892420.0,
+ "step": 2850
+ },
+ {
+ "entropy": 0.15510446103289724,
+ "epoch": 7.532813515269655,
+ "grad_norm": 0.45904669165611267,
+ "learning_rate": 4.7659079225272575e-05,
+ "loss": 0.08021929740905762,
+ "mean_token_accuracy": 0.9704085075855255,
+ "num_tokens": 7009857.0,
+ "step": 2900
+ },
+ {
+ "entropy": 0.14695645896717907,
+ "epoch": 7.662768031189084,
+ "grad_norm": 0.2910745441913605,
+ "learning_rate": 4.305484673065923e-05,
+ "loss": 0.07789832592010498,
+ "mean_token_accuracy": 0.9713015568256378,
+ "num_tokens": 7132184.0,
+ "step": 2950
+ },
+ {
+ "entropy": 0.15163813527673484,
+ "epoch": 7.792722547108512,
+ "grad_norm": 0.26081109046936035,
+ "learning_rate": 3.8642532738751e-05,
+ "loss": 0.07821487426757813,
+ "mean_token_accuracy": 0.9711482360959053,
+ "num_tokens": 7250668.0,
+ "step": 3000
+ },
+ {
+ "entropy": 0.15378343284130097,
+ "epoch": 7.92267706302794,
+ "grad_norm": 0.35973286628723145,
+ "learning_rate": 3.443120344982667e-05,
+ "loss": 0.07847892284393311,
+ "mean_token_accuracy": 0.9707216301560402,
+ "num_tokens": 7373911.0,
+ "step": 3050
+ },
+ {
+ "epoch": 8.0,
+ "eval_entropy": 0.30457046585014236,
+ "eval_loss": 1.0573656558990479,
+ "eval_mean_token_accuracy": 0.8123422947067481,
+ "eval_num_tokens": 7448696.0,
+ "eval_runtime": 94.7627,
+ "eval_samples_per_second": 17.486,
+ "eval_steps_per_second": 2.195,
+ "step": 3080
+ },
+ {
+ "entropy": 0.14268321996957214,
+ "epoch": 8.05198180636777,
+ "grad_norm": 0.17200376093387604,
+ "learning_rate": 3.0429512090932775e-05,
+ "loss": 0.06907873630523681,
+ "mean_token_accuracy": 0.9740184904941961,
+ "num_tokens": 7497025.0,
+ "step": 3100
+ },
+ {
+ "entropy": 0.13242442483082414,
+ "epoch": 8.1819363222872,
+ "grad_norm": 0.24965617060661316,
+ "learning_rate": 2.6645681135669307e-05,
+ "loss": 0.06381748676300049,
+ "mean_token_accuracy": 0.9758135312795639,
+ "num_tokens": 7615828.0,
+ "step": 3150
+ },
+ {
+ "entropy": 0.13472190242260695,
+ "epoch": 8.311890838206628,
+ "grad_norm": 0.24549734592437744,
+ "learning_rate": 2.3087485409065462e-05,
+ "loss": 0.06418798446655273,
+ "mean_token_accuracy": 0.9750474771857262,
+ "num_tokens": 7734498.0,
+ "step": 3200
+ },
+ {
+ "entropy": 0.12980990454554558,
+ "epoch": 8.441845354126055,
+ "grad_norm": 0.16283877193927765,
+ "learning_rate": 1.9762236112261417e-05,
+ "loss": 0.06304348945617676,
+ "mean_token_accuracy": 0.9762365907430649,
+ "num_tokens": 7855780.0,
+ "step": 3250
+ },
+ {
+ "entropy": 0.14281578866764902,
+ "epoch": 8.571799870045485,
+ "grad_norm": 0.25190046429634094,
+ "learning_rate": 1.667676579982073e-05,
+ "loss": 0.06881052494049072,
+ "mean_token_accuracy": 0.9739083594083786,
+ "num_tokens": 7968890.0,
+ "step": 3300
+ },
+ {
+ "entropy": 0.1275078719109297,
+ "epoch": 8.701754385964913,
+ "grad_norm": 0.14581522345542908,
+ "learning_rate": 1.3837414340541901e-05,
+ "loss": 0.06218691825866699,
+ "mean_token_accuracy": 0.9765720379352569,
+ "num_tokens": 8094439.0,
+ "step": 3350
+ },
+ {
+ "entropy": 0.12800191594287752,
+ "epoch": 8.83170890188434,
+ "grad_norm": 0.15930895507335663,
+ "learning_rate": 1.1250015890615666e-05,
+ "loss": 0.061544852256774904,
+ "mean_token_accuracy": 0.9760145637392997,
+ "num_tokens": 8219578.0,
+ "step": 3400
+ },
+ {
+ "entropy": 0.12469653191044927,
+ "epoch": 8.961663417803768,
+ "grad_norm": 0.3019782304763794,
+ "learning_rate": 8.91988690589514e-06,
+ "loss": 0.06210709571838379,
+ "mean_token_accuracy": 0.9761517831683159,
+ "num_tokens": 8345441.0,
+ "step": 3450
+ },
+ {
+ "epoch": 9.0,
+ "eval_entropy": 0.2802187033140889,
+ "eval_loss": 1.1696693897247314,
+ "eval_mean_token_accuracy": 0.8132876937205975,
+ "eval_num_tokens": 8379783.0,
+ "eval_runtime": 94.6664,
+ "eval_samples_per_second": 17.504,
+ "eval_steps_per_second": 2.197,
+ "step": 3465
+ },
+ {
+ "entropy": 0.11983221430499949,
+ "epoch": 9.0909681611436,
+ "grad_norm": 0.30692893266677856,
+ "learning_rate": 6.8518152179102544e-06,
+ "loss": 0.059758267402648925,
+ "mean_token_accuracy": 0.9791372684977162,
+ "num_tokens": 8466213.0,
+ "step": 3500
+ },
+ {
+ "entropy": 0.11968761058524251,
+ "epoch": 9.220922677063028,
+ "grad_norm": 0.12794272601604462,
+ "learning_rate": 5.050050196073183e-06,
+ "loss": 0.05675140857696533,
+ "mean_token_accuracy": 0.9780975925922394,
+ "num_tokens": 8591232.0,
+ "step": 3550
+ },
+ {
+ "entropy": 0.1262346575409174,
+ "epoch": 9.350877192982455,
+ "grad_norm": 0.20644903182983398,
+ "learning_rate": 3.5182940162881887e-06,
+ "loss": 0.05778531074523926,
+ "mean_token_accuracy": 0.9765303027629852,
+ "num_tokens": 8714639.0,
+ "step": 3600
+ },
+ {
+ "entropy": 0.13082476263865828,
+ "epoch": 9.480831708901885,
+ "grad_norm": 0.14169760048389435,
+ "learning_rate": 2.259694053907472e-06,
+ "loss": 0.06011291027069092,
+ "mean_token_accuracy": 0.9766910344362258,
+ "num_tokens": 8833066.0,
+ "step": 3650
+ },
+ {
+ "entropy": 0.1270183707214892,
+ "epoch": 9.610786224821313,
+ "grad_norm": 0.19377458095550537,
+ "learning_rate": 1.2768364166628255e-06,
+ "loss": 0.06239473819732666,
+ "mean_token_accuracy": 0.9762869846820831,
+ "num_tokens": 8949213.0,
+ "step": 3700
+ },
+ {
+ "entropy": 0.12354950185865164,
+ "epoch": 9.74074074074074,
+ "grad_norm": 0.12469393759965897,
+ "learning_rate": 5.717406308619949e-07,
+ "loss": 0.060302953720092776,
+ "mean_token_accuracy": 0.9774003893136978,
+ "num_tokens": 9069096.0,
+ "step": 3750
+ },
+ {
+ "entropy": 0.12386846126988531,
+ "epoch": 9.870695256660168,
+ "grad_norm": 0.1613702028989792,
+ "learning_rate": 1.4585549176784973e-07,
+ "loss": 0.06068916320800781,
+ "mean_token_accuracy": 0.9767120435833931,
+ "num_tokens": 9188485.0,
+ "step": 3800
+ },
+ {
+ "entropy": 0.1248967067230886,
+ "epoch": 10.0,
+ "grad_norm": 0.21104362607002258,
+ "learning_rate": 5.60866869303437e-11,
+ "loss": 0.05824681282043457,
+ "mean_token_accuracy": 0.9775403708069768,
+ "num_tokens": 9310870.0,
+ "step": 3850
+ },
+ {
+ "epoch": 10.0,
+ "eval_entropy": 0.27227992163254666,
+ "eval_loss": 1.2223320007324219,
+ "eval_mean_token_accuracy": 0.8136088647521459,
+ "eval_num_tokens": 9310870.0,
+ "eval_runtime": 94.3261,
+ "eval_samples_per_second": 17.567,
+ "eval_steps_per_second": 2.205,
+ "step": 3850
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3850,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": true
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 4.068648008425697e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-770/README.md b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-770/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-770/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-770/adapter_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-770/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..26742a540eca0befcc61f09a270940ded2abaca9
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-770/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 128,
+ "lora_bias": false,
+ "lora_dropout": 0.07386918480460901,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "k_proj",
+ "down_proj",
+ "q_proj",
+ "v_proj",
+ "up_proj",
+ "gate_proj",
+ "o_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-770/chat_template.jinja b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-770/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-770/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-770/tokenizer_config.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-770/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-770/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-770/trainer_state.json b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-770/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..f25b99ca1c02305ab9890c8da1ad20d68237a39b
--- /dev/null
+++ b/DBCA_original_Estonian/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-770/trainer_state.json
@@ -0,0 +1,206 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 2.0,
+ "eval_steps": 500,
+ "global_step": 770,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5888759356737137,
+ "epoch": 0.1299545159194282,
+ "grad_norm": 1.1323111057281494,
+ "learning_rate": 3.473456711878286e-05,
+ "loss": 1.4242234802246094,
+ "mean_token_accuracy": 0.6861670060455799,
+ "num_tokens": 127115.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.8638093116879463,
+ "epoch": 0.2599090318388564,
+ "grad_norm": 0.7135503888130188,
+ "learning_rate": 7.017800295427557e-05,
+ "loss": 0.7731611633300781,
+ "mean_token_accuracy": 0.7900564879179001,
+ "num_tokens": 251796.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.8165761995315551,
+ "epoch": 0.3898635477582846,
+ "grad_norm": 0.5835468173027039,
+ "learning_rate": 0.0001056214387897683,
+ "loss": 0.7186955261230469,
+ "mean_token_accuracy": 0.7992442399263382,
+ "num_tokens": 366505.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.7586414662003517,
+ "epoch": 0.5198180636777128,
+ "grad_norm": 0.5721960663795471,
+ "learning_rate": 0.00014106487462526099,
+ "loss": 0.6786507415771484,
+ "mean_token_accuracy": 0.8069057819247246,
+ "num_tokens": 493150.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.7753578001260757,
+ "epoch": 0.649772579597141,
+ "grad_norm": 0.6396453380584717,
+ "learning_rate": 0.00017650831046075373,
+ "loss": 0.6813118743896485,
+ "mean_token_accuracy": 0.806209616959095,
+ "num_tokens": 607197.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.7591985522210598,
+ "epoch": 0.7797270955165692,
+ "grad_norm": 0.577964186668396,
+ "learning_rate": 0.00021195174629624644,
+ "loss": 0.67177490234375,
+ "mean_token_accuracy": 0.8090321889519692,
+ "num_tokens": 726579.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.731777046918869,
+ "epoch": 0.9096816114359974,
+ "grad_norm": 0.597808301448822,
+ "learning_rate": 0.00024739518213173913,
+ "loss": 0.6601863861083984,
+ "mean_token_accuracy": 0.8120808112621307,
+ "num_tokens": 848944.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.7157932419616443,
+ "eval_loss": 0.7318932414054871,
+ "eval_mean_token_accuracy": 0.7979064600972029,
+ "eval_num_tokens": 931087.0,
+ "eval_runtime": 94.7419,
+ "eval_samples_per_second": 17.49,
+ "eval_steps_per_second": 2.195,
+ "step": 385
+ },
+ {
+ "entropy": 0.7265488104005555,
+ "epoch": 1.0389863547758285,
+ "grad_norm": 0.7224313020706177,
+ "learning_rate": 0.00027290346308950203,
+ "loss": 0.6452214050292969,
+ "mean_token_accuracy": 0.8156503406002293,
+ "num_tokens": 968580.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.6715995308756828,
+ "epoch": 1.1689408706952567,
+ "grad_norm": 0.6290106177330017,
+ "learning_rate": 0.000272684789300892,
+ "loss": 0.6042086791992187,
+ "mean_token_accuracy": 0.8257312250137329,
+ "num_tokens": 1090002.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.6747516925632954,
+ "epoch": 1.2988953866146848,
+ "grad_norm": 0.66269850730896,
+ "learning_rate": 0.00027218620198917387,
+ "loss": 0.6056245803833008,
+ "mean_token_accuracy": 0.8240418726205826,
+ "num_tokens": 1212161.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.6872365647554397,
+ "epoch": 1.428849902534113,
+ "grad_norm": 0.4894042909145355,
+ "learning_rate": 0.00027140872562641226,
+ "loss": 0.615659523010254,
+ "mean_token_accuracy": 0.8227636247873307,
+ "num_tokens": 1332852.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.6712110449373722,
+ "epoch": 1.5588044184535412,
+ "grad_norm": 0.9492104649543762,
+ "learning_rate": 0.00027035395773182937,
+ "loss": 0.595902099609375,
+ "mean_token_accuracy": 0.8245656326413154,
+ "num_tokens": 1453569.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.6576551312208175,
+ "epoch": 1.6887589343729694,
+ "grad_norm": 0.6480705142021179,
+ "learning_rate": 0.00026902406558930333,
+ "loss": 0.5864276885986328,
+ "mean_token_accuracy": 0.8273968940973282,
+ "num_tokens": 1577152.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.6850327280163765,
+ "epoch": 1.8187134502923976,
+ "grad_norm": 0.516665518283844,
+ "learning_rate": 0.0002674217817941425,
+ "loss": 0.6083594512939453,
+ "mean_token_accuracy": 0.8232863625884056,
+ "num_tokens": 1691926.0,
+ "step": 700
+ },
+ {
+ "entropy": 0.6665814685821533,
+ "epoch": 1.9486679662118258,
+ "grad_norm": 0.6762926578521729,
+ "learning_rate": 0.00026555039863828637,
+ "loss": 0.6011666488647461,
+ "mean_token_accuracy": 0.8256330332159996,
+ "num_tokens": 1816059.0,
+ "step": 750
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6573245783264821,
+ "eval_loss": 0.6852125525474548,
+ "eval_mean_token_accuracy": 0.808370346060166,
+ "eval_num_tokens": 1862174.0,
+ "eval_runtime": 94.2081,
+ "eval_samples_per_second": 17.589,
+ "eval_steps_per_second": 2.208,
+ "step": 770
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3850,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 8.15020951668265e+16,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..7924d5d0f7e14bbf8603df4c998cd8aab78346da
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/README.md
@@ -0,0 +1,58 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: transformers
+model_name: Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1
+tags:
+- generated_from_trainer
+- trl
+- sft
+licence: license
+---
+
+# Model Card for Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1
+
+This model is a fine-tuned version of [Qwen/Qwen3.5-4B-Base](https://huggingface.co/Qwen/Qwen3.5-4B-Base).
+It has been trained using [TRL](https://github.com/huggingface/trl).
+
+## Quick start
+
+```python
+from transformers import pipeline
+
+question = "If you had a time machine, but could only go to the past or the future once and never return, which would you choose and why?"
+generator = pipeline("text-generation", model="None", device="cuda")
+output = generator([{"role": "user", "content": question}], max_new_tokens=128, return_full_text=False)[0]
+print(output["generated_text"])
+```
+
+## Training procedure
+
+[
](https://wandb.ai/katriin-kukk/Cross_lingual_morphological_generalization/runs/wgwowzj9)
+
+
+
+This model was trained with SFT.
+
+### Framework versions
+
+- TRL: 0.29.0
+- Transformers: 5.5.4
+- Pytorch: 2.10.0
+- Datasets: 4.6.1
+- Tokenizers: 0.22.2
+
+## Citations
+
+
+
+Cite TRL as:
+
+```bibtex
+@software{vonwerra2020trl,
+ title = {{TRL: Transformers Reinforcement Learning}},
+ author = {von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and Beeching, Edward and Thrush, Tristan and Lambert, Nathan and Huang, Shengyi and Rasul, Kashif and Gallouédec, Quentin},
+ license = {Apache-2.0},
+ url = {https://github.com/huggingface/trl},
+ year = {2020}
+}
+```
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1122/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1122/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1122/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1122/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1122/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..aaa5315fd59e264ef47a82c1fe8efb39be4ce81b
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1122/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.0024401823792587042,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1122/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1122/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1122/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1122/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1122/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1122/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1122/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1122/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..677ff6bfa67120354e1a324cd4946d6403be8773
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1122/trainer_state.json
@@ -0,0 +1,287 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 3.0,
+ "eval_steps": 500,
+ "global_step": 1122,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5728709718585014,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.7528014183044434,
+ "learning_rate": 1.188290252952284e-05,
+ "loss": 1.354185791015625,
+ "mean_token_accuracy": 0.7051774816215038,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.7372898015379906,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.4441661834716797,
+ "learning_rate": 2.40083132739339e-05,
+ "loss": 0.6506095123291016,
+ "mean_token_accuracy": 0.8222023022174835,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6791197884082795,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.9964745044708252,
+ "learning_rate": 3.613372401834496e-05,
+ "loss": 0.6006810760498047,
+ "mean_token_accuracy": 0.8327413862943649,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6392384143173695,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.8044542074203491,
+ "learning_rate": 4.825913476275602e-05,
+ "loss": 0.5597672653198242,
+ "mean_token_accuracy": 0.8428809601068497,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6373909455537796,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.9033862948417664,
+ "learning_rate": 6.038454550716708e-05,
+ "loss": 0.5607868576049805,
+ "mean_token_accuracy": 0.8415500408411026,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6142180925607681,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8247693777084351,
+ "learning_rate": 7.250995625157815e-05,
+ "loss": 0.5359848785400391,
+ "mean_token_accuracy": 0.8485029026865959,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.600434636771679,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9828518033027649,
+ "learning_rate": 8.46353669959892e-05,
+ "loss": 0.5246802139282226,
+ "mean_token_accuracy": 0.8504493471980095,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6199610435962677,
+ "eval_loss": 0.5767685174942017,
+ "eval_mean_token_accuracy": 0.8355379092693329,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 98.478,
+ "eval_samples_per_second": 16.237,
+ "eval_steps_per_second": 2.031,
+ "step": 374
+ },
+ {
+ "entropy": 0.5812384999460645,
+ "epoch": 1.069611780455154,
+ "grad_norm": 0.7832330465316772,
+ "learning_rate": 9.068572797060741e-05,
+ "loss": 0.5035604858398437,
+ "mean_token_accuracy": 0.8548897610168265,
+ "num_tokens": 939587.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.5494279730319976,
+ "epoch": 1.2034805890227578,
+ "grad_norm": 0.6947465538978577,
+ "learning_rate": 9.058701310926409e-05,
+ "loss": 0.4750754928588867,
+ "mean_token_accuracy": 0.8600558358430862,
+ "num_tokens": 1058209.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.5506373339891434,
+ "epoch": 1.3373493975903614,
+ "grad_norm": 0.7013292908668518,
+ "learning_rate": 9.038979832559184e-05,
+ "loss": 0.4743109893798828,
+ "mean_token_accuracy": 0.8615988761186599,
+ "num_tokens": 1178618.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.5413181179761887,
+ "epoch": 1.4712182061579653,
+ "grad_norm": 0.8555615544319153,
+ "learning_rate": 9.009451302961728e-05,
+ "loss": 0.4731967544555664,
+ "mean_token_accuracy": 0.8605538108944892,
+ "num_tokens": 1299124.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.5346147198975086,
+ "epoch": 1.605087014725569,
+ "grad_norm": 0.8122630715370178,
+ "learning_rate": 8.970180016739357e-05,
+ "loss": 0.46506633758544924,
+ "mean_token_accuracy": 0.8624722695350647,
+ "num_tokens": 1422003.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.530957569181919,
+ "epoch": 1.7389558232931726,
+ "grad_norm": 0.6519771218299866,
+ "learning_rate": 8.921251482106754e-05,
+ "loss": 0.46199348449707034,
+ "mean_token_accuracy": 0.8636697196960449,
+ "num_tokens": 1542283.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.5241960515081883,
+ "epoch": 1.8728246318607764,
+ "grad_norm": 0.5876154899597168,
+ "learning_rate": 8.862772234704737e-05,
+ "loss": 0.45739620208740234,
+ "mean_token_accuracy": 0.866628175675869,
+ "num_tokens": 1656994.0,
+ "step": 700
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.5676956498622894,
+ "eval_loss": 0.545108437538147,
+ "eval_mean_token_accuracy": 0.8408122035861015,
+ "eval_num_tokens": 1763438.0,
+ "eval_runtime": 97.562,
+ "eval_samples_per_second": 16.39,
+ "eval_steps_per_second": 2.05,
+ "step": 748
+ },
+ {
+ "entropy": 0.5400724257483627,
+ "epoch": 2.005354752342704,
+ "grad_norm": 0.5679355263710022,
+ "learning_rate": 8.794869605632493e-05,
+ "loss": 0.46216156005859377,
+ "mean_token_accuracy": 0.8634762803111413,
+ "num_tokens": 1768737.0,
+ "step": 750
+ },
+ {
+ "entropy": 0.4396473586559296,
+ "epoch": 2.139223560910308,
+ "grad_norm": 0.8222826719284058,
+ "learning_rate": 8.717691444200361e-05,
+ "loss": 0.36421436309814453,
+ "mean_token_accuracy": 0.8855674511194229,
+ "num_tokens": 1888608.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.4505964604765177,
+ "epoch": 2.2730923694779115,
+ "grad_norm": 0.6411374807357788,
+ "learning_rate": 8.631405796006825e-05,
+ "loss": 0.37386734008789063,
+ "mean_token_accuracy": 0.8835385522246361,
+ "num_tokens": 2000711.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.4520136344432831,
+ "epoch": 2.4069611780455156,
+ "grad_norm": 0.665027379989624,
+ "learning_rate": 8.536200537040663e-05,
+ "loss": 0.37581691741943357,
+ "mean_token_accuracy": 0.884546649158001,
+ "num_tokens": 2117216.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.4466621845960617,
+ "epoch": 2.540829986613119,
+ "grad_norm": 0.6488580107688904,
+ "learning_rate": 8.432282964604958e-05,
+ "loss": 0.37471736907958986,
+ "mean_token_accuracy": 0.8835149678587914,
+ "num_tokens": 2233392.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.4457112967967987,
+ "epoch": 2.674698795180723,
+ "grad_norm": 0.9634214043617249,
+ "learning_rate": 8.31987934595367e-05,
+ "loss": 0.37740741729736327,
+ "mean_token_accuracy": 0.8837782579660416,
+ "num_tokens": 2349527.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.4474602049589157,
+ "epoch": 2.8085676037483265,
+ "grad_norm": 0.6158106923103333,
+ "learning_rate": 8.199234425623558e-05,
+ "loss": 0.3742681121826172,
+ "mean_token_accuracy": 0.8844600400328636,
+ "num_tokens": 2473970.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.45432978719472883,
+ "epoch": 2.9424364123159306,
+ "grad_norm": 0.6572170853614807,
+ "learning_rate": 8.070610892534193e-05,
+ "loss": 0.3823386001586914,
+ "mean_token_accuracy": 0.8826713335514068,
+ "num_tokens": 2593727.0,
+ "step": 1100
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.4939903026819229,
+ "eval_loss": 0.561303436756134,
+ "eval_mean_token_accuracy": 0.8428723356127739,
+ "eval_num_tokens": 2645157.0,
+ "eval_runtime": 97.2032,
+ "eval_samples_per_second": 16.45,
+ "eval_steps_per_second": 2.058,
+ "step": 1122
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.120904934916608e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1496/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1496/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1496/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1496/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1496/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..aaa5315fd59e264ef47a82c1fe8efb39be4ce81b
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1496/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.0024401823792587042,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1496/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1496/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1496/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1496/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1496/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1496/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1496/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1496/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..aff86c52433579bd147ca9e8d1614d3ade53b1a8
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1496/trainer_state.json
@@ -0,0 +1,368 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 4.0,
+ "eval_steps": 500,
+ "global_step": 1496,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5728709718585014,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.7528014183044434,
+ "learning_rate": 1.188290252952284e-05,
+ "loss": 1.354185791015625,
+ "mean_token_accuracy": 0.7051774816215038,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.7372898015379906,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.4441661834716797,
+ "learning_rate": 2.40083132739339e-05,
+ "loss": 0.6506095123291016,
+ "mean_token_accuracy": 0.8222023022174835,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6791197884082795,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.9964745044708252,
+ "learning_rate": 3.613372401834496e-05,
+ "loss": 0.6006810760498047,
+ "mean_token_accuracy": 0.8327413862943649,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6392384143173695,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.8044542074203491,
+ "learning_rate": 4.825913476275602e-05,
+ "loss": 0.5597672653198242,
+ "mean_token_accuracy": 0.8428809601068497,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6373909455537796,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.9033862948417664,
+ "learning_rate": 6.038454550716708e-05,
+ "loss": 0.5607868576049805,
+ "mean_token_accuracy": 0.8415500408411026,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6142180925607681,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8247693777084351,
+ "learning_rate": 7.250995625157815e-05,
+ "loss": 0.5359848785400391,
+ "mean_token_accuracy": 0.8485029026865959,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.600434636771679,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9828518033027649,
+ "learning_rate": 8.46353669959892e-05,
+ "loss": 0.5246802139282226,
+ "mean_token_accuracy": 0.8504493471980095,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6199610435962677,
+ "eval_loss": 0.5767685174942017,
+ "eval_mean_token_accuracy": 0.8355379092693329,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 98.478,
+ "eval_samples_per_second": 16.237,
+ "eval_steps_per_second": 2.031,
+ "step": 374
+ },
+ {
+ "entropy": 0.5812384999460645,
+ "epoch": 1.069611780455154,
+ "grad_norm": 0.7832330465316772,
+ "learning_rate": 9.068572797060741e-05,
+ "loss": 0.5035604858398437,
+ "mean_token_accuracy": 0.8548897610168265,
+ "num_tokens": 939587.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.5494279730319976,
+ "epoch": 1.2034805890227578,
+ "grad_norm": 0.6947465538978577,
+ "learning_rate": 9.058701310926409e-05,
+ "loss": 0.4750754928588867,
+ "mean_token_accuracy": 0.8600558358430862,
+ "num_tokens": 1058209.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.5506373339891434,
+ "epoch": 1.3373493975903614,
+ "grad_norm": 0.7013292908668518,
+ "learning_rate": 9.038979832559184e-05,
+ "loss": 0.4743109893798828,
+ "mean_token_accuracy": 0.8615988761186599,
+ "num_tokens": 1178618.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.5413181179761887,
+ "epoch": 1.4712182061579653,
+ "grad_norm": 0.8555615544319153,
+ "learning_rate": 9.009451302961728e-05,
+ "loss": 0.4731967544555664,
+ "mean_token_accuracy": 0.8605538108944892,
+ "num_tokens": 1299124.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.5346147198975086,
+ "epoch": 1.605087014725569,
+ "grad_norm": 0.8122630715370178,
+ "learning_rate": 8.970180016739357e-05,
+ "loss": 0.46506633758544924,
+ "mean_token_accuracy": 0.8624722695350647,
+ "num_tokens": 1422003.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.530957569181919,
+ "epoch": 1.7389558232931726,
+ "grad_norm": 0.6519771218299866,
+ "learning_rate": 8.921251482106754e-05,
+ "loss": 0.46199348449707034,
+ "mean_token_accuracy": 0.8636697196960449,
+ "num_tokens": 1542283.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.5241960515081883,
+ "epoch": 1.8728246318607764,
+ "grad_norm": 0.5876154899597168,
+ "learning_rate": 8.862772234704737e-05,
+ "loss": 0.45739620208740234,
+ "mean_token_accuracy": 0.866628175675869,
+ "num_tokens": 1656994.0,
+ "step": 700
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.5676956498622894,
+ "eval_loss": 0.545108437538147,
+ "eval_mean_token_accuracy": 0.8408122035861015,
+ "eval_num_tokens": 1763438.0,
+ "eval_runtime": 97.562,
+ "eval_samples_per_second": 16.39,
+ "eval_steps_per_second": 2.05,
+ "step": 748
+ },
+ {
+ "entropy": 0.5400724257483627,
+ "epoch": 2.005354752342704,
+ "grad_norm": 0.5679355263710022,
+ "learning_rate": 8.794869605632493e-05,
+ "loss": 0.46216156005859377,
+ "mean_token_accuracy": 0.8634762803111413,
+ "num_tokens": 1768737.0,
+ "step": 750
+ },
+ {
+ "entropy": 0.4396473586559296,
+ "epoch": 2.139223560910308,
+ "grad_norm": 0.8222826719284058,
+ "learning_rate": 8.717691444200361e-05,
+ "loss": 0.36421436309814453,
+ "mean_token_accuracy": 0.8855674511194229,
+ "num_tokens": 1888608.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.4505964604765177,
+ "epoch": 2.2730923694779115,
+ "grad_norm": 0.6411374807357788,
+ "learning_rate": 8.631405796006825e-05,
+ "loss": 0.37386734008789063,
+ "mean_token_accuracy": 0.8835385522246361,
+ "num_tokens": 2000711.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.4520136344432831,
+ "epoch": 2.4069611780455156,
+ "grad_norm": 0.665027379989624,
+ "learning_rate": 8.536200537040663e-05,
+ "loss": 0.37581691741943357,
+ "mean_token_accuracy": 0.884546649158001,
+ "num_tokens": 2117216.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.4466621845960617,
+ "epoch": 2.540829986613119,
+ "grad_norm": 0.6488580107688904,
+ "learning_rate": 8.432282964604958e-05,
+ "loss": 0.37471736907958986,
+ "mean_token_accuracy": 0.8835149678587914,
+ "num_tokens": 2233392.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.4457112967967987,
+ "epoch": 2.674698795180723,
+ "grad_norm": 0.9634214043617249,
+ "learning_rate": 8.31987934595367e-05,
+ "loss": 0.37740741729736327,
+ "mean_token_accuracy": 0.8837782579660416,
+ "num_tokens": 2349527.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.4474602049589157,
+ "epoch": 2.8085676037483265,
+ "grad_norm": 0.6158106923103333,
+ "learning_rate": 8.199234425623558e-05,
+ "loss": 0.3742681121826172,
+ "mean_token_accuracy": 0.8844600400328636,
+ "num_tokens": 2473970.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.45432978719472883,
+ "epoch": 2.9424364123159306,
+ "grad_norm": 0.6572170853614807,
+ "learning_rate": 8.070610892534193e-05,
+ "loss": 0.3823386001586914,
+ "mean_token_accuracy": 0.8826713335514068,
+ "num_tokens": 2593727.0,
+ "step": 1100
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.4939903026819229,
+ "eval_loss": 0.561303436756134,
+ "eval_mean_token_accuracy": 0.8428723356127739,
+ "eval_num_tokens": 2645157.0,
+ "eval_runtime": 97.2032,
+ "eval_samples_per_second": 16.45,
+ "eval_steps_per_second": 2.058,
+ "step": 1122
+ },
+ {
+ "entropy": 0.39069811209584726,
+ "epoch": 3.074966532797858,
+ "grad_norm": 0.8949297070503235,
+ "learning_rate": 7.934288808016343e-05,
+ "loss": 0.31472652435302734,
+ "mean_token_accuracy": 0.9009616712127069,
+ "num_tokens": 2715241.0,
+ "step": 1150
+ },
+ {
+ "entropy": 0.3619236435741186,
+ "epoch": 3.208835341365462,
+ "grad_norm": 0.8773637413978577,
+ "learning_rate": 7.790564996014168e-05,
+ "loss": 0.273685131072998,
+ "mean_token_accuracy": 0.9100899347662925,
+ "num_tokens": 2828619.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.35230451248586175,
+ "epoch": 3.3427041499330654,
+ "grad_norm": 0.6625242829322815,
+ "learning_rate": 7.639752396788952e-05,
+ "loss": 0.27389230728149416,
+ "mean_token_accuracy": 0.9103762644529343,
+ "num_tokens": 2950688.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.35693769574165346,
+ "epoch": 3.4765729585006695,
+ "grad_norm": 0.7137399315834045,
+ "learning_rate": 7.482179385531625e-05,
+ "loss": 0.27805461883544924,
+ "mean_token_accuracy": 0.9091088688373565,
+ "num_tokens": 3071943.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.3644157887250185,
+ "epoch": 3.610441767068273,
+ "grad_norm": 0.9558516144752502,
+ "learning_rate": 7.318189057367674e-05,
+ "loss": 0.2868117141723633,
+ "mean_token_accuracy": 0.9085465380549431,
+ "num_tokens": 3187873.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.3672615347057581,
+ "epoch": 3.7443105756358768,
+ "grad_norm": 0.8229874968528748,
+ "learning_rate": 7.14813848031129e-05,
+ "loss": 0.28742366790771484,
+ "mean_token_accuracy": 0.9074980768561364,
+ "num_tokens": 3299002.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.3580747715383768,
+ "epoch": 3.878179384203481,
+ "grad_norm": 0.8101242184638977,
+ "learning_rate": 6.972397917795341e-05,
+ "loss": 0.28595699310302736,
+ "mean_token_accuracy": 0.9081357708573341,
+ "num_tokens": 3422804.0,
+ "step": 1450
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.4521959410607815,
+ "eval_loss": 0.592692494392395,
+ "eval_mean_token_accuracy": 0.841086029112339,
+ "eval_num_tokens": 3526876.0,
+ "eval_runtime": 96.9985,
+ "eval_samples_per_second": 16.485,
+ "eval_steps_per_second": 2.062,
+ "step": 1496
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.494594216880988e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1870/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1870/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1870/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1870/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1870/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..aaa5315fd59e264ef47a82c1fe8efb39be4ce81b
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1870/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.0024401823792587042,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1870/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1870/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1870/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1870/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1870/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1870/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1870/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1870/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..eec6092cb144354f1541c6bae5271e74adcd8e90
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-1870/trainer_state.json
@@ -0,0 +1,459 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 5.0,
+ "eval_steps": 500,
+ "global_step": 1870,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5728709718585014,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.7528014183044434,
+ "learning_rate": 1.188290252952284e-05,
+ "loss": 1.354185791015625,
+ "mean_token_accuracy": 0.7051774816215038,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.7372898015379906,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.4441661834716797,
+ "learning_rate": 2.40083132739339e-05,
+ "loss": 0.6506095123291016,
+ "mean_token_accuracy": 0.8222023022174835,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6791197884082795,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.9964745044708252,
+ "learning_rate": 3.613372401834496e-05,
+ "loss": 0.6006810760498047,
+ "mean_token_accuracy": 0.8327413862943649,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6392384143173695,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.8044542074203491,
+ "learning_rate": 4.825913476275602e-05,
+ "loss": 0.5597672653198242,
+ "mean_token_accuracy": 0.8428809601068497,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6373909455537796,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.9033862948417664,
+ "learning_rate": 6.038454550716708e-05,
+ "loss": 0.5607868576049805,
+ "mean_token_accuracy": 0.8415500408411026,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6142180925607681,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8247693777084351,
+ "learning_rate": 7.250995625157815e-05,
+ "loss": 0.5359848785400391,
+ "mean_token_accuracy": 0.8485029026865959,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.600434636771679,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9828518033027649,
+ "learning_rate": 8.46353669959892e-05,
+ "loss": 0.5246802139282226,
+ "mean_token_accuracy": 0.8504493471980095,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6199610435962677,
+ "eval_loss": 0.5767685174942017,
+ "eval_mean_token_accuracy": 0.8355379092693329,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 98.478,
+ "eval_samples_per_second": 16.237,
+ "eval_steps_per_second": 2.031,
+ "step": 374
+ },
+ {
+ "entropy": 0.5812384999460645,
+ "epoch": 1.069611780455154,
+ "grad_norm": 0.7832330465316772,
+ "learning_rate": 9.068572797060741e-05,
+ "loss": 0.5035604858398437,
+ "mean_token_accuracy": 0.8548897610168265,
+ "num_tokens": 939587.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.5494279730319976,
+ "epoch": 1.2034805890227578,
+ "grad_norm": 0.6947465538978577,
+ "learning_rate": 9.058701310926409e-05,
+ "loss": 0.4750754928588867,
+ "mean_token_accuracy": 0.8600558358430862,
+ "num_tokens": 1058209.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.5506373339891434,
+ "epoch": 1.3373493975903614,
+ "grad_norm": 0.7013292908668518,
+ "learning_rate": 9.038979832559184e-05,
+ "loss": 0.4743109893798828,
+ "mean_token_accuracy": 0.8615988761186599,
+ "num_tokens": 1178618.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.5413181179761887,
+ "epoch": 1.4712182061579653,
+ "grad_norm": 0.8555615544319153,
+ "learning_rate": 9.009451302961728e-05,
+ "loss": 0.4731967544555664,
+ "mean_token_accuracy": 0.8605538108944892,
+ "num_tokens": 1299124.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.5346147198975086,
+ "epoch": 1.605087014725569,
+ "grad_norm": 0.8122630715370178,
+ "learning_rate": 8.970180016739357e-05,
+ "loss": 0.46506633758544924,
+ "mean_token_accuracy": 0.8624722695350647,
+ "num_tokens": 1422003.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.530957569181919,
+ "epoch": 1.7389558232931726,
+ "grad_norm": 0.6519771218299866,
+ "learning_rate": 8.921251482106754e-05,
+ "loss": 0.46199348449707034,
+ "mean_token_accuracy": 0.8636697196960449,
+ "num_tokens": 1542283.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.5241960515081883,
+ "epoch": 1.8728246318607764,
+ "grad_norm": 0.5876154899597168,
+ "learning_rate": 8.862772234704737e-05,
+ "loss": 0.45739620208740234,
+ "mean_token_accuracy": 0.866628175675869,
+ "num_tokens": 1656994.0,
+ "step": 700
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.5676956498622894,
+ "eval_loss": 0.545108437538147,
+ "eval_mean_token_accuracy": 0.8408122035861015,
+ "eval_num_tokens": 1763438.0,
+ "eval_runtime": 97.562,
+ "eval_samples_per_second": 16.39,
+ "eval_steps_per_second": 2.05,
+ "step": 748
+ },
+ {
+ "entropy": 0.5400724257483627,
+ "epoch": 2.005354752342704,
+ "grad_norm": 0.5679355263710022,
+ "learning_rate": 8.794869605632493e-05,
+ "loss": 0.46216156005859377,
+ "mean_token_accuracy": 0.8634762803111413,
+ "num_tokens": 1768737.0,
+ "step": 750
+ },
+ {
+ "entropy": 0.4396473586559296,
+ "epoch": 2.139223560910308,
+ "grad_norm": 0.8222826719284058,
+ "learning_rate": 8.717691444200361e-05,
+ "loss": 0.36421436309814453,
+ "mean_token_accuracy": 0.8855674511194229,
+ "num_tokens": 1888608.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.4505964604765177,
+ "epoch": 2.2730923694779115,
+ "grad_norm": 0.6411374807357788,
+ "learning_rate": 8.631405796006825e-05,
+ "loss": 0.37386734008789063,
+ "mean_token_accuracy": 0.8835385522246361,
+ "num_tokens": 2000711.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.4520136344432831,
+ "epoch": 2.4069611780455156,
+ "grad_norm": 0.665027379989624,
+ "learning_rate": 8.536200537040663e-05,
+ "loss": 0.37581691741943357,
+ "mean_token_accuracy": 0.884546649158001,
+ "num_tokens": 2117216.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.4466621845960617,
+ "epoch": 2.540829986613119,
+ "grad_norm": 0.6488580107688904,
+ "learning_rate": 8.432282964604958e-05,
+ "loss": 0.37471736907958986,
+ "mean_token_accuracy": 0.8835149678587914,
+ "num_tokens": 2233392.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.4457112967967987,
+ "epoch": 2.674698795180723,
+ "grad_norm": 0.9634214043617249,
+ "learning_rate": 8.31987934595367e-05,
+ "loss": 0.37740741729736327,
+ "mean_token_accuracy": 0.8837782579660416,
+ "num_tokens": 2349527.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.4474602049589157,
+ "epoch": 2.8085676037483265,
+ "grad_norm": 0.6158106923103333,
+ "learning_rate": 8.199234425623558e-05,
+ "loss": 0.3742681121826172,
+ "mean_token_accuracy": 0.8844600400328636,
+ "num_tokens": 2473970.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.45432978719472883,
+ "epoch": 2.9424364123159306,
+ "grad_norm": 0.6572170853614807,
+ "learning_rate": 8.070610892534193e-05,
+ "loss": 0.3823386001586914,
+ "mean_token_accuracy": 0.8826713335514068,
+ "num_tokens": 2593727.0,
+ "step": 1100
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.4939903026819229,
+ "eval_loss": 0.561303436756134,
+ "eval_mean_token_accuracy": 0.8428723356127739,
+ "eval_num_tokens": 2645157.0,
+ "eval_runtime": 97.2032,
+ "eval_samples_per_second": 16.45,
+ "eval_steps_per_second": 2.058,
+ "step": 1122
+ },
+ {
+ "entropy": 0.39069811209584726,
+ "epoch": 3.074966532797858,
+ "grad_norm": 0.8949297070503235,
+ "learning_rate": 7.934288808016343e-05,
+ "loss": 0.31472652435302734,
+ "mean_token_accuracy": 0.9009616712127069,
+ "num_tokens": 2715241.0,
+ "step": 1150
+ },
+ {
+ "entropy": 0.3619236435741186,
+ "epoch": 3.208835341365462,
+ "grad_norm": 0.8773637413978577,
+ "learning_rate": 7.790564996014168e-05,
+ "loss": 0.273685131072998,
+ "mean_token_accuracy": 0.9100899347662925,
+ "num_tokens": 2828619.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.35230451248586175,
+ "epoch": 3.3427041499330654,
+ "grad_norm": 0.6625242829322815,
+ "learning_rate": 7.639752396788952e-05,
+ "loss": 0.27389230728149416,
+ "mean_token_accuracy": 0.9103762644529343,
+ "num_tokens": 2950688.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.35693769574165346,
+ "epoch": 3.4765729585006695,
+ "grad_norm": 0.7137399315834045,
+ "learning_rate": 7.482179385531625e-05,
+ "loss": 0.27805461883544924,
+ "mean_token_accuracy": 0.9091088688373565,
+ "num_tokens": 3071943.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.3644157887250185,
+ "epoch": 3.610441767068273,
+ "grad_norm": 0.9558516144752502,
+ "learning_rate": 7.318189057367674e-05,
+ "loss": 0.2868117141723633,
+ "mean_token_accuracy": 0.9085465380549431,
+ "num_tokens": 3187873.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.3672615347057581,
+ "epoch": 3.7443105756358768,
+ "grad_norm": 0.8229874968528748,
+ "learning_rate": 7.14813848031129e-05,
+ "loss": 0.28742366790771484,
+ "mean_token_accuracy": 0.9074980768561364,
+ "num_tokens": 3299002.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.3580747715383768,
+ "epoch": 3.878179384203481,
+ "grad_norm": 0.8101242184638977,
+ "learning_rate": 6.972397917795341e-05,
+ "loss": 0.28595699310302736,
+ "mean_token_accuracy": 0.9081357708573341,
+ "num_tokens": 3422804.0,
+ "step": 1450
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.4521959410607815,
+ "eval_loss": 0.592692494392395,
+ "eval_mean_token_accuracy": 0.841086029112339,
+ "eval_num_tokens": 3526876.0,
+ "eval_runtime": 96.9985,
+ "eval_samples_per_second": 16.485,
+ "eval_steps_per_second": 2.062,
+ "step": 1496
+ },
+ {
+ "entropy": 0.36107602956319096,
+ "epoch": 4.010709504685408,
+ "grad_norm": 0.8092170357704163,
+ "learning_rate": 6.791350022469971e-05,
+ "loss": 0.28187707901000975,
+ "mean_token_accuracy": 0.9096371344845704,
+ "num_tokens": 3536363.0,
+ "step": 1500
+ },
+ {
+ "entropy": 0.2679470381140709,
+ "epoch": 4.144578313253012,
+ "grad_norm": 1.0743237733840942,
+ "learning_rate": 6.605389003025307e-05,
+ "loss": 0.18246625900268554,
+ "mean_token_accuracy": 0.9399469995498657,
+ "num_tokens": 3653560.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.2658273372799158,
+ "epoch": 4.278447121820616,
+ "grad_norm": 0.9514140486717224,
+ "learning_rate": 6.41491976585231e-05,
+ "loss": 0.18667778015136718,
+ "mean_token_accuracy": 0.9373631486296654,
+ "num_tokens": 3776281.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.267765996158123,
+ "epoch": 4.412315930388219,
+ "grad_norm": 0.8411495089530945,
+ "learning_rate": 6.220357033410804e-05,
+ "loss": 0.1890517807006836,
+ "mean_token_accuracy": 0.9369775268435478,
+ "num_tokens": 3896081.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.27206149250268935,
+ "epoch": 4.546184738955823,
+ "grad_norm": 1.1611251831054688,
+ "learning_rate": 6.022124441224217e-05,
+ "loss": 0.1913918685913086,
+ "mean_token_accuracy": 0.9364263373613357,
+ "num_tokens": 4010469.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.2756531854718924,
+ "epoch": 4.680053547523427,
+ "grad_norm": 0.7768945097923279,
+ "learning_rate": 5.820653615467293e-05,
+ "loss": 0.1946771240234375,
+ "mean_token_accuracy": 0.9357997533679009,
+ "num_tokens": 4125858.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.2766863085329533,
+ "epoch": 4.813922356091031,
+ "grad_norm": 0.986640453338623,
+ "learning_rate": 5.6163832331551755e-05,
+ "loss": 0.19539087295532226,
+ "mean_token_accuracy": 0.9345261418819427,
+ "num_tokens": 4243193.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.26934320479631424,
+ "epoch": 4.947791164658635,
+ "grad_norm": 0.8279967308044434,
+ "learning_rate": 5.4097580669801786e-05,
+ "loss": 0.18980844497680663,
+ "mean_token_accuracy": 0.9369515025615692,
+ "num_tokens": 4365873.0,
+ "step": 1850
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.38973631352186205,
+ "eval_loss": 0.6432627439498901,
+ "eval_mean_token_accuracy": 0.8423123425245285,
+ "eval_num_tokens": 4408595.0,
+ "eval_runtime": 97.135,
+ "eval_samples_per_second": 16.462,
+ "eval_steps_per_second": 2.059,
+ "step": 1870
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.8661573535249203e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2244/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2244/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2244/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2244/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2244/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..aaa5315fd59e264ef47a82c1fe8efb39be4ce81b
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2244/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.0024401823792587042,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2244/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2244/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2244/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2244/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2244/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2244/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2244/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2244/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..cf274cb6d5347ffdfa88f1d46605e8e767be97c4
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2244/trainer_state.json
@@ -0,0 +1,540 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 6.0,
+ "eval_steps": 500,
+ "global_step": 2244,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5728709718585014,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.7528014183044434,
+ "learning_rate": 1.188290252952284e-05,
+ "loss": 1.354185791015625,
+ "mean_token_accuracy": 0.7051774816215038,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.7372898015379906,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.4441661834716797,
+ "learning_rate": 2.40083132739339e-05,
+ "loss": 0.6506095123291016,
+ "mean_token_accuracy": 0.8222023022174835,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6791197884082795,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.9964745044708252,
+ "learning_rate": 3.613372401834496e-05,
+ "loss": 0.6006810760498047,
+ "mean_token_accuracy": 0.8327413862943649,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6392384143173695,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.8044542074203491,
+ "learning_rate": 4.825913476275602e-05,
+ "loss": 0.5597672653198242,
+ "mean_token_accuracy": 0.8428809601068497,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6373909455537796,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.9033862948417664,
+ "learning_rate": 6.038454550716708e-05,
+ "loss": 0.5607868576049805,
+ "mean_token_accuracy": 0.8415500408411026,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6142180925607681,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8247693777084351,
+ "learning_rate": 7.250995625157815e-05,
+ "loss": 0.5359848785400391,
+ "mean_token_accuracy": 0.8485029026865959,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.600434636771679,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9828518033027649,
+ "learning_rate": 8.46353669959892e-05,
+ "loss": 0.5246802139282226,
+ "mean_token_accuracy": 0.8504493471980095,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6199610435962677,
+ "eval_loss": 0.5767685174942017,
+ "eval_mean_token_accuracy": 0.8355379092693329,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 98.478,
+ "eval_samples_per_second": 16.237,
+ "eval_steps_per_second": 2.031,
+ "step": 374
+ },
+ {
+ "entropy": 0.5812384999460645,
+ "epoch": 1.069611780455154,
+ "grad_norm": 0.7832330465316772,
+ "learning_rate": 9.068572797060741e-05,
+ "loss": 0.5035604858398437,
+ "mean_token_accuracy": 0.8548897610168265,
+ "num_tokens": 939587.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.5494279730319976,
+ "epoch": 1.2034805890227578,
+ "grad_norm": 0.6947465538978577,
+ "learning_rate": 9.058701310926409e-05,
+ "loss": 0.4750754928588867,
+ "mean_token_accuracy": 0.8600558358430862,
+ "num_tokens": 1058209.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.5506373339891434,
+ "epoch": 1.3373493975903614,
+ "grad_norm": 0.7013292908668518,
+ "learning_rate": 9.038979832559184e-05,
+ "loss": 0.4743109893798828,
+ "mean_token_accuracy": 0.8615988761186599,
+ "num_tokens": 1178618.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.5413181179761887,
+ "epoch": 1.4712182061579653,
+ "grad_norm": 0.8555615544319153,
+ "learning_rate": 9.009451302961728e-05,
+ "loss": 0.4731967544555664,
+ "mean_token_accuracy": 0.8605538108944892,
+ "num_tokens": 1299124.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.5346147198975086,
+ "epoch": 1.605087014725569,
+ "grad_norm": 0.8122630715370178,
+ "learning_rate": 8.970180016739357e-05,
+ "loss": 0.46506633758544924,
+ "mean_token_accuracy": 0.8624722695350647,
+ "num_tokens": 1422003.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.530957569181919,
+ "epoch": 1.7389558232931726,
+ "grad_norm": 0.6519771218299866,
+ "learning_rate": 8.921251482106754e-05,
+ "loss": 0.46199348449707034,
+ "mean_token_accuracy": 0.8636697196960449,
+ "num_tokens": 1542283.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.5241960515081883,
+ "epoch": 1.8728246318607764,
+ "grad_norm": 0.5876154899597168,
+ "learning_rate": 8.862772234704737e-05,
+ "loss": 0.45739620208740234,
+ "mean_token_accuracy": 0.866628175675869,
+ "num_tokens": 1656994.0,
+ "step": 700
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.5676956498622894,
+ "eval_loss": 0.545108437538147,
+ "eval_mean_token_accuracy": 0.8408122035861015,
+ "eval_num_tokens": 1763438.0,
+ "eval_runtime": 97.562,
+ "eval_samples_per_second": 16.39,
+ "eval_steps_per_second": 2.05,
+ "step": 748
+ },
+ {
+ "entropy": 0.5400724257483627,
+ "epoch": 2.005354752342704,
+ "grad_norm": 0.5679355263710022,
+ "learning_rate": 8.794869605632493e-05,
+ "loss": 0.46216156005859377,
+ "mean_token_accuracy": 0.8634762803111413,
+ "num_tokens": 1768737.0,
+ "step": 750
+ },
+ {
+ "entropy": 0.4396473586559296,
+ "epoch": 2.139223560910308,
+ "grad_norm": 0.8222826719284058,
+ "learning_rate": 8.717691444200361e-05,
+ "loss": 0.36421436309814453,
+ "mean_token_accuracy": 0.8855674511194229,
+ "num_tokens": 1888608.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.4505964604765177,
+ "epoch": 2.2730923694779115,
+ "grad_norm": 0.6411374807357788,
+ "learning_rate": 8.631405796006825e-05,
+ "loss": 0.37386734008789063,
+ "mean_token_accuracy": 0.8835385522246361,
+ "num_tokens": 2000711.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.4520136344432831,
+ "epoch": 2.4069611780455156,
+ "grad_norm": 0.665027379989624,
+ "learning_rate": 8.536200537040663e-05,
+ "loss": 0.37581691741943357,
+ "mean_token_accuracy": 0.884546649158001,
+ "num_tokens": 2117216.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.4466621845960617,
+ "epoch": 2.540829986613119,
+ "grad_norm": 0.6488580107688904,
+ "learning_rate": 8.432282964604958e-05,
+ "loss": 0.37471736907958986,
+ "mean_token_accuracy": 0.8835149678587914,
+ "num_tokens": 2233392.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.4457112967967987,
+ "epoch": 2.674698795180723,
+ "grad_norm": 0.9634214043617249,
+ "learning_rate": 8.31987934595367e-05,
+ "loss": 0.37740741729736327,
+ "mean_token_accuracy": 0.8837782579660416,
+ "num_tokens": 2349527.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.4474602049589157,
+ "epoch": 2.8085676037483265,
+ "grad_norm": 0.6158106923103333,
+ "learning_rate": 8.199234425623558e-05,
+ "loss": 0.3742681121826172,
+ "mean_token_accuracy": 0.8844600400328636,
+ "num_tokens": 2473970.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.45432978719472883,
+ "epoch": 2.9424364123159306,
+ "grad_norm": 0.6572170853614807,
+ "learning_rate": 8.070610892534193e-05,
+ "loss": 0.3823386001586914,
+ "mean_token_accuracy": 0.8826713335514068,
+ "num_tokens": 2593727.0,
+ "step": 1100
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.4939903026819229,
+ "eval_loss": 0.561303436756134,
+ "eval_mean_token_accuracy": 0.8428723356127739,
+ "eval_num_tokens": 2645157.0,
+ "eval_runtime": 97.2032,
+ "eval_samples_per_second": 16.45,
+ "eval_steps_per_second": 2.058,
+ "step": 1122
+ },
+ {
+ "entropy": 0.39069811209584726,
+ "epoch": 3.074966532797858,
+ "grad_norm": 0.8949297070503235,
+ "learning_rate": 7.934288808016343e-05,
+ "loss": 0.31472652435302734,
+ "mean_token_accuracy": 0.9009616712127069,
+ "num_tokens": 2715241.0,
+ "step": 1150
+ },
+ {
+ "entropy": 0.3619236435741186,
+ "epoch": 3.208835341365462,
+ "grad_norm": 0.8773637413978577,
+ "learning_rate": 7.790564996014168e-05,
+ "loss": 0.273685131072998,
+ "mean_token_accuracy": 0.9100899347662925,
+ "num_tokens": 2828619.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.35230451248586175,
+ "epoch": 3.3427041499330654,
+ "grad_norm": 0.6625242829322815,
+ "learning_rate": 7.639752396788952e-05,
+ "loss": 0.27389230728149416,
+ "mean_token_accuracy": 0.9103762644529343,
+ "num_tokens": 2950688.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.35693769574165346,
+ "epoch": 3.4765729585006695,
+ "grad_norm": 0.7137399315834045,
+ "learning_rate": 7.482179385531625e-05,
+ "loss": 0.27805461883544924,
+ "mean_token_accuracy": 0.9091088688373565,
+ "num_tokens": 3071943.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.3644157887250185,
+ "epoch": 3.610441767068273,
+ "grad_norm": 0.9558516144752502,
+ "learning_rate": 7.318189057367674e-05,
+ "loss": 0.2868117141723633,
+ "mean_token_accuracy": 0.9085465380549431,
+ "num_tokens": 3187873.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.3672615347057581,
+ "epoch": 3.7443105756358768,
+ "grad_norm": 0.8229874968528748,
+ "learning_rate": 7.14813848031129e-05,
+ "loss": 0.28742366790771484,
+ "mean_token_accuracy": 0.9074980768561364,
+ "num_tokens": 3299002.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.3580747715383768,
+ "epoch": 3.878179384203481,
+ "grad_norm": 0.8101242184638977,
+ "learning_rate": 6.972397917795341e-05,
+ "loss": 0.28595699310302736,
+ "mean_token_accuracy": 0.9081357708573341,
+ "num_tokens": 3422804.0,
+ "step": 1450
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.4521959410607815,
+ "eval_loss": 0.592692494392395,
+ "eval_mean_token_accuracy": 0.841086029112339,
+ "eval_num_tokens": 3526876.0,
+ "eval_runtime": 96.9985,
+ "eval_samples_per_second": 16.485,
+ "eval_steps_per_second": 2.062,
+ "step": 1496
+ },
+ {
+ "entropy": 0.36107602956319096,
+ "epoch": 4.010709504685408,
+ "grad_norm": 0.8092170357704163,
+ "learning_rate": 6.791350022469971e-05,
+ "loss": 0.28187707901000975,
+ "mean_token_accuracy": 0.9096371344845704,
+ "num_tokens": 3536363.0,
+ "step": 1500
+ },
+ {
+ "entropy": 0.2679470381140709,
+ "epoch": 4.144578313253012,
+ "grad_norm": 1.0743237733840942,
+ "learning_rate": 6.605389003025307e-05,
+ "loss": 0.18246625900268554,
+ "mean_token_accuracy": 0.9399469995498657,
+ "num_tokens": 3653560.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.2658273372799158,
+ "epoch": 4.278447121820616,
+ "grad_norm": 0.9514140486717224,
+ "learning_rate": 6.41491976585231e-05,
+ "loss": 0.18667778015136718,
+ "mean_token_accuracy": 0.9373631486296654,
+ "num_tokens": 3776281.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.267765996158123,
+ "epoch": 4.412315930388219,
+ "grad_norm": 0.8411495089530945,
+ "learning_rate": 6.220357033410804e-05,
+ "loss": 0.1890517807006836,
+ "mean_token_accuracy": 0.9369775268435478,
+ "num_tokens": 3896081.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.27206149250268935,
+ "epoch": 4.546184738955823,
+ "grad_norm": 1.1611251831054688,
+ "learning_rate": 6.022124441224217e-05,
+ "loss": 0.1913918685913086,
+ "mean_token_accuracy": 0.9364263373613357,
+ "num_tokens": 4010469.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.2756531854718924,
+ "epoch": 4.680053547523427,
+ "grad_norm": 0.7768945097923279,
+ "learning_rate": 5.820653615467293e-05,
+ "loss": 0.1946771240234375,
+ "mean_token_accuracy": 0.9357997533679009,
+ "num_tokens": 4125858.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.2766863085329533,
+ "epoch": 4.813922356091031,
+ "grad_norm": 0.986640453338623,
+ "learning_rate": 5.6163832331551755e-05,
+ "loss": 0.19539087295532226,
+ "mean_token_accuracy": 0.9345261418819427,
+ "num_tokens": 4243193.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.26934320479631424,
+ "epoch": 4.947791164658635,
+ "grad_norm": 0.8279967308044434,
+ "learning_rate": 5.4097580669801786e-05,
+ "loss": 0.18980844497680663,
+ "mean_token_accuracy": 0.9369515025615692,
+ "num_tokens": 4365873.0,
+ "step": 1850
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.38973631352186205,
+ "eval_loss": 0.6432627439498901,
+ "eval_mean_token_accuracy": 0.8423123425245285,
+ "eval_num_tokens": 4408595.0,
+ "eval_runtime": 97.135,
+ "eval_samples_per_second": 16.462,
+ "eval_steps_per_second": 2.059,
+ "step": 1870
+ },
+ {
+ "entropy": 0.22773028695673653,
+ "epoch": 5.080321285140562,
+ "grad_norm": 0.6459256410598755,
+ "learning_rate": 5.2012280168760034e-05,
+ "loss": 0.14593751907348632,
+ "mean_token_accuracy": 0.9516990624292933,
+ "num_tokens": 4482540.0,
+ "step": 1900
+ },
+ {
+ "entropy": 0.20054464034736155,
+ "epoch": 5.214190093708166,
+ "grad_norm": 0.8127110004425049,
+ "learning_rate": 4.9912471304180415e-05,
+ "loss": 0.1171640968322754,
+ "mean_token_accuracy": 0.9613950234651566,
+ "num_tokens": 4598191.0,
+ "step": 1950
+ },
+ {
+ "entropy": 0.19813392639160157,
+ "epoch": 5.34805890227577,
+ "grad_norm": 0.9870916604995728,
+ "learning_rate": 4.780272614192717e-05,
+ "loss": 0.11783818244934081,
+ "mean_token_accuracy": 0.9606451898813247,
+ "num_tokens": 4716333.0,
+ "step": 2000
+ },
+ {
+ "entropy": 0.20009616125375032,
+ "epoch": 5.481927710843373,
+ "grad_norm": 0.9596051573753357,
+ "learning_rate": 4.568763838288482e-05,
+ "loss": 0.12061534881591797,
+ "mean_token_accuracy": 0.9601768556237221,
+ "num_tokens": 4830262.0,
+ "step": 2050
+ },
+ {
+ "entropy": 0.19717911910265684,
+ "epoch": 5.615796519410977,
+ "grad_norm": 0.9989078044891357,
+ "learning_rate": 4.357181336076072e-05,
+ "loss": 0.11976994514465332,
+ "mean_token_accuracy": 0.9600497630238533,
+ "num_tokens": 4951007.0,
+ "step": 2100
+ },
+ {
+ "entropy": 0.1992201693728566,
+ "epoch": 5.749665327978581,
+ "grad_norm": 0.8984506130218506,
+ "learning_rate": 4.145985801455844e-05,
+ "loss": 0.12019011497497559,
+ "mean_token_accuracy": 0.9582142195105553,
+ "num_tokens": 5067535.0,
+ "step": 2150
+ },
+ {
+ "entropy": 0.1871098166331649,
+ "epoch": 5.883534136546185,
+ "grad_norm": 0.9786492586135864,
+ "learning_rate": 3.9356370857556064e-05,
+ "loss": 0.1156222915649414,
+ "mean_token_accuracy": 0.9616268679499627,
+ "num_tokens": 5185751.0,
+ "step": 2200
+ },
+ {
+ "epoch": 6.0,
+ "eval_entropy": 0.3187944942712784,
+ "eval_loss": 0.7801651358604431,
+ "eval_mean_token_accuracy": 0.837693527340889,
+ "eval_num_tokens": 5290314.0,
+ "eval_runtime": 97.5774,
+ "eval_samples_per_second": 16.387,
+ "eval_steps_per_second": 2.05,
+ "step": 2244
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.2368361483527168e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2618/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2618/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2618/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2618/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2618/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..aaa5315fd59e264ef47a82c1fe8efb39be4ce81b
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2618/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.0024401823792587042,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2618/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2618/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2618/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2618/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2618/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2618/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2618/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2618/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..6c56cb559803b645551dab96c25c2b59bde4490f
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2618/trainer_state.json
@@ -0,0 +1,631 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 7.0,
+ "eval_steps": 500,
+ "global_step": 2618,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5728709718585014,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.7528014183044434,
+ "learning_rate": 1.188290252952284e-05,
+ "loss": 1.354185791015625,
+ "mean_token_accuracy": 0.7051774816215038,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.7372898015379906,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.4441661834716797,
+ "learning_rate": 2.40083132739339e-05,
+ "loss": 0.6506095123291016,
+ "mean_token_accuracy": 0.8222023022174835,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6791197884082795,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.9964745044708252,
+ "learning_rate": 3.613372401834496e-05,
+ "loss": 0.6006810760498047,
+ "mean_token_accuracy": 0.8327413862943649,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6392384143173695,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.8044542074203491,
+ "learning_rate": 4.825913476275602e-05,
+ "loss": 0.5597672653198242,
+ "mean_token_accuracy": 0.8428809601068497,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6373909455537796,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.9033862948417664,
+ "learning_rate": 6.038454550716708e-05,
+ "loss": 0.5607868576049805,
+ "mean_token_accuracy": 0.8415500408411026,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6142180925607681,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8247693777084351,
+ "learning_rate": 7.250995625157815e-05,
+ "loss": 0.5359848785400391,
+ "mean_token_accuracy": 0.8485029026865959,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.600434636771679,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9828518033027649,
+ "learning_rate": 8.46353669959892e-05,
+ "loss": 0.5246802139282226,
+ "mean_token_accuracy": 0.8504493471980095,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6199610435962677,
+ "eval_loss": 0.5767685174942017,
+ "eval_mean_token_accuracy": 0.8355379092693329,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 98.478,
+ "eval_samples_per_second": 16.237,
+ "eval_steps_per_second": 2.031,
+ "step": 374
+ },
+ {
+ "entropy": 0.5812384999460645,
+ "epoch": 1.069611780455154,
+ "grad_norm": 0.7832330465316772,
+ "learning_rate": 9.068572797060741e-05,
+ "loss": 0.5035604858398437,
+ "mean_token_accuracy": 0.8548897610168265,
+ "num_tokens": 939587.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.5494279730319976,
+ "epoch": 1.2034805890227578,
+ "grad_norm": 0.6947465538978577,
+ "learning_rate": 9.058701310926409e-05,
+ "loss": 0.4750754928588867,
+ "mean_token_accuracy": 0.8600558358430862,
+ "num_tokens": 1058209.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.5506373339891434,
+ "epoch": 1.3373493975903614,
+ "grad_norm": 0.7013292908668518,
+ "learning_rate": 9.038979832559184e-05,
+ "loss": 0.4743109893798828,
+ "mean_token_accuracy": 0.8615988761186599,
+ "num_tokens": 1178618.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.5413181179761887,
+ "epoch": 1.4712182061579653,
+ "grad_norm": 0.8555615544319153,
+ "learning_rate": 9.009451302961728e-05,
+ "loss": 0.4731967544555664,
+ "mean_token_accuracy": 0.8605538108944892,
+ "num_tokens": 1299124.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.5346147198975086,
+ "epoch": 1.605087014725569,
+ "grad_norm": 0.8122630715370178,
+ "learning_rate": 8.970180016739357e-05,
+ "loss": 0.46506633758544924,
+ "mean_token_accuracy": 0.8624722695350647,
+ "num_tokens": 1422003.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.530957569181919,
+ "epoch": 1.7389558232931726,
+ "grad_norm": 0.6519771218299866,
+ "learning_rate": 8.921251482106754e-05,
+ "loss": 0.46199348449707034,
+ "mean_token_accuracy": 0.8636697196960449,
+ "num_tokens": 1542283.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.5241960515081883,
+ "epoch": 1.8728246318607764,
+ "grad_norm": 0.5876154899597168,
+ "learning_rate": 8.862772234704737e-05,
+ "loss": 0.45739620208740234,
+ "mean_token_accuracy": 0.866628175675869,
+ "num_tokens": 1656994.0,
+ "step": 700
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.5676956498622894,
+ "eval_loss": 0.545108437538147,
+ "eval_mean_token_accuracy": 0.8408122035861015,
+ "eval_num_tokens": 1763438.0,
+ "eval_runtime": 97.562,
+ "eval_samples_per_second": 16.39,
+ "eval_steps_per_second": 2.05,
+ "step": 748
+ },
+ {
+ "entropy": 0.5400724257483627,
+ "epoch": 2.005354752342704,
+ "grad_norm": 0.5679355263710022,
+ "learning_rate": 8.794869605632493e-05,
+ "loss": 0.46216156005859377,
+ "mean_token_accuracy": 0.8634762803111413,
+ "num_tokens": 1768737.0,
+ "step": 750
+ },
+ {
+ "entropy": 0.4396473586559296,
+ "epoch": 2.139223560910308,
+ "grad_norm": 0.8222826719284058,
+ "learning_rate": 8.717691444200361e-05,
+ "loss": 0.36421436309814453,
+ "mean_token_accuracy": 0.8855674511194229,
+ "num_tokens": 1888608.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.4505964604765177,
+ "epoch": 2.2730923694779115,
+ "grad_norm": 0.6411374807357788,
+ "learning_rate": 8.631405796006825e-05,
+ "loss": 0.37386734008789063,
+ "mean_token_accuracy": 0.8835385522246361,
+ "num_tokens": 2000711.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.4520136344432831,
+ "epoch": 2.4069611780455156,
+ "grad_norm": 0.665027379989624,
+ "learning_rate": 8.536200537040663e-05,
+ "loss": 0.37581691741943357,
+ "mean_token_accuracy": 0.884546649158001,
+ "num_tokens": 2117216.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.4466621845960617,
+ "epoch": 2.540829986613119,
+ "grad_norm": 0.6488580107688904,
+ "learning_rate": 8.432282964604958e-05,
+ "loss": 0.37471736907958986,
+ "mean_token_accuracy": 0.8835149678587914,
+ "num_tokens": 2233392.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.4457112967967987,
+ "epoch": 2.674698795180723,
+ "grad_norm": 0.9634214043617249,
+ "learning_rate": 8.31987934595367e-05,
+ "loss": 0.37740741729736327,
+ "mean_token_accuracy": 0.8837782579660416,
+ "num_tokens": 2349527.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.4474602049589157,
+ "epoch": 2.8085676037483265,
+ "grad_norm": 0.6158106923103333,
+ "learning_rate": 8.199234425623558e-05,
+ "loss": 0.3742681121826172,
+ "mean_token_accuracy": 0.8844600400328636,
+ "num_tokens": 2473970.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.45432978719472883,
+ "epoch": 2.9424364123159306,
+ "grad_norm": 0.6572170853614807,
+ "learning_rate": 8.070610892534193e-05,
+ "loss": 0.3823386001586914,
+ "mean_token_accuracy": 0.8826713335514068,
+ "num_tokens": 2593727.0,
+ "step": 1100
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.4939903026819229,
+ "eval_loss": 0.561303436756134,
+ "eval_mean_token_accuracy": 0.8428723356127739,
+ "eval_num_tokens": 2645157.0,
+ "eval_runtime": 97.2032,
+ "eval_samples_per_second": 16.45,
+ "eval_steps_per_second": 2.058,
+ "step": 1122
+ },
+ {
+ "entropy": 0.39069811209584726,
+ "epoch": 3.074966532797858,
+ "grad_norm": 0.8949297070503235,
+ "learning_rate": 7.934288808016343e-05,
+ "loss": 0.31472652435302734,
+ "mean_token_accuracy": 0.9009616712127069,
+ "num_tokens": 2715241.0,
+ "step": 1150
+ },
+ {
+ "entropy": 0.3619236435741186,
+ "epoch": 3.208835341365462,
+ "grad_norm": 0.8773637413978577,
+ "learning_rate": 7.790564996014168e-05,
+ "loss": 0.273685131072998,
+ "mean_token_accuracy": 0.9100899347662925,
+ "num_tokens": 2828619.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.35230451248586175,
+ "epoch": 3.3427041499330654,
+ "grad_norm": 0.6625242829322815,
+ "learning_rate": 7.639752396788952e-05,
+ "loss": 0.27389230728149416,
+ "mean_token_accuracy": 0.9103762644529343,
+ "num_tokens": 2950688.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.35693769574165346,
+ "epoch": 3.4765729585006695,
+ "grad_norm": 0.7137399315834045,
+ "learning_rate": 7.482179385531625e-05,
+ "loss": 0.27805461883544924,
+ "mean_token_accuracy": 0.9091088688373565,
+ "num_tokens": 3071943.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.3644157887250185,
+ "epoch": 3.610441767068273,
+ "grad_norm": 0.9558516144752502,
+ "learning_rate": 7.318189057367674e-05,
+ "loss": 0.2868117141723633,
+ "mean_token_accuracy": 0.9085465380549431,
+ "num_tokens": 3187873.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.3672615347057581,
+ "epoch": 3.7443105756358768,
+ "grad_norm": 0.8229874968528748,
+ "learning_rate": 7.14813848031129e-05,
+ "loss": 0.28742366790771484,
+ "mean_token_accuracy": 0.9074980768561364,
+ "num_tokens": 3299002.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.3580747715383768,
+ "epoch": 3.878179384203481,
+ "grad_norm": 0.8101242184638977,
+ "learning_rate": 6.972397917795341e-05,
+ "loss": 0.28595699310302736,
+ "mean_token_accuracy": 0.9081357708573341,
+ "num_tokens": 3422804.0,
+ "step": 1450
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.4521959410607815,
+ "eval_loss": 0.592692494392395,
+ "eval_mean_token_accuracy": 0.841086029112339,
+ "eval_num_tokens": 3526876.0,
+ "eval_runtime": 96.9985,
+ "eval_samples_per_second": 16.485,
+ "eval_steps_per_second": 2.062,
+ "step": 1496
+ },
+ {
+ "entropy": 0.36107602956319096,
+ "epoch": 4.010709504685408,
+ "grad_norm": 0.8092170357704163,
+ "learning_rate": 6.791350022469971e-05,
+ "loss": 0.28187707901000975,
+ "mean_token_accuracy": 0.9096371344845704,
+ "num_tokens": 3536363.0,
+ "step": 1500
+ },
+ {
+ "entropy": 0.2679470381140709,
+ "epoch": 4.144578313253012,
+ "grad_norm": 1.0743237733840942,
+ "learning_rate": 6.605389003025307e-05,
+ "loss": 0.18246625900268554,
+ "mean_token_accuracy": 0.9399469995498657,
+ "num_tokens": 3653560.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.2658273372799158,
+ "epoch": 4.278447121820616,
+ "grad_norm": 0.9514140486717224,
+ "learning_rate": 6.41491976585231e-05,
+ "loss": 0.18667778015136718,
+ "mean_token_accuracy": 0.9373631486296654,
+ "num_tokens": 3776281.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.267765996158123,
+ "epoch": 4.412315930388219,
+ "grad_norm": 0.8411495089530945,
+ "learning_rate": 6.220357033410804e-05,
+ "loss": 0.1890517807006836,
+ "mean_token_accuracy": 0.9369775268435478,
+ "num_tokens": 3896081.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.27206149250268935,
+ "epoch": 4.546184738955823,
+ "grad_norm": 1.1611251831054688,
+ "learning_rate": 6.022124441224217e-05,
+ "loss": 0.1913918685913086,
+ "mean_token_accuracy": 0.9364263373613357,
+ "num_tokens": 4010469.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.2756531854718924,
+ "epoch": 4.680053547523427,
+ "grad_norm": 0.7768945097923279,
+ "learning_rate": 5.820653615467293e-05,
+ "loss": 0.1946771240234375,
+ "mean_token_accuracy": 0.9357997533679009,
+ "num_tokens": 4125858.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.2766863085329533,
+ "epoch": 4.813922356091031,
+ "grad_norm": 0.986640453338623,
+ "learning_rate": 5.6163832331551755e-05,
+ "loss": 0.19539087295532226,
+ "mean_token_accuracy": 0.9345261418819427,
+ "num_tokens": 4243193.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.26934320479631424,
+ "epoch": 4.947791164658635,
+ "grad_norm": 0.8279967308044434,
+ "learning_rate": 5.4097580669801786e-05,
+ "loss": 0.18980844497680663,
+ "mean_token_accuracy": 0.9369515025615692,
+ "num_tokens": 4365873.0,
+ "step": 1850
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.38973631352186205,
+ "eval_loss": 0.6432627439498901,
+ "eval_mean_token_accuracy": 0.8423123425245285,
+ "eval_num_tokens": 4408595.0,
+ "eval_runtime": 97.135,
+ "eval_samples_per_second": 16.462,
+ "eval_steps_per_second": 2.059,
+ "step": 1870
+ },
+ {
+ "entropy": 0.22773028695673653,
+ "epoch": 5.080321285140562,
+ "grad_norm": 0.6459256410598755,
+ "learning_rate": 5.2012280168760034e-05,
+ "loss": 0.14593751907348632,
+ "mean_token_accuracy": 0.9516990624292933,
+ "num_tokens": 4482540.0,
+ "step": 1900
+ },
+ {
+ "entropy": 0.20054464034736155,
+ "epoch": 5.214190093708166,
+ "grad_norm": 0.8127110004425049,
+ "learning_rate": 4.9912471304180415e-05,
+ "loss": 0.1171640968322754,
+ "mean_token_accuracy": 0.9613950234651566,
+ "num_tokens": 4598191.0,
+ "step": 1950
+ },
+ {
+ "entropy": 0.19813392639160157,
+ "epoch": 5.34805890227577,
+ "grad_norm": 0.9870916604995728,
+ "learning_rate": 4.780272614192717e-05,
+ "loss": 0.11783818244934081,
+ "mean_token_accuracy": 0.9606451898813247,
+ "num_tokens": 4716333.0,
+ "step": 2000
+ },
+ {
+ "entropy": 0.20009616125375032,
+ "epoch": 5.481927710843373,
+ "grad_norm": 0.9596051573753357,
+ "learning_rate": 4.568763838288482e-05,
+ "loss": 0.12061534881591797,
+ "mean_token_accuracy": 0.9601768556237221,
+ "num_tokens": 4830262.0,
+ "step": 2050
+ },
+ {
+ "entropy": 0.19717911910265684,
+ "epoch": 5.615796519410977,
+ "grad_norm": 0.9989078044891357,
+ "learning_rate": 4.357181336076072e-05,
+ "loss": 0.11976994514465332,
+ "mean_token_accuracy": 0.9600497630238533,
+ "num_tokens": 4951007.0,
+ "step": 2100
+ },
+ {
+ "entropy": 0.1992201693728566,
+ "epoch": 5.749665327978581,
+ "grad_norm": 0.8984506130218506,
+ "learning_rate": 4.145985801455844e-05,
+ "loss": 0.12019011497497559,
+ "mean_token_accuracy": 0.9582142195105553,
+ "num_tokens": 5067535.0,
+ "step": 2150
+ },
+ {
+ "entropy": 0.1871098166331649,
+ "epoch": 5.883534136546185,
+ "grad_norm": 0.9786492586135864,
+ "learning_rate": 3.9356370857556064e-05,
+ "loss": 0.1156222915649414,
+ "mean_token_accuracy": 0.9616268679499627,
+ "num_tokens": 5185751.0,
+ "step": 2200
+ },
+ {
+ "epoch": 6.0,
+ "eval_entropy": 0.3187944942712784,
+ "eval_loss": 0.7801651358604431,
+ "eval_mean_token_accuracy": 0.837693527340889,
+ "eval_num_tokens": 5290314.0,
+ "eval_runtime": 97.5774,
+ "eval_samples_per_second": 16.387,
+ "eval_steps_per_second": 2.05,
+ "step": 2244
+ },
+ {
+ "entropy": 0.18930522743800673,
+ "epoch": 6.016064257028113,
+ "grad_norm": 0.36866649985313416,
+ "learning_rate": 3.72659319646302e-05,
+ "loss": 0.1124226188659668,
+ "mean_token_accuracy": 0.9628059541938281,
+ "num_tokens": 5303628.0,
+ "step": 2250
+ },
+ {
+ "entropy": 0.14962079100310802,
+ "epoch": 6.149933065595716,
+ "grad_norm": 0.5094484090805054,
+ "learning_rate": 3.519309299972745e-05,
+ "loss": 0.07826479434967042,
+ "mean_token_accuracy": 0.9736301329731941,
+ "num_tokens": 5424681.0,
+ "step": 2300
+ },
+ {
+ "entropy": 0.1441676900163293,
+ "epoch": 6.28380187416332,
+ "grad_norm": 0.7995980978012085,
+ "learning_rate": 3.314236730519691e-05,
+ "loss": 0.07800994396209716,
+ "mean_token_accuracy": 0.9741961327195168,
+ "num_tokens": 5546681.0,
+ "step": 2350
+ },
+ {
+ "entropy": 0.16022401936352254,
+ "epoch": 6.417670682730924,
+ "grad_norm": 0.7507748007774353,
+ "learning_rate": 3.1118220074563075e-05,
+ "loss": 0.0822414779663086,
+ "mean_token_accuracy": 0.9718792167305946,
+ "num_tokens": 5660189.0,
+ "step": 2400
+ },
+ {
+ "entropy": 0.15378836765885354,
+ "epoch": 6.551539491298527,
+ "grad_norm": 0.6187227964401245,
+ "learning_rate": 2.912505863013681e-05,
+ "loss": 0.07948014259338379,
+ "mean_token_accuracy": 0.972696775496006,
+ "num_tokens": 5780047.0,
+ "step": 2450
+ },
+ {
+ "entropy": 0.15355918522924183,
+ "epoch": 6.685408299866131,
+ "grad_norm": 0.8529348373413086,
+ "learning_rate": 2.716722282663332e-05,
+ "loss": 0.08225400924682617,
+ "mean_token_accuracy": 0.9728327831625938,
+ "num_tokens": 5897170.0,
+ "step": 2500
+ },
+ {
+ "entropy": 0.15673237202689053,
+ "epoch": 6.8192771084337345,
+ "grad_norm": 0.7052093148231506,
+ "learning_rate": 2.5248975601692297e-05,
+ "loss": 0.0835498332977295,
+ "mean_token_accuracy": 0.9724654936790467,
+ "num_tokens": 6012704.0,
+ "step": 2550
+ },
+ {
+ "entropy": 0.15278195016086102,
+ "epoch": 6.953145917001339,
+ "grad_norm": 0.5294437408447266,
+ "learning_rate": 2.337449369387515e-05,
+ "loss": 0.08099061012268066,
+ "mean_token_accuracy": 0.9720051056146621,
+ "num_tokens": 6132181.0,
+ "step": 2600
+ },
+ {
+ "epoch": 7.0,
+ "eval_entropy": 0.2793829473108053,
+ "eval_loss": 0.8685178160667419,
+ "eval_mean_token_accuracy": 0.8362820941209793,
+ "eval_num_tokens": 6172033.0,
+ "eval_runtime": 97.4409,
+ "eval_samples_per_second": 16.41,
+ "eval_steps_per_second": 2.053,
+ "step": 2618
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.612589418143744e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2992/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2992/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2992/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2992/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2992/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..aaa5315fd59e264ef47a82c1fe8efb39be4ce81b
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2992/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.0024401823792587042,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2992/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2992/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2992/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2992/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2992/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2992/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2992/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2992/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..a894f4162f929d70e2053e02cc3bdaac75ba9de2
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-2992/trainer_state.json
@@ -0,0 +1,712 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 8.0,
+ "eval_steps": 500,
+ "global_step": 2992,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5728709718585014,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.7528014183044434,
+ "learning_rate": 1.188290252952284e-05,
+ "loss": 1.354185791015625,
+ "mean_token_accuracy": 0.7051774816215038,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.7372898015379906,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.4441661834716797,
+ "learning_rate": 2.40083132739339e-05,
+ "loss": 0.6506095123291016,
+ "mean_token_accuracy": 0.8222023022174835,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6791197884082795,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.9964745044708252,
+ "learning_rate": 3.613372401834496e-05,
+ "loss": 0.6006810760498047,
+ "mean_token_accuracy": 0.8327413862943649,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6392384143173695,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.8044542074203491,
+ "learning_rate": 4.825913476275602e-05,
+ "loss": 0.5597672653198242,
+ "mean_token_accuracy": 0.8428809601068497,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6373909455537796,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.9033862948417664,
+ "learning_rate": 6.038454550716708e-05,
+ "loss": 0.5607868576049805,
+ "mean_token_accuracy": 0.8415500408411026,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6142180925607681,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8247693777084351,
+ "learning_rate": 7.250995625157815e-05,
+ "loss": 0.5359848785400391,
+ "mean_token_accuracy": 0.8485029026865959,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.600434636771679,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9828518033027649,
+ "learning_rate": 8.46353669959892e-05,
+ "loss": 0.5246802139282226,
+ "mean_token_accuracy": 0.8504493471980095,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6199610435962677,
+ "eval_loss": 0.5767685174942017,
+ "eval_mean_token_accuracy": 0.8355379092693329,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 98.478,
+ "eval_samples_per_second": 16.237,
+ "eval_steps_per_second": 2.031,
+ "step": 374
+ },
+ {
+ "entropy": 0.5812384999460645,
+ "epoch": 1.069611780455154,
+ "grad_norm": 0.7832330465316772,
+ "learning_rate": 9.068572797060741e-05,
+ "loss": 0.5035604858398437,
+ "mean_token_accuracy": 0.8548897610168265,
+ "num_tokens": 939587.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.5494279730319976,
+ "epoch": 1.2034805890227578,
+ "grad_norm": 0.6947465538978577,
+ "learning_rate": 9.058701310926409e-05,
+ "loss": 0.4750754928588867,
+ "mean_token_accuracy": 0.8600558358430862,
+ "num_tokens": 1058209.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.5506373339891434,
+ "epoch": 1.3373493975903614,
+ "grad_norm": 0.7013292908668518,
+ "learning_rate": 9.038979832559184e-05,
+ "loss": 0.4743109893798828,
+ "mean_token_accuracy": 0.8615988761186599,
+ "num_tokens": 1178618.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.5413181179761887,
+ "epoch": 1.4712182061579653,
+ "grad_norm": 0.8555615544319153,
+ "learning_rate": 9.009451302961728e-05,
+ "loss": 0.4731967544555664,
+ "mean_token_accuracy": 0.8605538108944892,
+ "num_tokens": 1299124.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.5346147198975086,
+ "epoch": 1.605087014725569,
+ "grad_norm": 0.8122630715370178,
+ "learning_rate": 8.970180016739357e-05,
+ "loss": 0.46506633758544924,
+ "mean_token_accuracy": 0.8624722695350647,
+ "num_tokens": 1422003.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.530957569181919,
+ "epoch": 1.7389558232931726,
+ "grad_norm": 0.6519771218299866,
+ "learning_rate": 8.921251482106754e-05,
+ "loss": 0.46199348449707034,
+ "mean_token_accuracy": 0.8636697196960449,
+ "num_tokens": 1542283.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.5241960515081883,
+ "epoch": 1.8728246318607764,
+ "grad_norm": 0.5876154899597168,
+ "learning_rate": 8.862772234704737e-05,
+ "loss": 0.45739620208740234,
+ "mean_token_accuracy": 0.866628175675869,
+ "num_tokens": 1656994.0,
+ "step": 700
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.5676956498622894,
+ "eval_loss": 0.545108437538147,
+ "eval_mean_token_accuracy": 0.8408122035861015,
+ "eval_num_tokens": 1763438.0,
+ "eval_runtime": 97.562,
+ "eval_samples_per_second": 16.39,
+ "eval_steps_per_second": 2.05,
+ "step": 748
+ },
+ {
+ "entropy": 0.5400724257483627,
+ "epoch": 2.005354752342704,
+ "grad_norm": 0.5679355263710022,
+ "learning_rate": 8.794869605632493e-05,
+ "loss": 0.46216156005859377,
+ "mean_token_accuracy": 0.8634762803111413,
+ "num_tokens": 1768737.0,
+ "step": 750
+ },
+ {
+ "entropy": 0.4396473586559296,
+ "epoch": 2.139223560910308,
+ "grad_norm": 0.8222826719284058,
+ "learning_rate": 8.717691444200361e-05,
+ "loss": 0.36421436309814453,
+ "mean_token_accuracy": 0.8855674511194229,
+ "num_tokens": 1888608.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.4505964604765177,
+ "epoch": 2.2730923694779115,
+ "grad_norm": 0.6411374807357788,
+ "learning_rate": 8.631405796006825e-05,
+ "loss": 0.37386734008789063,
+ "mean_token_accuracy": 0.8835385522246361,
+ "num_tokens": 2000711.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.4520136344432831,
+ "epoch": 2.4069611780455156,
+ "grad_norm": 0.665027379989624,
+ "learning_rate": 8.536200537040663e-05,
+ "loss": 0.37581691741943357,
+ "mean_token_accuracy": 0.884546649158001,
+ "num_tokens": 2117216.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.4466621845960617,
+ "epoch": 2.540829986613119,
+ "grad_norm": 0.6488580107688904,
+ "learning_rate": 8.432282964604958e-05,
+ "loss": 0.37471736907958986,
+ "mean_token_accuracy": 0.8835149678587914,
+ "num_tokens": 2233392.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.4457112967967987,
+ "epoch": 2.674698795180723,
+ "grad_norm": 0.9634214043617249,
+ "learning_rate": 8.31987934595367e-05,
+ "loss": 0.37740741729736327,
+ "mean_token_accuracy": 0.8837782579660416,
+ "num_tokens": 2349527.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.4474602049589157,
+ "epoch": 2.8085676037483265,
+ "grad_norm": 0.6158106923103333,
+ "learning_rate": 8.199234425623558e-05,
+ "loss": 0.3742681121826172,
+ "mean_token_accuracy": 0.8844600400328636,
+ "num_tokens": 2473970.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.45432978719472883,
+ "epoch": 2.9424364123159306,
+ "grad_norm": 0.6572170853614807,
+ "learning_rate": 8.070610892534193e-05,
+ "loss": 0.3823386001586914,
+ "mean_token_accuracy": 0.8826713335514068,
+ "num_tokens": 2593727.0,
+ "step": 1100
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.4939903026819229,
+ "eval_loss": 0.561303436756134,
+ "eval_mean_token_accuracy": 0.8428723356127739,
+ "eval_num_tokens": 2645157.0,
+ "eval_runtime": 97.2032,
+ "eval_samples_per_second": 16.45,
+ "eval_steps_per_second": 2.058,
+ "step": 1122
+ },
+ {
+ "entropy": 0.39069811209584726,
+ "epoch": 3.074966532797858,
+ "grad_norm": 0.8949297070503235,
+ "learning_rate": 7.934288808016343e-05,
+ "loss": 0.31472652435302734,
+ "mean_token_accuracy": 0.9009616712127069,
+ "num_tokens": 2715241.0,
+ "step": 1150
+ },
+ {
+ "entropy": 0.3619236435741186,
+ "epoch": 3.208835341365462,
+ "grad_norm": 0.8773637413978577,
+ "learning_rate": 7.790564996014168e-05,
+ "loss": 0.273685131072998,
+ "mean_token_accuracy": 0.9100899347662925,
+ "num_tokens": 2828619.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.35230451248586175,
+ "epoch": 3.3427041499330654,
+ "grad_norm": 0.6625242829322815,
+ "learning_rate": 7.639752396788952e-05,
+ "loss": 0.27389230728149416,
+ "mean_token_accuracy": 0.9103762644529343,
+ "num_tokens": 2950688.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.35693769574165346,
+ "epoch": 3.4765729585006695,
+ "grad_norm": 0.7137399315834045,
+ "learning_rate": 7.482179385531625e-05,
+ "loss": 0.27805461883544924,
+ "mean_token_accuracy": 0.9091088688373565,
+ "num_tokens": 3071943.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.3644157887250185,
+ "epoch": 3.610441767068273,
+ "grad_norm": 0.9558516144752502,
+ "learning_rate": 7.318189057367674e-05,
+ "loss": 0.2868117141723633,
+ "mean_token_accuracy": 0.9085465380549431,
+ "num_tokens": 3187873.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.3672615347057581,
+ "epoch": 3.7443105756358768,
+ "grad_norm": 0.8229874968528748,
+ "learning_rate": 7.14813848031129e-05,
+ "loss": 0.28742366790771484,
+ "mean_token_accuracy": 0.9074980768561364,
+ "num_tokens": 3299002.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.3580747715383768,
+ "epoch": 3.878179384203481,
+ "grad_norm": 0.8101242184638977,
+ "learning_rate": 6.972397917795341e-05,
+ "loss": 0.28595699310302736,
+ "mean_token_accuracy": 0.9081357708573341,
+ "num_tokens": 3422804.0,
+ "step": 1450
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.4521959410607815,
+ "eval_loss": 0.592692494392395,
+ "eval_mean_token_accuracy": 0.841086029112339,
+ "eval_num_tokens": 3526876.0,
+ "eval_runtime": 96.9985,
+ "eval_samples_per_second": 16.485,
+ "eval_steps_per_second": 2.062,
+ "step": 1496
+ },
+ {
+ "entropy": 0.36107602956319096,
+ "epoch": 4.010709504685408,
+ "grad_norm": 0.8092170357704163,
+ "learning_rate": 6.791350022469971e-05,
+ "loss": 0.28187707901000975,
+ "mean_token_accuracy": 0.9096371344845704,
+ "num_tokens": 3536363.0,
+ "step": 1500
+ },
+ {
+ "entropy": 0.2679470381140709,
+ "epoch": 4.144578313253012,
+ "grad_norm": 1.0743237733840942,
+ "learning_rate": 6.605389003025307e-05,
+ "loss": 0.18246625900268554,
+ "mean_token_accuracy": 0.9399469995498657,
+ "num_tokens": 3653560.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.2658273372799158,
+ "epoch": 4.278447121820616,
+ "grad_norm": 0.9514140486717224,
+ "learning_rate": 6.41491976585231e-05,
+ "loss": 0.18667778015136718,
+ "mean_token_accuracy": 0.9373631486296654,
+ "num_tokens": 3776281.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.267765996158123,
+ "epoch": 4.412315930388219,
+ "grad_norm": 0.8411495089530945,
+ "learning_rate": 6.220357033410804e-05,
+ "loss": 0.1890517807006836,
+ "mean_token_accuracy": 0.9369775268435478,
+ "num_tokens": 3896081.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.27206149250268935,
+ "epoch": 4.546184738955823,
+ "grad_norm": 1.1611251831054688,
+ "learning_rate": 6.022124441224217e-05,
+ "loss": 0.1913918685913086,
+ "mean_token_accuracy": 0.9364263373613357,
+ "num_tokens": 4010469.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.2756531854718924,
+ "epoch": 4.680053547523427,
+ "grad_norm": 0.7768945097923279,
+ "learning_rate": 5.820653615467293e-05,
+ "loss": 0.1946771240234375,
+ "mean_token_accuracy": 0.9357997533679009,
+ "num_tokens": 4125858.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.2766863085329533,
+ "epoch": 4.813922356091031,
+ "grad_norm": 0.986640453338623,
+ "learning_rate": 5.6163832331551755e-05,
+ "loss": 0.19539087295532226,
+ "mean_token_accuracy": 0.9345261418819427,
+ "num_tokens": 4243193.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.26934320479631424,
+ "epoch": 4.947791164658635,
+ "grad_norm": 0.8279967308044434,
+ "learning_rate": 5.4097580669801786e-05,
+ "loss": 0.18980844497680663,
+ "mean_token_accuracy": 0.9369515025615692,
+ "num_tokens": 4365873.0,
+ "step": 1850
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.38973631352186205,
+ "eval_loss": 0.6432627439498901,
+ "eval_mean_token_accuracy": 0.8423123425245285,
+ "eval_num_tokens": 4408595.0,
+ "eval_runtime": 97.135,
+ "eval_samples_per_second": 16.462,
+ "eval_steps_per_second": 2.059,
+ "step": 1870
+ },
+ {
+ "entropy": 0.22773028695673653,
+ "epoch": 5.080321285140562,
+ "grad_norm": 0.6459256410598755,
+ "learning_rate": 5.2012280168760034e-05,
+ "loss": 0.14593751907348632,
+ "mean_token_accuracy": 0.9516990624292933,
+ "num_tokens": 4482540.0,
+ "step": 1900
+ },
+ {
+ "entropy": 0.20054464034736155,
+ "epoch": 5.214190093708166,
+ "grad_norm": 0.8127110004425049,
+ "learning_rate": 4.9912471304180415e-05,
+ "loss": 0.1171640968322754,
+ "mean_token_accuracy": 0.9613950234651566,
+ "num_tokens": 4598191.0,
+ "step": 1950
+ },
+ {
+ "entropy": 0.19813392639160157,
+ "epoch": 5.34805890227577,
+ "grad_norm": 0.9870916604995728,
+ "learning_rate": 4.780272614192717e-05,
+ "loss": 0.11783818244934081,
+ "mean_token_accuracy": 0.9606451898813247,
+ "num_tokens": 4716333.0,
+ "step": 2000
+ },
+ {
+ "entropy": 0.20009616125375032,
+ "epoch": 5.481927710843373,
+ "grad_norm": 0.9596051573753357,
+ "learning_rate": 4.568763838288482e-05,
+ "loss": 0.12061534881591797,
+ "mean_token_accuracy": 0.9601768556237221,
+ "num_tokens": 4830262.0,
+ "step": 2050
+ },
+ {
+ "entropy": 0.19717911910265684,
+ "epoch": 5.615796519410977,
+ "grad_norm": 0.9989078044891357,
+ "learning_rate": 4.357181336076072e-05,
+ "loss": 0.11976994514465332,
+ "mean_token_accuracy": 0.9600497630238533,
+ "num_tokens": 4951007.0,
+ "step": 2100
+ },
+ {
+ "entropy": 0.1992201693728566,
+ "epoch": 5.749665327978581,
+ "grad_norm": 0.8984506130218506,
+ "learning_rate": 4.145985801455844e-05,
+ "loss": 0.12019011497497559,
+ "mean_token_accuracy": 0.9582142195105553,
+ "num_tokens": 5067535.0,
+ "step": 2150
+ },
+ {
+ "entropy": 0.1871098166331649,
+ "epoch": 5.883534136546185,
+ "grad_norm": 0.9786492586135864,
+ "learning_rate": 3.9356370857556064e-05,
+ "loss": 0.1156222915649414,
+ "mean_token_accuracy": 0.9616268679499627,
+ "num_tokens": 5185751.0,
+ "step": 2200
+ },
+ {
+ "epoch": 6.0,
+ "eval_entropy": 0.3187944942712784,
+ "eval_loss": 0.7801651358604431,
+ "eval_mean_token_accuracy": 0.837693527340889,
+ "eval_num_tokens": 5290314.0,
+ "eval_runtime": 97.5774,
+ "eval_samples_per_second": 16.387,
+ "eval_steps_per_second": 2.05,
+ "step": 2244
+ },
+ {
+ "entropy": 0.18930522743800673,
+ "epoch": 6.016064257028113,
+ "grad_norm": 0.36866649985313416,
+ "learning_rate": 3.72659319646302e-05,
+ "loss": 0.1124226188659668,
+ "mean_token_accuracy": 0.9628059541938281,
+ "num_tokens": 5303628.0,
+ "step": 2250
+ },
+ {
+ "entropy": 0.14962079100310802,
+ "epoch": 6.149933065595716,
+ "grad_norm": 0.5094484090805054,
+ "learning_rate": 3.519309299972745e-05,
+ "loss": 0.07826479434967042,
+ "mean_token_accuracy": 0.9736301329731941,
+ "num_tokens": 5424681.0,
+ "step": 2300
+ },
+ {
+ "entropy": 0.1441676900163293,
+ "epoch": 6.28380187416332,
+ "grad_norm": 0.7995980978012085,
+ "learning_rate": 3.314236730519691e-05,
+ "loss": 0.07800994396209716,
+ "mean_token_accuracy": 0.9741961327195168,
+ "num_tokens": 5546681.0,
+ "step": 2350
+ },
+ {
+ "entropy": 0.16022401936352254,
+ "epoch": 6.417670682730924,
+ "grad_norm": 0.7507748007774353,
+ "learning_rate": 3.1118220074563075e-05,
+ "loss": 0.0822414779663086,
+ "mean_token_accuracy": 0.9718792167305946,
+ "num_tokens": 5660189.0,
+ "step": 2400
+ },
+ {
+ "entropy": 0.15378836765885354,
+ "epoch": 6.551539491298527,
+ "grad_norm": 0.6187227964401245,
+ "learning_rate": 2.912505863013681e-05,
+ "loss": 0.07948014259338379,
+ "mean_token_accuracy": 0.972696775496006,
+ "num_tokens": 5780047.0,
+ "step": 2450
+ },
+ {
+ "entropy": 0.15355918522924183,
+ "epoch": 6.685408299866131,
+ "grad_norm": 0.8529348373413086,
+ "learning_rate": 2.716722282663332e-05,
+ "loss": 0.08225400924682617,
+ "mean_token_accuracy": 0.9728327831625938,
+ "num_tokens": 5897170.0,
+ "step": 2500
+ },
+ {
+ "entropy": 0.15673237202689053,
+ "epoch": 6.8192771084337345,
+ "grad_norm": 0.7052093148231506,
+ "learning_rate": 2.5248975601692297e-05,
+ "loss": 0.0835498332977295,
+ "mean_token_accuracy": 0.9724654936790467,
+ "num_tokens": 6012704.0,
+ "step": 2550
+ },
+ {
+ "entropy": 0.15278195016086102,
+ "epoch": 6.953145917001339,
+ "grad_norm": 0.5294437408447266,
+ "learning_rate": 2.337449369387515e-05,
+ "loss": 0.08099061012268066,
+ "mean_token_accuracy": 0.9720051056146621,
+ "num_tokens": 6132181.0,
+ "step": 2600
+ },
+ {
+ "epoch": 7.0,
+ "eval_entropy": 0.2793829473108053,
+ "eval_loss": 0.8685178160667419,
+ "eval_mean_token_accuracy": 0.8362820941209793,
+ "eval_num_tokens": 6172033.0,
+ "eval_runtime": 97.4409,
+ "eval_samples_per_second": 16.41,
+ "eval_steps_per_second": 2.053,
+ "step": 2618
+ },
+ {
+ "entropy": 0.14510226490521672,
+ "epoch": 7.085676037483267,
+ "grad_norm": 0.408651202917099,
+ "learning_rate": 2.154785854834981e-05,
+ "loss": 0.0697246789932251,
+ "mean_token_accuracy": 0.975282986055721,
+ "num_tokens": 6250634.0,
+ "step": 2650
+ },
+ {
+ "entropy": 0.13072173111140728,
+ "epoch": 7.21954484605087,
+ "grad_norm": 0.45500648021698,
+ "learning_rate": 1.9773047430064584e-05,
+ "loss": 0.06559439182281494,
+ "mean_token_accuracy": 0.977269931435585,
+ "num_tokens": 6371882.0,
+ "step": 2700
+ },
+ {
+ "entropy": 0.14414370791986586,
+ "epoch": 7.353413654618474,
+ "grad_norm": 0.3971934914588928,
+ "learning_rate": 1.8053924763761286e-05,
+ "loss": 0.06887433052062988,
+ "mean_token_accuracy": 0.9752715587615967,
+ "num_tokens": 6483611.0,
+ "step": 2750
+ },
+ {
+ "entropy": 0.13290089815855027,
+ "epoch": 7.4872824631860775,
+ "grad_norm": 0.29701998829841614,
+ "learning_rate": 1.639423371968347e-05,
+ "loss": 0.06695588111877442,
+ "mean_token_accuracy": 0.9771433094143868,
+ "num_tokens": 6602543.0,
+ "step": 2800
+ },
+ {
+ "entropy": 0.12676143053919076,
+ "epoch": 7.621151271753681,
+ "grad_norm": 0.40760618448257446,
+ "learning_rate": 1.4797588063300879e-05,
+ "loss": 0.063039231300354,
+ "mean_token_accuracy": 0.9780234199762344,
+ "num_tokens": 6727899.0,
+ "step": 2850
+ },
+ {
+ "entropy": 0.13205574000254272,
+ "epoch": 7.755020080321285,
+ "grad_norm": 0.32809528708457947,
+ "learning_rate": 1.3267464286796153e-05,
+ "loss": 0.06675319194793701,
+ "mean_token_accuracy": 0.977043402493,
+ "num_tokens": 6845842.0,
+ "step": 2900
+ },
+ {
+ "entropy": 0.13014700911939145,
+ "epoch": 7.888888888888889,
+ "grad_norm": 0.28722622990608215,
+ "learning_rate": 1.1807194039446814e-05,
+ "loss": 0.06771249294281007,
+ "mean_token_accuracy": 0.9773513314127922,
+ "num_tokens": 6961910.0,
+ "step": 2950
+ },
+ {
+ "epoch": 8.0,
+ "eval_entropy": 0.2620891720056534,
+ "eval_loss": 0.9525583982467651,
+ "eval_mean_token_accuracy": 0.838210754096508,
+ "eval_num_tokens": 7053752.0,
+ "eval_runtime": 96.9459,
+ "eval_samples_per_second": 16.494,
+ "eval_steps_per_second": 2.063,
+ "step": 2992
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.987999811940086e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3366/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3366/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3366/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3366/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3366/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..aaa5315fd59e264ef47a82c1fe8efb39be4ce81b
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3366/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.0024401823792587042,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3366/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3366/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3366/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3366/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3366/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3366/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3366/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3366/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..151fc8bdb3e994f84d1dcaf88b607ec816d7208e
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3366/trainer_state.json
@@ -0,0 +1,803 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 9.0,
+ "eval_steps": 500,
+ "global_step": 3366,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5728709718585014,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.7528014183044434,
+ "learning_rate": 1.188290252952284e-05,
+ "loss": 1.354185791015625,
+ "mean_token_accuracy": 0.7051774816215038,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.7372898015379906,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.4441661834716797,
+ "learning_rate": 2.40083132739339e-05,
+ "loss": 0.6506095123291016,
+ "mean_token_accuracy": 0.8222023022174835,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6791197884082795,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.9964745044708252,
+ "learning_rate": 3.613372401834496e-05,
+ "loss": 0.6006810760498047,
+ "mean_token_accuracy": 0.8327413862943649,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6392384143173695,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.8044542074203491,
+ "learning_rate": 4.825913476275602e-05,
+ "loss": 0.5597672653198242,
+ "mean_token_accuracy": 0.8428809601068497,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6373909455537796,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.9033862948417664,
+ "learning_rate": 6.038454550716708e-05,
+ "loss": 0.5607868576049805,
+ "mean_token_accuracy": 0.8415500408411026,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6142180925607681,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8247693777084351,
+ "learning_rate": 7.250995625157815e-05,
+ "loss": 0.5359848785400391,
+ "mean_token_accuracy": 0.8485029026865959,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.600434636771679,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9828518033027649,
+ "learning_rate": 8.46353669959892e-05,
+ "loss": 0.5246802139282226,
+ "mean_token_accuracy": 0.8504493471980095,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6199610435962677,
+ "eval_loss": 0.5767685174942017,
+ "eval_mean_token_accuracy": 0.8355379092693329,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 98.478,
+ "eval_samples_per_second": 16.237,
+ "eval_steps_per_second": 2.031,
+ "step": 374
+ },
+ {
+ "entropy": 0.5812384999460645,
+ "epoch": 1.069611780455154,
+ "grad_norm": 0.7832330465316772,
+ "learning_rate": 9.068572797060741e-05,
+ "loss": 0.5035604858398437,
+ "mean_token_accuracy": 0.8548897610168265,
+ "num_tokens": 939587.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.5494279730319976,
+ "epoch": 1.2034805890227578,
+ "grad_norm": 0.6947465538978577,
+ "learning_rate": 9.058701310926409e-05,
+ "loss": 0.4750754928588867,
+ "mean_token_accuracy": 0.8600558358430862,
+ "num_tokens": 1058209.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.5506373339891434,
+ "epoch": 1.3373493975903614,
+ "grad_norm": 0.7013292908668518,
+ "learning_rate": 9.038979832559184e-05,
+ "loss": 0.4743109893798828,
+ "mean_token_accuracy": 0.8615988761186599,
+ "num_tokens": 1178618.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.5413181179761887,
+ "epoch": 1.4712182061579653,
+ "grad_norm": 0.8555615544319153,
+ "learning_rate": 9.009451302961728e-05,
+ "loss": 0.4731967544555664,
+ "mean_token_accuracy": 0.8605538108944892,
+ "num_tokens": 1299124.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.5346147198975086,
+ "epoch": 1.605087014725569,
+ "grad_norm": 0.8122630715370178,
+ "learning_rate": 8.970180016739357e-05,
+ "loss": 0.46506633758544924,
+ "mean_token_accuracy": 0.8624722695350647,
+ "num_tokens": 1422003.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.530957569181919,
+ "epoch": 1.7389558232931726,
+ "grad_norm": 0.6519771218299866,
+ "learning_rate": 8.921251482106754e-05,
+ "loss": 0.46199348449707034,
+ "mean_token_accuracy": 0.8636697196960449,
+ "num_tokens": 1542283.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.5241960515081883,
+ "epoch": 1.8728246318607764,
+ "grad_norm": 0.5876154899597168,
+ "learning_rate": 8.862772234704737e-05,
+ "loss": 0.45739620208740234,
+ "mean_token_accuracy": 0.866628175675869,
+ "num_tokens": 1656994.0,
+ "step": 700
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.5676956498622894,
+ "eval_loss": 0.545108437538147,
+ "eval_mean_token_accuracy": 0.8408122035861015,
+ "eval_num_tokens": 1763438.0,
+ "eval_runtime": 97.562,
+ "eval_samples_per_second": 16.39,
+ "eval_steps_per_second": 2.05,
+ "step": 748
+ },
+ {
+ "entropy": 0.5400724257483627,
+ "epoch": 2.005354752342704,
+ "grad_norm": 0.5679355263710022,
+ "learning_rate": 8.794869605632493e-05,
+ "loss": 0.46216156005859377,
+ "mean_token_accuracy": 0.8634762803111413,
+ "num_tokens": 1768737.0,
+ "step": 750
+ },
+ {
+ "entropy": 0.4396473586559296,
+ "epoch": 2.139223560910308,
+ "grad_norm": 0.8222826719284058,
+ "learning_rate": 8.717691444200361e-05,
+ "loss": 0.36421436309814453,
+ "mean_token_accuracy": 0.8855674511194229,
+ "num_tokens": 1888608.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.4505964604765177,
+ "epoch": 2.2730923694779115,
+ "grad_norm": 0.6411374807357788,
+ "learning_rate": 8.631405796006825e-05,
+ "loss": 0.37386734008789063,
+ "mean_token_accuracy": 0.8835385522246361,
+ "num_tokens": 2000711.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.4520136344432831,
+ "epoch": 2.4069611780455156,
+ "grad_norm": 0.665027379989624,
+ "learning_rate": 8.536200537040663e-05,
+ "loss": 0.37581691741943357,
+ "mean_token_accuracy": 0.884546649158001,
+ "num_tokens": 2117216.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.4466621845960617,
+ "epoch": 2.540829986613119,
+ "grad_norm": 0.6488580107688904,
+ "learning_rate": 8.432282964604958e-05,
+ "loss": 0.37471736907958986,
+ "mean_token_accuracy": 0.8835149678587914,
+ "num_tokens": 2233392.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.4457112967967987,
+ "epoch": 2.674698795180723,
+ "grad_norm": 0.9634214043617249,
+ "learning_rate": 8.31987934595367e-05,
+ "loss": 0.37740741729736327,
+ "mean_token_accuracy": 0.8837782579660416,
+ "num_tokens": 2349527.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.4474602049589157,
+ "epoch": 2.8085676037483265,
+ "grad_norm": 0.6158106923103333,
+ "learning_rate": 8.199234425623558e-05,
+ "loss": 0.3742681121826172,
+ "mean_token_accuracy": 0.8844600400328636,
+ "num_tokens": 2473970.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.45432978719472883,
+ "epoch": 2.9424364123159306,
+ "grad_norm": 0.6572170853614807,
+ "learning_rate": 8.070610892534193e-05,
+ "loss": 0.3823386001586914,
+ "mean_token_accuracy": 0.8826713335514068,
+ "num_tokens": 2593727.0,
+ "step": 1100
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.4939903026819229,
+ "eval_loss": 0.561303436756134,
+ "eval_mean_token_accuracy": 0.8428723356127739,
+ "eval_num_tokens": 2645157.0,
+ "eval_runtime": 97.2032,
+ "eval_samples_per_second": 16.45,
+ "eval_steps_per_second": 2.058,
+ "step": 1122
+ },
+ {
+ "entropy": 0.39069811209584726,
+ "epoch": 3.074966532797858,
+ "grad_norm": 0.8949297070503235,
+ "learning_rate": 7.934288808016343e-05,
+ "loss": 0.31472652435302734,
+ "mean_token_accuracy": 0.9009616712127069,
+ "num_tokens": 2715241.0,
+ "step": 1150
+ },
+ {
+ "entropy": 0.3619236435741186,
+ "epoch": 3.208835341365462,
+ "grad_norm": 0.8773637413978577,
+ "learning_rate": 7.790564996014168e-05,
+ "loss": 0.273685131072998,
+ "mean_token_accuracy": 0.9100899347662925,
+ "num_tokens": 2828619.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.35230451248586175,
+ "epoch": 3.3427041499330654,
+ "grad_norm": 0.6625242829322815,
+ "learning_rate": 7.639752396788952e-05,
+ "loss": 0.27389230728149416,
+ "mean_token_accuracy": 0.9103762644529343,
+ "num_tokens": 2950688.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.35693769574165346,
+ "epoch": 3.4765729585006695,
+ "grad_norm": 0.7137399315834045,
+ "learning_rate": 7.482179385531625e-05,
+ "loss": 0.27805461883544924,
+ "mean_token_accuracy": 0.9091088688373565,
+ "num_tokens": 3071943.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.3644157887250185,
+ "epoch": 3.610441767068273,
+ "grad_norm": 0.9558516144752502,
+ "learning_rate": 7.318189057367674e-05,
+ "loss": 0.2868117141723633,
+ "mean_token_accuracy": 0.9085465380549431,
+ "num_tokens": 3187873.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.3672615347057581,
+ "epoch": 3.7443105756358768,
+ "grad_norm": 0.8229874968528748,
+ "learning_rate": 7.14813848031129e-05,
+ "loss": 0.28742366790771484,
+ "mean_token_accuracy": 0.9074980768561364,
+ "num_tokens": 3299002.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.3580747715383768,
+ "epoch": 3.878179384203481,
+ "grad_norm": 0.8101242184638977,
+ "learning_rate": 6.972397917795341e-05,
+ "loss": 0.28595699310302736,
+ "mean_token_accuracy": 0.9081357708573341,
+ "num_tokens": 3422804.0,
+ "step": 1450
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.4521959410607815,
+ "eval_loss": 0.592692494392395,
+ "eval_mean_token_accuracy": 0.841086029112339,
+ "eval_num_tokens": 3526876.0,
+ "eval_runtime": 96.9985,
+ "eval_samples_per_second": 16.485,
+ "eval_steps_per_second": 2.062,
+ "step": 1496
+ },
+ {
+ "entropy": 0.36107602956319096,
+ "epoch": 4.010709504685408,
+ "grad_norm": 0.8092170357704163,
+ "learning_rate": 6.791350022469971e-05,
+ "loss": 0.28187707901000975,
+ "mean_token_accuracy": 0.9096371344845704,
+ "num_tokens": 3536363.0,
+ "step": 1500
+ },
+ {
+ "entropy": 0.2679470381140709,
+ "epoch": 4.144578313253012,
+ "grad_norm": 1.0743237733840942,
+ "learning_rate": 6.605389003025307e-05,
+ "loss": 0.18246625900268554,
+ "mean_token_accuracy": 0.9399469995498657,
+ "num_tokens": 3653560.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.2658273372799158,
+ "epoch": 4.278447121820616,
+ "grad_norm": 0.9514140486717224,
+ "learning_rate": 6.41491976585231e-05,
+ "loss": 0.18667778015136718,
+ "mean_token_accuracy": 0.9373631486296654,
+ "num_tokens": 3776281.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.267765996158123,
+ "epoch": 4.412315930388219,
+ "grad_norm": 0.8411495089530945,
+ "learning_rate": 6.220357033410804e-05,
+ "loss": 0.1890517807006836,
+ "mean_token_accuracy": 0.9369775268435478,
+ "num_tokens": 3896081.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.27206149250268935,
+ "epoch": 4.546184738955823,
+ "grad_norm": 1.1611251831054688,
+ "learning_rate": 6.022124441224217e-05,
+ "loss": 0.1913918685913086,
+ "mean_token_accuracy": 0.9364263373613357,
+ "num_tokens": 4010469.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.2756531854718924,
+ "epoch": 4.680053547523427,
+ "grad_norm": 0.7768945097923279,
+ "learning_rate": 5.820653615467293e-05,
+ "loss": 0.1946771240234375,
+ "mean_token_accuracy": 0.9357997533679009,
+ "num_tokens": 4125858.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.2766863085329533,
+ "epoch": 4.813922356091031,
+ "grad_norm": 0.986640453338623,
+ "learning_rate": 5.6163832331551755e-05,
+ "loss": 0.19539087295532226,
+ "mean_token_accuracy": 0.9345261418819427,
+ "num_tokens": 4243193.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.26934320479631424,
+ "epoch": 4.947791164658635,
+ "grad_norm": 0.8279967308044434,
+ "learning_rate": 5.4097580669801786e-05,
+ "loss": 0.18980844497680663,
+ "mean_token_accuracy": 0.9369515025615692,
+ "num_tokens": 4365873.0,
+ "step": 1850
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.38973631352186205,
+ "eval_loss": 0.6432627439498901,
+ "eval_mean_token_accuracy": 0.8423123425245285,
+ "eval_num_tokens": 4408595.0,
+ "eval_runtime": 97.135,
+ "eval_samples_per_second": 16.462,
+ "eval_steps_per_second": 2.059,
+ "step": 1870
+ },
+ {
+ "entropy": 0.22773028695673653,
+ "epoch": 5.080321285140562,
+ "grad_norm": 0.6459256410598755,
+ "learning_rate": 5.2012280168760034e-05,
+ "loss": 0.14593751907348632,
+ "mean_token_accuracy": 0.9516990624292933,
+ "num_tokens": 4482540.0,
+ "step": 1900
+ },
+ {
+ "entropy": 0.20054464034736155,
+ "epoch": 5.214190093708166,
+ "grad_norm": 0.8127110004425049,
+ "learning_rate": 4.9912471304180415e-05,
+ "loss": 0.1171640968322754,
+ "mean_token_accuracy": 0.9613950234651566,
+ "num_tokens": 4598191.0,
+ "step": 1950
+ },
+ {
+ "entropy": 0.19813392639160157,
+ "epoch": 5.34805890227577,
+ "grad_norm": 0.9870916604995728,
+ "learning_rate": 4.780272614192717e-05,
+ "loss": 0.11783818244934081,
+ "mean_token_accuracy": 0.9606451898813247,
+ "num_tokens": 4716333.0,
+ "step": 2000
+ },
+ {
+ "entropy": 0.20009616125375032,
+ "epoch": 5.481927710843373,
+ "grad_norm": 0.9596051573753357,
+ "learning_rate": 4.568763838288482e-05,
+ "loss": 0.12061534881591797,
+ "mean_token_accuracy": 0.9601768556237221,
+ "num_tokens": 4830262.0,
+ "step": 2050
+ },
+ {
+ "entropy": 0.19717911910265684,
+ "epoch": 5.615796519410977,
+ "grad_norm": 0.9989078044891357,
+ "learning_rate": 4.357181336076072e-05,
+ "loss": 0.11976994514465332,
+ "mean_token_accuracy": 0.9600497630238533,
+ "num_tokens": 4951007.0,
+ "step": 2100
+ },
+ {
+ "entropy": 0.1992201693728566,
+ "epoch": 5.749665327978581,
+ "grad_norm": 0.8984506130218506,
+ "learning_rate": 4.145985801455844e-05,
+ "loss": 0.12019011497497559,
+ "mean_token_accuracy": 0.9582142195105553,
+ "num_tokens": 5067535.0,
+ "step": 2150
+ },
+ {
+ "entropy": 0.1871098166331649,
+ "epoch": 5.883534136546185,
+ "grad_norm": 0.9786492586135864,
+ "learning_rate": 3.9356370857556064e-05,
+ "loss": 0.1156222915649414,
+ "mean_token_accuracy": 0.9616268679499627,
+ "num_tokens": 5185751.0,
+ "step": 2200
+ },
+ {
+ "epoch": 6.0,
+ "eval_entropy": 0.3187944942712784,
+ "eval_loss": 0.7801651358604431,
+ "eval_mean_token_accuracy": 0.837693527340889,
+ "eval_num_tokens": 5290314.0,
+ "eval_runtime": 97.5774,
+ "eval_samples_per_second": 16.387,
+ "eval_steps_per_second": 2.05,
+ "step": 2244
+ },
+ {
+ "entropy": 0.18930522743800673,
+ "epoch": 6.016064257028113,
+ "grad_norm": 0.36866649985313416,
+ "learning_rate": 3.72659319646302e-05,
+ "loss": 0.1124226188659668,
+ "mean_token_accuracy": 0.9628059541938281,
+ "num_tokens": 5303628.0,
+ "step": 2250
+ },
+ {
+ "entropy": 0.14962079100310802,
+ "epoch": 6.149933065595716,
+ "grad_norm": 0.5094484090805054,
+ "learning_rate": 3.519309299972745e-05,
+ "loss": 0.07826479434967042,
+ "mean_token_accuracy": 0.9736301329731941,
+ "num_tokens": 5424681.0,
+ "step": 2300
+ },
+ {
+ "entropy": 0.1441676900163293,
+ "epoch": 6.28380187416332,
+ "grad_norm": 0.7995980978012085,
+ "learning_rate": 3.314236730519691e-05,
+ "loss": 0.07800994396209716,
+ "mean_token_accuracy": 0.9741961327195168,
+ "num_tokens": 5546681.0,
+ "step": 2350
+ },
+ {
+ "entropy": 0.16022401936352254,
+ "epoch": 6.417670682730924,
+ "grad_norm": 0.7507748007774353,
+ "learning_rate": 3.1118220074563075e-05,
+ "loss": 0.0822414779663086,
+ "mean_token_accuracy": 0.9718792167305946,
+ "num_tokens": 5660189.0,
+ "step": 2400
+ },
+ {
+ "entropy": 0.15378836765885354,
+ "epoch": 6.551539491298527,
+ "grad_norm": 0.6187227964401245,
+ "learning_rate": 2.912505863013681e-05,
+ "loss": 0.07948014259338379,
+ "mean_token_accuracy": 0.972696775496006,
+ "num_tokens": 5780047.0,
+ "step": 2450
+ },
+ {
+ "entropy": 0.15355918522924183,
+ "epoch": 6.685408299866131,
+ "grad_norm": 0.8529348373413086,
+ "learning_rate": 2.716722282663332e-05,
+ "loss": 0.08225400924682617,
+ "mean_token_accuracy": 0.9728327831625938,
+ "num_tokens": 5897170.0,
+ "step": 2500
+ },
+ {
+ "entropy": 0.15673237202689053,
+ "epoch": 6.8192771084337345,
+ "grad_norm": 0.7052093148231506,
+ "learning_rate": 2.5248975601692297e-05,
+ "loss": 0.0835498332977295,
+ "mean_token_accuracy": 0.9724654936790467,
+ "num_tokens": 6012704.0,
+ "step": 2550
+ },
+ {
+ "entropy": 0.15278195016086102,
+ "epoch": 6.953145917001339,
+ "grad_norm": 0.5294437408447266,
+ "learning_rate": 2.337449369387515e-05,
+ "loss": 0.08099061012268066,
+ "mean_token_accuracy": 0.9720051056146621,
+ "num_tokens": 6132181.0,
+ "step": 2600
+ },
+ {
+ "epoch": 7.0,
+ "eval_entropy": 0.2793829473108053,
+ "eval_loss": 0.8685178160667419,
+ "eval_mean_token_accuracy": 0.8362820941209793,
+ "eval_num_tokens": 6172033.0,
+ "eval_runtime": 97.4409,
+ "eval_samples_per_second": 16.41,
+ "eval_steps_per_second": 2.053,
+ "step": 2618
+ },
+ {
+ "entropy": 0.14510226490521672,
+ "epoch": 7.085676037483267,
+ "grad_norm": 0.408651202917099,
+ "learning_rate": 2.154785854834981e-05,
+ "loss": 0.0697246789932251,
+ "mean_token_accuracy": 0.975282986055721,
+ "num_tokens": 6250634.0,
+ "step": 2650
+ },
+ {
+ "entropy": 0.13072173111140728,
+ "epoch": 7.21954484605087,
+ "grad_norm": 0.45500648021698,
+ "learning_rate": 1.9773047430064584e-05,
+ "loss": 0.06559439182281494,
+ "mean_token_accuracy": 0.977269931435585,
+ "num_tokens": 6371882.0,
+ "step": 2700
+ },
+ {
+ "entropy": 0.14414370791986586,
+ "epoch": 7.353413654618474,
+ "grad_norm": 0.3971934914588928,
+ "learning_rate": 1.8053924763761286e-05,
+ "loss": 0.06887433052062988,
+ "mean_token_accuracy": 0.9752715587615967,
+ "num_tokens": 6483611.0,
+ "step": 2750
+ },
+ {
+ "entropy": 0.13290089815855027,
+ "epoch": 7.4872824631860775,
+ "grad_norm": 0.29701998829841614,
+ "learning_rate": 1.639423371968347e-05,
+ "loss": 0.06695588111877442,
+ "mean_token_accuracy": 0.9771433094143868,
+ "num_tokens": 6602543.0,
+ "step": 2800
+ },
+ {
+ "entropy": 0.12676143053919076,
+ "epoch": 7.621151271753681,
+ "grad_norm": 0.40760618448257446,
+ "learning_rate": 1.4797588063300879e-05,
+ "loss": 0.063039231300354,
+ "mean_token_accuracy": 0.9780234199762344,
+ "num_tokens": 6727899.0,
+ "step": 2850
+ },
+ {
+ "entropy": 0.13205574000254272,
+ "epoch": 7.755020080321285,
+ "grad_norm": 0.32809528708457947,
+ "learning_rate": 1.3267464286796153e-05,
+ "loss": 0.06675319194793701,
+ "mean_token_accuracy": 0.977043402493,
+ "num_tokens": 6845842.0,
+ "step": 2900
+ },
+ {
+ "entropy": 0.13014700911939145,
+ "epoch": 7.888888888888889,
+ "grad_norm": 0.28722622990608215,
+ "learning_rate": 1.1807194039446814e-05,
+ "loss": 0.06771249294281007,
+ "mean_token_accuracy": 0.9773513314127922,
+ "num_tokens": 6961910.0,
+ "step": 2950
+ },
+ {
+ "epoch": 8.0,
+ "eval_entropy": 0.2620891720056534,
+ "eval_loss": 0.9525583982467651,
+ "eval_mean_token_accuracy": 0.838210754096508,
+ "eval_num_tokens": 7053752.0,
+ "eval_runtime": 96.9459,
+ "eval_samples_per_second": 16.494,
+ "eval_steps_per_second": 2.063,
+ "step": 2992
+ },
+ {
+ "entropy": 0.13656818576984936,
+ "epoch": 8.021419009370817,
+ "grad_norm": 0.18831250071525574,
+ "learning_rate": 1.0419956873384076e-05,
+ "loss": 0.06974420547485352,
+ "mean_token_accuracy": 0.9758493298231953,
+ "num_tokens": 7073570.0,
+ "step": 3000
+ },
+ {
+ "entropy": 0.13666721584275365,
+ "epoch": 8.15528781793842,
+ "grad_norm": 0.28443825244903564,
+ "learning_rate": 9.108773320523731e-06,
+ "loss": 0.063851957321167,
+ "mean_token_accuracy": 0.9773794823884964,
+ "num_tokens": 7186955.0,
+ "step": 3050
+ },
+ {
+ "entropy": 0.12880131481215357,
+ "epoch": 8.289156626506024,
+ "grad_norm": 0.19322320818901062,
+ "learning_rate": 7.87649831574337e-06,
+ "loss": 0.05904998779296875,
+ "mean_token_accuracy": 0.9780682101845741,
+ "num_tokens": 7307543.0,
+ "step": 3100
+ },
+ {
+ "entropy": 0.12289242332801223,
+ "epoch": 8.423025435073628,
+ "grad_norm": 0.22629190981388092,
+ "learning_rate": 6.725814980625996e-06,
+ "loss": 0.060590991973876955,
+ "mean_token_accuracy": 0.9784254813194275,
+ "num_tokens": 7425896.0,
+ "step": 3150
+ },
+ {
+ "entropy": 0.12430635459721089,
+ "epoch": 8.556894243641231,
+ "grad_norm": 0.29989469051361084,
+ "learning_rate": 5.659228781305109e-06,
+ "loss": 0.06014200210571289,
+ "mean_token_accuracy": 0.9776277858018875,
+ "num_tokens": 7546590.0,
+ "step": 3200
+ },
+ {
+ "entropy": 0.12310879753902554,
+ "epoch": 8.690763052208835,
+ "grad_norm": 0.13988597691059113,
+ "learning_rate": 4.6790620731319124e-06,
+ "loss": 0.060741419792175295,
+ "mean_token_accuracy": 0.9787649729847908,
+ "num_tokens": 7666771.0,
+ "step": 3250
+ },
+ {
+ "entropy": 0.13028676200658082,
+ "epoch": 8.824631860776439,
+ "grad_norm": 0.19891025125980377,
+ "learning_rate": 3.787449044042906e-06,
+ "loss": 0.06288023948669434,
+ "mean_token_accuracy": 0.9772285357117653,
+ "num_tokens": 7782159.0,
+ "step": 3300
+ },
+ {
+ "entropy": 0.13001586498692633,
+ "epoch": 8.958500669344042,
+ "grad_norm": 0.1798337697982788,
+ "learning_rate": 2.9863310676379143e-06,
+ "loss": 0.06166903972625733,
+ "mean_token_accuracy": 0.9771321558952332,
+ "num_tokens": 7900010.0,
+ "step": 3350
+ },
+ {
+ "epoch": 9.0,
+ "eval_entropy": 0.24724020630121232,
+ "eval_loss": 1.0340324640274048,
+ "eval_mean_token_accuracy": 0.837886828482151,
+ "eval_num_tokens": 7935471.0,
+ "eval_runtime": 97.1913,
+ "eval_samples_per_second": 16.452,
+ "eval_steps_per_second": 2.058,
+ "step": 3366
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 3.3602237926968115e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-374/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-374/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-374/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-374/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-374/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..aaa5315fd59e264ef47a82c1fe8efb39be4ce81b
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-374/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.0024401823792587042,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-374/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-374/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-374/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-374/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-374/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-374/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-374/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-374/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..a9367bfad57d97c6648aacdcdf19a5ef538d3fd5
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-374/trainer_state.json
@@ -0,0 +1,115 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 1.0,
+ "eval_steps": 500,
+ "global_step": 374,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5728709718585014,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.7528014183044434,
+ "learning_rate": 1.188290252952284e-05,
+ "loss": 1.354185791015625,
+ "mean_token_accuracy": 0.7051774816215038,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.7372898015379906,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.4441661834716797,
+ "learning_rate": 2.40083132739339e-05,
+ "loss": 0.6506095123291016,
+ "mean_token_accuracy": 0.8222023022174835,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6791197884082795,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.9964745044708252,
+ "learning_rate": 3.613372401834496e-05,
+ "loss": 0.6006810760498047,
+ "mean_token_accuracy": 0.8327413862943649,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6392384143173695,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.8044542074203491,
+ "learning_rate": 4.825913476275602e-05,
+ "loss": 0.5597672653198242,
+ "mean_token_accuracy": 0.8428809601068497,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6373909455537796,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.9033862948417664,
+ "learning_rate": 6.038454550716708e-05,
+ "loss": 0.5607868576049805,
+ "mean_token_accuracy": 0.8415500408411026,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6142180925607681,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8247693777084351,
+ "learning_rate": 7.250995625157815e-05,
+ "loss": 0.5359848785400391,
+ "mean_token_accuracy": 0.8485029026865959,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.600434636771679,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9828518033027649,
+ "learning_rate": 8.46353669959892e-05,
+ "loss": 0.5246802139282226,
+ "mean_token_accuracy": 0.8504493471980095,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6199610435962677,
+ "eval_loss": 0.5767685174942017,
+ "eval_mean_token_accuracy": 0.8355379092693329,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 98.478,
+ "eval_samples_per_second": 16.237,
+ "eval_steps_per_second": 2.031,
+ "step": 374
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 3.709458252163277e+16,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3740/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3740/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3740/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3740/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3740/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..aaa5315fd59e264ef47a82c1fe8efb39be4ce81b
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3740/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.0024401823792587042,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3740/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3740/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3740/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3740/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3740/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3740/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3740/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3740/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..3ad613c29fd11918789fe794ab3ae8d7a46ec76f
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-3740/trainer_state.json
@@ -0,0 +1,884 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 10.0,
+ "eval_steps": 500,
+ "global_step": 3740,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5728709718585014,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.7528014183044434,
+ "learning_rate": 1.188290252952284e-05,
+ "loss": 1.354185791015625,
+ "mean_token_accuracy": 0.7051774816215038,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.7372898015379906,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.4441661834716797,
+ "learning_rate": 2.40083132739339e-05,
+ "loss": 0.6506095123291016,
+ "mean_token_accuracy": 0.8222023022174835,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6791197884082795,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.9964745044708252,
+ "learning_rate": 3.613372401834496e-05,
+ "loss": 0.6006810760498047,
+ "mean_token_accuracy": 0.8327413862943649,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6392384143173695,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.8044542074203491,
+ "learning_rate": 4.825913476275602e-05,
+ "loss": 0.5597672653198242,
+ "mean_token_accuracy": 0.8428809601068497,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6373909455537796,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.9033862948417664,
+ "learning_rate": 6.038454550716708e-05,
+ "loss": 0.5607868576049805,
+ "mean_token_accuracy": 0.8415500408411026,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6142180925607681,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8247693777084351,
+ "learning_rate": 7.250995625157815e-05,
+ "loss": 0.5359848785400391,
+ "mean_token_accuracy": 0.8485029026865959,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.600434636771679,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9828518033027649,
+ "learning_rate": 8.46353669959892e-05,
+ "loss": 0.5246802139282226,
+ "mean_token_accuracy": 0.8504493471980095,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6199610435962677,
+ "eval_loss": 0.5767685174942017,
+ "eval_mean_token_accuracy": 0.8355379092693329,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 98.478,
+ "eval_samples_per_second": 16.237,
+ "eval_steps_per_second": 2.031,
+ "step": 374
+ },
+ {
+ "entropy": 0.5812384999460645,
+ "epoch": 1.069611780455154,
+ "grad_norm": 0.7832330465316772,
+ "learning_rate": 9.068572797060741e-05,
+ "loss": 0.5035604858398437,
+ "mean_token_accuracy": 0.8548897610168265,
+ "num_tokens": 939587.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.5494279730319976,
+ "epoch": 1.2034805890227578,
+ "grad_norm": 0.6947465538978577,
+ "learning_rate": 9.058701310926409e-05,
+ "loss": 0.4750754928588867,
+ "mean_token_accuracy": 0.8600558358430862,
+ "num_tokens": 1058209.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.5506373339891434,
+ "epoch": 1.3373493975903614,
+ "grad_norm": 0.7013292908668518,
+ "learning_rate": 9.038979832559184e-05,
+ "loss": 0.4743109893798828,
+ "mean_token_accuracy": 0.8615988761186599,
+ "num_tokens": 1178618.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.5413181179761887,
+ "epoch": 1.4712182061579653,
+ "grad_norm": 0.8555615544319153,
+ "learning_rate": 9.009451302961728e-05,
+ "loss": 0.4731967544555664,
+ "mean_token_accuracy": 0.8605538108944892,
+ "num_tokens": 1299124.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.5346147198975086,
+ "epoch": 1.605087014725569,
+ "grad_norm": 0.8122630715370178,
+ "learning_rate": 8.970180016739357e-05,
+ "loss": 0.46506633758544924,
+ "mean_token_accuracy": 0.8624722695350647,
+ "num_tokens": 1422003.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.530957569181919,
+ "epoch": 1.7389558232931726,
+ "grad_norm": 0.6519771218299866,
+ "learning_rate": 8.921251482106754e-05,
+ "loss": 0.46199348449707034,
+ "mean_token_accuracy": 0.8636697196960449,
+ "num_tokens": 1542283.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.5241960515081883,
+ "epoch": 1.8728246318607764,
+ "grad_norm": 0.5876154899597168,
+ "learning_rate": 8.862772234704737e-05,
+ "loss": 0.45739620208740234,
+ "mean_token_accuracy": 0.866628175675869,
+ "num_tokens": 1656994.0,
+ "step": 700
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.5676956498622894,
+ "eval_loss": 0.545108437538147,
+ "eval_mean_token_accuracy": 0.8408122035861015,
+ "eval_num_tokens": 1763438.0,
+ "eval_runtime": 97.562,
+ "eval_samples_per_second": 16.39,
+ "eval_steps_per_second": 2.05,
+ "step": 748
+ },
+ {
+ "entropy": 0.5400724257483627,
+ "epoch": 2.005354752342704,
+ "grad_norm": 0.5679355263710022,
+ "learning_rate": 8.794869605632493e-05,
+ "loss": 0.46216156005859377,
+ "mean_token_accuracy": 0.8634762803111413,
+ "num_tokens": 1768737.0,
+ "step": 750
+ },
+ {
+ "entropy": 0.4396473586559296,
+ "epoch": 2.139223560910308,
+ "grad_norm": 0.8222826719284058,
+ "learning_rate": 8.717691444200361e-05,
+ "loss": 0.36421436309814453,
+ "mean_token_accuracy": 0.8855674511194229,
+ "num_tokens": 1888608.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.4505964604765177,
+ "epoch": 2.2730923694779115,
+ "grad_norm": 0.6411374807357788,
+ "learning_rate": 8.631405796006825e-05,
+ "loss": 0.37386734008789063,
+ "mean_token_accuracy": 0.8835385522246361,
+ "num_tokens": 2000711.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.4520136344432831,
+ "epoch": 2.4069611780455156,
+ "grad_norm": 0.665027379989624,
+ "learning_rate": 8.536200537040663e-05,
+ "loss": 0.37581691741943357,
+ "mean_token_accuracy": 0.884546649158001,
+ "num_tokens": 2117216.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.4466621845960617,
+ "epoch": 2.540829986613119,
+ "grad_norm": 0.6488580107688904,
+ "learning_rate": 8.432282964604958e-05,
+ "loss": 0.37471736907958986,
+ "mean_token_accuracy": 0.8835149678587914,
+ "num_tokens": 2233392.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.4457112967967987,
+ "epoch": 2.674698795180723,
+ "grad_norm": 0.9634214043617249,
+ "learning_rate": 8.31987934595367e-05,
+ "loss": 0.37740741729736327,
+ "mean_token_accuracy": 0.8837782579660416,
+ "num_tokens": 2349527.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.4474602049589157,
+ "epoch": 2.8085676037483265,
+ "grad_norm": 0.6158106923103333,
+ "learning_rate": 8.199234425623558e-05,
+ "loss": 0.3742681121826172,
+ "mean_token_accuracy": 0.8844600400328636,
+ "num_tokens": 2473970.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.45432978719472883,
+ "epoch": 2.9424364123159306,
+ "grad_norm": 0.6572170853614807,
+ "learning_rate": 8.070610892534193e-05,
+ "loss": 0.3823386001586914,
+ "mean_token_accuracy": 0.8826713335514068,
+ "num_tokens": 2593727.0,
+ "step": 1100
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.4939903026819229,
+ "eval_loss": 0.561303436756134,
+ "eval_mean_token_accuracy": 0.8428723356127739,
+ "eval_num_tokens": 2645157.0,
+ "eval_runtime": 97.2032,
+ "eval_samples_per_second": 16.45,
+ "eval_steps_per_second": 2.058,
+ "step": 1122
+ },
+ {
+ "entropy": 0.39069811209584726,
+ "epoch": 3.074966532797858,
+ "grad_norm": 0.8949297070503235,
+ "learning_rate": 7.934288808016343e-05,
+ "loss": 0.31472652435302734,
+ "mean_token_accuracy": 0.9009616712127069,
+ "num_tokens": 2715241.0,
+ "step": 1150
+ },
+ {
+ "entropy": 0.3619236435741186,
+ "epoch": 3.208835341365462,
+ "grad_norm": 0.8773637413978577,
+ "learning_rate": 7.790564996014168e-05,
+ "loss": 0.273685131072998,
+ "mean_token_accuracy": 0.9100899347662925,
+ "num_tokens": 2828619.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.35230451248586175,
+ "epoch": 3.3427041499330654,
+ "grad_norm": 0.6625242829322815,
+ "learning_rate": 7.639752396788952e-05,
+ "loss": 0.27389230728149416,
+ "mean_token_accuracy": 0.9103762644529343,
+ "num_tokens": 2950688.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.35693769574165346,
+ "epoch": 3.4765729585006695,
+ "grad_norm": 0.7137399315834045,
+ "learning_rate": 7.482179385531625e-05,
+ "loss": 0.27805461883544924,
+ "mean_token_accuracy": 0.9091088688373565,
+ "num_tokens": 3071943.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.3644157887250185,
+ "epoch": 3.610441767068273,
+ "grad_norm": 0.9558516144752502,
+ "learning_rate": 7.318189057367674e-05,
+ "loss": 0.2868117141723633,
+ "mean_token_accuracy": 0.9085465380549431,
+ "num_tokens": 3187873.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.3672615347057581,
+ "epoch": 3.7443105756358768,
+ "grad_norm": 0.8229874968528748,
+ "learning_rate": 7.14813848031129e-05,
+ "loss": 0.28742366790771484,
+ "mean_token_accuracy": 0.9074980768561364,
+ "num_tokens": 3299002.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.3580747715383768,
+ "epoch": 3.878179384203481,
+ "grad_norm": 0.8101242184638977,
+ "learning_rate": 6.972397917795341e-05,
+ "loss": 0.28595699310302736,
+ "mean_token_accuracy": 0.9081357708573341,
+ "num_tokens": 3422804.0,
+ "step": 1450
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.4521959410607815,
+ "eval_loss": 0.592692494392395,
+ "eval_mean_token_accuracy": 0.841086029112339,
+ "eval_num_tokens": 3526876.0,
+ "eval_runtime": 96.9985,
+ "eval_samples_per_second": 16.485,
+ "eval_steps_per_second": 2.062,
+ "step": 1496
+ },
+ {
+ "entropy": 0.36107602956319096,
+ "epoch": 4.010709504685408,
+ "grad_norm": 0.8092170357704163,
+ "learning_rate": 6.791350022469971e-05,
+ "loss": 0.28187707901000975,
+ "mean_token_accuracy": 0.9096371344845704,
+ "num_tokens": 3536363.0,
+ "step": 1500
+ },
+ {
+ "entropy": 0.2679470381140709,
+ "epoch": 4.144578313253012,
+ "grad_norm": 1.0743237733840942,
+ "learning_rate": 6.605389003025307e-05,
+ "loss": 0.18246625900268554,
+ "mean_token_accuracy": 0.9399469995498657,
+ "num_tokens": 3653560.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.2658273372799158,
+ "epoch": 4.278447121820616,
+ "grad_norm": 0.9514140486717224,
+ "learning_rate": 6.41491976585231e-05,
+ "loss": 0.18667778015136718,
+ "mean_token_accuracy": 0.9373631486296654,
+ "num_tokens": 3776281.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.267765996158123,
+ "epoch": 4.412315930388219,
+ "grad_norm": 0.8411495089530945,
+ "learning_rate": 6.220357033410804e-05,
+ "loss": 0.1890517807006836,
+ "mean_token_accuracy": 0.9369775268435478,
+ "num_tokens": 3896081.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.27206149250268935,
+ "epoch": 4.546184738955823,
+ "grad_norm": 1.1611251831054688,
+ "learning_rate": 6.022124441224217e-05,
+ "loss": 0.1913918685913086,
+ "mean_token_accuracy": 0.9364263373613357,
+ "num_tokens": 4010469.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.2756531854718924,
+ "epoch": 4.680053547523427,
+ "grad_norm": 0.7768945097923279,
+ "learning_rate": 5.820653615467293e-05,
+ "loss": 0.1946771240234375,
+ "mean_token_accuracy": 0.9357997533679009,
+ "num_tokens": 4125858.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.2766863085329533,
+ "epoch": 4.813922356091031,
+ "grad_norm": 0.986640453338623,
+ "learning_rate": 5.6163832331551755e-05,
+ "loss": 0.19539087295532226,
+ "mean_token_accuracy": 0.9345261418819427,
+ "num_tokens": 4243193.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.26934320479631424,
+ "epoch": 4.947791164658635,
+ "grad_norm": 0.8279967308044434,
+ "learning_rate": 5.4097580669801786e-05,
+ "loss": 0.18980844497680663,
+ "mean_token_accuracy": 0.9369515025615692,
+ "num_tokens": 4365873.0,
+ "step": 1850
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.38973631352186205,
+ "eval_loss": 0.6432627439498901,
+ "eval_mean_token_accuracy": 0.8423123425245285,
+ "eval_num_tokens": 4408595.0,
+ "eval_runtime": 97.135,
+ "eval_samples_per_second": 16.462,
+ "eval_steps_per_second": 2.059,
+ "step": 1870
+ },
+ {
+ "entropy": 0.22773028695673653,
+ "epoch": 5.080321285140562,
+ "grad_norm": 0.6459256410598755,
+ "learning_rate": 5.2012280168760034e-05,
+ "loss": 0.14593751907348632,
+ "mean_token_accuracy": 0.9516990624292933,
+ "num_tokens": 4482540.0,
+ "step": 1900
+ },
+ {
+ "entropy": 0.20054464034736155,
+ "epoch": 5.214190093708166,
+ "grad_norm": 0.8127110004425049,
+ "learning_rate": 4.9912471304180415e-05,
+ "loss": 0.1171640968322754,
+ "mean_token_accuracy": 0.9613950234651566,
+ "num_tokens": 4598191.0,
+ "step": 1950
+ },
+ {
+ "entropy": 0.19813392639160157,
+ "epoch": 5.34805890227577,
+ "grad_norm": 0.9870916604995728,
+ "learning_rate": 4.780272614192717e-05,
+ "loss": 0.11783818244934081,
+ "mean_token_accuracy": 0.9606451898813247,
+ "num_tokens": 4716333.0,
+ "step": 2000
+ },
+ {
+ "entropy": 0.20009616125375032,
+ "epoch": 5.481927710843373,
+ "grad_norm": 0.9596051573753357,
+ "learning_rate": 4.568763838288482e-05,
+ "loss": 0.12061534881591797,
+ "mean_token_accuracy": 0.9601768556237221,
+ "num_tokens": 4830262.0,
+ "step": 2050
+ },
+ {
+ "entropy": 0.19717911910265684,
+ "epoch": 5.615796519410977,
+ "grad_norm": 0.9989078044891357,
+ "learning_rate": 4.357181336076072e-05,
+ "loss": 0.11976994514465332,
+ "mean_token_accuracy": 0.9600497630238533,
+ "num_tokens": 4951007.0,
+ "step": 2100
+ },
+ {
+ "entropy": 0.1992201693728566,
+ "epoch": 5.749665327978581,
+ "grad_norm": 0.8984506130218506,
+ "learning_rate": 4.145985801455844e-05,
+ "loss": 0.12019011497497559,
+ "mean_token_accuracy": 0.9582142195105553,
+ "num_tokens": 5067535.0,
+ "step": 2150
+ },
+ {
+ "entropy": 0.1871098166331649,
+ "epoch": 5.883534136546185,
+ "grad_norm": 0.9786492586135864,
+ "learning_rate": 3.9356370857556064e-05,
+ "loss": 0.1156222915649414,
+ "mean_token_accuracy": 0.9616268679499627,
+ "num_tokens": 5185751.0,
+ "step": 2200
+ },
+ {
+ "epoch": 6.0,
+ "eval_entropy": 0.3187944942712784,
+ "eval_loss": 0.7801651358604431,
+ "eval_mean_token_accuracy": 0.837693527340889,
+ "eval_num_tokens": 5290314.0,
+ "eval_runtime": 97.5774,
+ "eval_samples_per_second": 16.387,
+ "eval_steps_per_second": 2.05,
+ "step": 2244
+ },
+ {
+ "entropy": 0.18930522743800673,
+ "epoch": 6.016064257028113,
+ "grad_norm": 0.36866649985313416,
+ "learning_rate": 3.72659319646302e-05,
+ "loss": 0.1124226188659668,
+ "mean_token_accuracy": 0.9628059541938281,
+ "num_tokens": 5303628.0,
+ "step": 2250
+ },
+ {
+ "entropy": 0.14962079100310802,
+ "epoch": 6.149933065595716,
+ "grad_norm": 0.5094484090805054,
+ "learning_rate": 3.519309299972745e-05,
+ "loss": 0.07826479434967042,
+ "mean_token_accuracy": 0.9736301329731941,
+ "num_tokens": 5424681.0,
+ "step": 2300
+ },
+ {
+ "entropy": 0.1441676900163293,
+ "epoch": 6.28380187416332,
+ "grad_norm": 0.7995980978012085,
+ "learning_rate": 3.314236730519691e-05,
+ "loss": 0.07800994396209716,
+ "mean_token_accuracy": 0.9741961327195168,
+ "num_tokens": 5546681.0,
+ "step": 2350
+ },
+ {
+ "entropy": 0.16022401936352254,
+ "epoch": 6.417670682730924,
+ "grad_norm": 0.7507748007774353,
+ "learning_rate": 3.1118220074563075e-05,
+ "loss": 0.0822414779663086,
+ "mean_token_accuracy": 0.9718792167305946,
+ "num_tokens": 5660189.0,
+ "step": 2400
+ },
+ {
+ "entropy": 0.15378836765885354,
+ "epoch": 6.551539491298527,
+ "grad_norm": 0.6187227964401245,
+ "learning_rate": 2.912505863013681e-05,
+ "loss": 0.07948014259338379,
+ "mean_token_accuracy": 0.972696775496006,
+ "num_tokens": 5780047.0,
+ "step": 2450
+ },
+ {
+ "entropy": 0.15355918522924183,
+ "epoch": 6.685408299866131,
+ "grad_norm": 0.8529348373413086,
+ "learning_rate": 2.716722282663332e-05,
+ "loss": 0.08225400924682617,
+ "mean_token_accuracy": 0.9728327831625938,
+ "num_tokens": 5897170.0,
+ "step": 2500
+ },
+ {
+ "entropy": 0.15673237202689053,
+ "epoch": 6.8192771084337345,
+ "grad_norm": 0.7052093148231506,
+ "learning_rate": 2.5248975601692297e-05,
+ "loss": 0.0835498332977295,
+ "mean_token_accuracy": 0.9724654936790467,
+ "num_tokens": 6012704.0,
+ "step": 2550
+ },
+ {
+ "entropy": 0.15278195016086102,
+ "epoch": 6.953145917001339,
+ "grad_norm": 0.5294437408447266,
+ "learning_rate": 2.337449369387515e-05,
+ "loss": 0.08099061012268066,
+ "mean_token_accuracy": 0.9720051056146621,
+ "num_tokens": 6132181.0,
+ "step": 2600
+ },
+ {
+ "epoch": 7.0,
+ "eval_entropy": 0.2793829473108053,
+ "eval_loss": 0.8685178160667419,
+ "eval_mean_token_accuracy": 0.8362820941209793,
+ "eval_num_tokens": 6172033.0,
+ "eval_runtime": 97.4409,
+ "eval_samples_per_second": 16.41,
+ "eval_steps_per_second": 2.053,
+ "step": 2618
+ },
+ {
+ "entropy": 0.14510226490521672,
+ "epoch": 7.085676037483267,
+ "grad_norm": 0.408651202917099,
+ "learning_rate": 2.154785854834981e-05,
+ "loss": 0.0697246789932251,
+ "mean_token_accuracy": 0.975282986055721,
+ "num_tokens": 6250634.0,
+ "step": 2650
+ },
+ {
+ "entropy": 0.13072173111140728,
+ "epoch": 7.21954484605087,
+ "grad_norm": 0.45500648021698,
+ "learning_rate": 1.9773047430064584e-05,
+ "loss": 0.06559439182281494,
+ "mean_token_accuracy": 0.977269931435585,
+ "num_tokens": 6371882.0,
+ "step": 2700
+ },
+ {
+ "entropy": 0.14414370791986586,
+ "epoch": 7.353413654618474,
+ "grad_norm": 0.3971934914588928,
+ "learning_rate": 1.8053924763761286e-05,
+ "loss": 0.06887433052062988,
+ "mean_token_accuracy": 0.9752715587615967,
+ "num_tokens": 6483611.0,
+ "step": 2750
+ },
+ {
+ "entropy": 0.13290089815855027,
+ "epoch": 7.4872824631860775,
+ "grad_norm": 0.29701998829841614,
+ "learning_rate": 1.639423371968347e-05,
+ "loss": 0.06695588111877442,
+ "mean_token_accuracy": 0.9771433094143868,
+ "num_tokens": 6602543.0,
+ "step": 2800
+ },
+ {
+ "entropy": 0.12676143053919076,
+ "epoch": 7.621151271753681,
+ "grad_norm": 0.40760618448257446,
+ "learning_rate": 1.4797588063300879e-05,
+ "loss": 0.063039231300354,
+ "mean_token_accuracy": 0.9780234199762344,
+ "num_tokens": 6727899.0,
+ "step": 2850
+ },
+ {
+ "entropy": 0.13205574000254272,
+ "epoch": 7.755020080321285,
+ "grad_norm": 0.32809528708457947,
+ "learning_rate": 1.3267464286796153e-05,
+ "loss": 0.06675319194793701,
+ "mean_token_accuracy": 0.977043402493,
+ "num_tokens": 6845842.0,
+ "step": 2900
+ },
+ {
+ "entropy": 0.13014700911939145,
+ "epoch": 7.888888888888889,
+ "grad_norm": 0.28722622990608215,
+ "learning_rate": 1.1807194039446814e-05,
+ "loss": 0.06771249294281007,
+ "mean_token_accuracy": 0.9773513314127922,
+ "num_tokens": 6961910.0,
+ "step": 2950
+ },
+ {
+ "epoch": 8.0,
+ "eval_entropy": 0.2620891720056534,
+ "eval_loss": 0.9525583982467651,
+ "eval_mean_token_accuracy": 0.838210754096508,
+ "eval_num_tokens": 7053752.0,
+ "eval_runtime": 96.9459,
+ "eval_samples_per_second": 16.494,
+ "eval_steps_per_second": 2.063,
+ "step": 2992
+ },
+ {
+ "entropy": 0.13656818576984936,
+ "epoch": 8.021419009370817,
+ "grad_norm": 0.18831250071525574,
+ "learning_rate": 1.0419956873384076e-05,
+ "loss": 0.06974420547485352,
+ "mean_token_accuracy": 0.9758493298231953,
+ "num_tokens": 7073570.0,
+ "step": 3000
+ },
+ {
+ "entropy": 0.13666721584275365,
+ "epoch": 8.15528781793842,
+ "grad_norm": 0.28443825244903564,
+ "learning_rate": 9.108773320523731e-06,
+ "loss": 0.063851957321167,
+ "mean_token_accuracy": 0.9773794823884964,
+ "num_tokens": 7186955.0,
+ "step": 3050
+ },
+ {
+ "entropy": 0.12880131481215357,
+ "epoch": 8.289156626506024,
+ "grad_norm": 0.19322320818901062,
+ "learning_rate": 7.87649831574337e-06,
+ "loss": 0.05904998779296875,
+ "mean_token_accuracy": 0.9780682101845741,
+ "num_tokens": 7307543.0,
+ "step": 3100
+ },
+ {
+ "entropy": 0.12289242332801223,
+ "epoch": 8.423025435073628,
+ "grad_norm": 0.22629190981388092,
+ "learning_rate": 6.725814980625996e-06,
+ "loss": 0.060590991973876955,
+ "mean_token_accuracy": 0.9784254813194275,
+ "num_tokens": 7425896.0,
+ "step": 3150
+ },
+ {
+ "entropy": 0.12430635459721089,
+ "epoch": 8.556894243641231,
+ "grad_norm": 0.29989469051361084,
+ "learning_rate": 5.659228781305109e-06,
+ "loss": 0.06014200210571289,
+ "mean_token_accuracy": 0.9776277858018875,
+ "num_tokens": 7546590.0,
+ "step": 3200
+ },
+ {
+ "entropy": 0.12310879753902554,
+ "epoch": 8.690763052208835,
+ "grad_norm": 0.13988597691059113,
+ "learning_rate": 4.6790620731319124e-06,
+ "loss": 0.060741419792175295,
+ "mean_token_accuracy": 0.9787649729847908,
+ "num_tokens": 7666771.0,
+ "step": 3250
+ },
+ {
+ "entropy": 0.13028676200658082,
+ "epoch": 8.824631860776439,
+ "grad_norm": 0.19891025125980377,
+ "learning_rate": 3.787449044042906e-06,
+ "loss": 0.06288023948669434,
+ "mean_token_accuracy": 0.9772285357117653,
+ "num_tokens": 7782159.0,
+ "step": 3300
+ },
+ {
+ "entropy": 0.13001586498692633,
+ "epoch": 8.958500669344042,
+ "grad_norm": 0.1798337697982788,
+ "learning_rate": 2.9863310676379143e-06,
+ "loss": 0.06166903972625733,
+ "mean_token_accuracy": 0.9771321558952332,
+ "num_tokens": 7900010.0,
+ "step": 3350
+ },
+ {
+ "epoch": 9.0,
+ "eval_entropy": 0.24724020630121232,
+ "eval_loss": 1.0340324640274048,
+ "eval_mean_token_accuracy": 0.837886828482151,
+ "eval_num_tokens": 7935471.0,
+ "eval_runtime": 97.1913,
+ "eval_samples_per_second": 16.452,
+ "eval_steps_per_second": 2.058,
+ "step": 3366
+ },
+ {
+ "entropy": 0.13675598613917828,
+ "epoch": 9.09103078982597,
+ "grad_norm": 0.28065016865730286,
+ "learning_rate": 2.2774524760868494e-06,
+ "loss": 0.06236181735992432,
+ "mean_token_accuracy": 0.9770608095809666,
+ "num_tokens": 8010103.0,
+ "step": 3400
+ },
+ {
+ "entropy": 0.12781766049563884,
+ "epoch": 9.224899598393574,
+ "grad_norm": 0.2494969218969345,
+ "learning_rate": 1.662356762068963e-06,
+ "loss": 0.06112825870513916,
+ "mean_token_accuracy": 0.9781736519932747,
+ "num_tokens": 8122846.0,
+ "step": 3450
+ },
+ {
+ "entropy": 0.1289342412352562,
+ "epoch": 9.358768406961179,
+ "grad_norm": 0.2757965326309204,
+ "learning_rate": 1.1423832180145196e-06,
+ "loss": 0.059611873626708986,
+ "mean_token_accuracy": 0.9781888082623482,
+ "num_tokens": 8239477.0,
+ "step": 3500
+ },
+ {
+ "entropy": 0.11920199872925878,
+ "epoch": 9.492637215528783,
+ "grad_norm": 0.17794811725616455,
+ "learning_rate": 7.186640199663721e-07,
+ "loss": 0.05627605438232422,
+ "mean_token_accuracy": 0.9795496875047683,
+ "num_tokens": 8363547.0,
+ "step": 3550
+ },
+ {
+ "entropy": 0.12675758374854923,
+ "epoch": 9.626506024096386,
+ "grad_norm": 0.2903261184692383,
+ "learning_rate": 3.921217624111389e-07,
+ "loss": 0.058483819961547855,
+ "mean_token_accuracy": 0.9784113824367523,
+ "num_tokens": 8482651.0,
+ "step": 3600
+ },
+ {
+ "entropy": 0.12302237752825022,
+ "epoch": 9.76037483266399,
+ "grad_norm": 0.3160693645477295,
+ "learning_rate": 1.6346744944743327e-07,
+ "loss": 0.057983202934265135,
+ "mean_token_accuracy": 0.9791911172866822,
+ "num_tokens": 8601806.0,
+ "step": 3650
+ },
+ {
+ "entropy": 0.11526903139427304,
+ "epoch": 9.894243641231594,
+ "grad_norm": 0.2653988301753998,
+ "learning_rate": 3.319894666513807e-08,
+ "loss": 0.055376110076904295,
+ "mean_token_accuracy": 0.9802410218119622,
+ "num_tokens": 8728357.0,
+ "step": 3700
+ },
+ {
+ "epoch": 10.0,
+ "eval_entropy": 0.2426315762847662,
+ "eval_loss": 1.072160243988037,
+ "eval_mean_token_accuracy": 0.8378600019216538,
+ "eval_num_tokens": 8817190.0,
+ "eval_runtime": 97.4885,
+ "eval_samples_per_second": 16.402,
+ "eval_steps_per_second": 2.052,
+ "step": 3740
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": true
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 3.7345927679778816e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-748/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-748/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-748/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-748/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-748/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..aaa5315fd59e264ef47a82c1fe8efb39be4ce81b
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-748/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.0024401823792587042,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-748/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-748/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-748/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-748/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-748/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-748/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-748/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-748/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..9f0f61cea85ffcc2295f70e47ee3539807a6714d
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test1/checkpoint-748/trainer_state.json
@@ -0,0 +1,196 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 2.0,
+ "eval_steps": 500,
+ "global_step": 748,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.5728709718585014,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.7528014183044434,
+ "learning_rate": 1.188290252952284e-05,
+ "loss": 1.354185791015625,
+ "mean_token_accuracy": 0.7051774816215038,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.7372898015379906,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.4441661834716797,
+ "learning_rate": 2.40083132739339e-05,
+ "loss": 0.6506095123291016,
+ "mean_token_accuracy": 0.8222023022174835,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6791197884082795,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.9964745044708252,
+ "learning_rate": 3.613372401834496e-05,
+ "loss": 0.6006810760498047,
+ "mean_token_accuracy": 0.8327413862943649,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6392384143173695,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.8044542074203491,
+ "learning_rate": 4.825913476275602e-05,
+ "loss": 0.5597672653198242,
+ "mean_token_accuracy": 0.8428809601068497,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6373909455537796,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.9033862948417664,
+ "learning_rate": 6.038454550716708e-05,
+ "loss": 0.5607868576049805,
+ "mean_token_accuracy": 0.8415500408411026,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6142180925607681,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8247693777084351,
+ "learning_rate": 7.250995625157815e-05,
+ "loss": 0.5359848785400391,
+ "mean_token_accuracy": 0.8485029026865959,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.600434636771679,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9828518033027649,
+ "learning_rate": 8.46353669959892e-05,
+ "loss": 0.5246802139282226,
+ "mean_token_accuracy": 0.8504493471980095,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6199610435962677,
+ "eval_loss": 0.5767685174942017,
+ "eval_mean_token_accuracy": 0.8355379092693329,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 98.478,
+ "eval_samples_per_second": 16.237,
+ "eval_steps_per_second": 2.031,
+ "step": 374
+ },
+ {
+ "entropy": 0.5812384999460645,
+ "epoch": 1.069611780455154,
+ "grad_norm": 0.7832330465316772,
+ "learning_rate": 9.068572797060741e-05,
+ "loss": 0.5035604858398437,
+ "mean_token_accuracy": 0.8548897610168265,
+ "num_tokens": 939587.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.5494279730319976,
+ "epoch": 1.2034805890227578,
+ "grad_norm": 0.6947465538978577,
+ "learning_rate": 9.058701310926409e-05,
+ "loss": 0.4750754928588867,
+ "mean_token_accuracy": 0.8600558358430862,
+ "num_tokens": 1058209.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.5506373339891434,
+ "epoch": 1.3373493975903614,
+ "grad_norm": 0.7013292908668518,
+ "learning_rate": 9.038979832559184e-05,
+ "loss": 0.4743109893798828,
+ "mean_token_accuracy": 0.8615988761186599,
+ "num_tokens": 1178618.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.5413181179761887,
+ "epoch": 1.4712182061579653,
+ "grad_norm": 0.8555615544319153,
+ "learning_rate": 9.009451302961728e-05,
+ "loss": 0.4731967544555664,
+ "mean_token_accuracy": 0.8605538108944892,
+ "num_tokens": 1299124.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.5346147198975086,
+ "epoch": 1.605087014725569,
+ "grad_norm": 0.8122630715370178,
+ "learning_rate": 8.970180016739357e-05,
+ "loss": 0.46506633758544924,
+ "mean_token_accuracy": 0.8624722695350647,
+ "num_tokens": 1422003.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.530957569181919,
+ "epoch": 1.7389558232931726,
+ "grad_norm": 0.6519771218299866,
+ "learning_rate": 8.921251482106754e-05,
+ "loss": 0.46199348449707034,
+ "mean_token_accuracy": 0.8636697196960449,
+ "num_tokens": 1542283.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.5241960515081883,
+ "epoch": 1.8728246318607764,
+ "grad_norm": 0.5876154899597168,
+ "learning_rate": 8.862772234704737e-05,
+ "loss": 0.45739620208740234,
+ "mean_token_accuracy": 0.866628175675869,
+ "num_tokens": 1656994.0,
+ "step": 700
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.5676956498622894,
+ "eval_loss": 0.545108437538147,
+ "eval_mean_token_accuracy": 0.8408122035861015,
+ "eval_num_tokens": 1763438.0,
+ "eval_runtime": 97.562,
+ "eval_samples_per_second": 16.39,
+ "eval_steps_per_second": 2.05,
+ "step": 748
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 7.45087263678935e+16,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..447d27c0356e6e2d236dd3623c145bef3a182b97
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/README.md
@@ -0,0 +1,58 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: transformers
+model_name: Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2
+tags:
+- generated_from_trainer
+- trl
+- sft
+licence: license
+---
+
+# Model Card for Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2
+
+This model is a fine-tuned version of [Qwen/Qwen3.5-4B-Base](https://huggingface.co/Qwen/Qwen3.5-4B-Base).
+It has been trained using [TRL](https://github.com/huggingface/trl).
+
+## Quick start
+
+```python
+from transformers import pipeline
+
+question = "If you had a time machine, but could only go to the past or the future once and never return, which would you choose and why?"
+generator = pipeline("text-generation", model="None", device="cuda")
+output = generator([{"role": "user", "content": question}], max_new_tokens=128, return_full_text=False)[0]
+print(output["generated_text"])
+```
+
+## Training procedure
+
+[
](https://wandb.ai/katriin-kukk/Cross_lingual_morphological_generalization/runs/uosbjdwb)
+
+
+
+This model was trained with SFT.
+
+### Framework versions
+
+- TRL: 0.29.0
+- Transformers: 5.5.4
+- Pytorch: 2.10.0
+- Datasets: 4.6.1
+- Tokenizers: 0.22.2
+
+## Citations
+
+
+
+Cite TRL as:
+
+```bibtex
+@software{vonwerra2020trl,
+ title = {{TRL: Transformers Reinforcement Learning}},
+ author = {von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and Beeching, Edward and Thrush, Tristan and Lambert, Nathan and Huang, Shengyi and Rasul, Kashif and Gallouédec, Quentin},
+ license = {Apache-2.0},
+ url = {https://github.com/huggingface/trl},
+ year = {2020}
+}
+```
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1122/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1122/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1122/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1122/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1122/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..6bd19552d78207fd6a10c40231faf153510828db
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1122/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.08713363658978694,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1122/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1122/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1122/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1122/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1122/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1122/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1122/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1122/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..4f4c71d987816caa138d194da9d085ef679811e8
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1122/trainer_state.json
@@ -0,0 +1,287 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 3.0,
+ "eval_steps": 500,
+ "global_step": 1122,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.3166805765032767,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.5876737833023071,
+ "learning_rate": 2.918822371674455e-05,
+ "loss": 1.1408894348144532,
+ "mean_token_accuracy": 0.7406156754493713,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.6964564691483974,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.1449841260910034,
+ "learning_rate": 5.897212546852471e-05,
+ "loss": 0.6171329116821289,
+ "mean_token_accuracy": 0.8284879937767983,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6569836577773094,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.923166811466217,
+ "learning_rate": 8.875602722030486e-05,
+ "loss": 0.5850146102905274,
+ "mean_token_accuracy": 0.8369373443722725,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6257135657966137,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.9985070824623108,
+ "learning_rate": 0.00011853992897208502,
+ "loss": 0.5537805557250977,
+ "mean_token_accuracy": 0.8436850249767304,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6312060031294823,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.8036398887634277,
+ "learning_rate": 0.00014832383072386517,
+ "loss": 0.5614564514160156,
+ "mean_token_accuracy": 0.8422514832019806,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6134552739560604,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8147028088569641,
+ "learning_rate": 0.00017810773247564532,
+ "loss": 0.5460160064697266,
+ "mean_token_accuracy": 0.8461242046952248,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.6121367686986923,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9814819693565369,
+ "learning_rate": 0.0002078916342274255,
+ "loss": 0.5481348037719727,
+ "mean_token_accuracy": 0.845237042605877,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6632432791590691,
+ "eval_loss": 0.6451922655105591,
+ "eval_mean_token_accuracy": 0.8261485403776169,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 99.2666,
+ "eval_samples_per_second": 16.098,
+ "eval_steps_per_second": 2.015,
+ "step": 374
+ },
+ {
+ "entropy": 0.6011885740239211,
+ "epoch": 1.069611780455154,
+ "grad_norm": 0.9249665141105652,
+ "learning_rate": 0.00022275326329957007,
+ "loss": 0.5334951400756835,
+ "mean_token_accuracy": 0.8480202983124088,
+ "num_tokens": 939587.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.5666237922012806,
+ "epoch": 1.2034805890227578,
+ "grad_norm": 0.7326717376708984,
+ "learning_rate": 0.00022251078790688734,
+ "loss": 0.5072500610351562,
+ "mean_token_accuracy": 0.8530087029933929,
+ "num_tokens": 1058209.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.5760806338489055,
+ "epoch": 1.3373493975903614,
+ "grad_norm": 0.7875131964683533,
+ "learning_rate": 0.00022202636508074927,
+ "loss": 0.5134101486206055,
+ "mean_token_accuracy": 0.8527275303006172,
+ "num_tokens": 1178618.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.5638319627940654,
+ "epoch": 1.4712182061579653,
+ "grad_norm": 0.9272266626358032,
+ "learning_rate": 0.00022130104959004678,
+ "loss": 0.5113970184326172,
+ "mean_token_accuracy": 0.8526131376624108,
+ "num_tokens": 1299124.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.5506764651834964,
+ "epoch": 1.605087014725569,
+ "grad_norm": 0.7173363566398621,
+ "learning_rate": 0.00022033642071670964,
+ "loss": 0.5000322723388672,
+ "mean_token_accuracy": 0.8560754117369652,
+ "num_tokens": 1422003.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.5522706513106823,
+ "epoch": 1.7389558232931726,
+ "grad_norm": 0.6712430715560913,
+ "learning_rate": 0.00021913457881702166,
+ "loss": 0.4997842788696289,
+ "mean_token_accuracy": 0.8543951433897018,
+ "num_tokens": 1542283.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.5442864654958248,
+ "epoch": 1.8728246318607764,
+ "grad_norm": 0.7373962998390198,
+ "learning_rate": 0.0002176981407483628,
+ "loss": 0.493228874206543,
+ "mean_token_accuracy": 0.8584304532408714,
+ "num_tokens": 1656994.0,
+ "step": 700
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6364928308129311,
+ "eval_loss": 0.603687584400177,
+ "eval_mean_token_accuracy": 0.8288862618803978,
+ "eval_num_tokens": 1763438.0,
+ "eval_runtime": 98.5299,
+ "eval_samples_per_second": 16.218,
+ "eval_steps_per_second": 2.03,
+ "step": 748
+ },
+ {
+ "entropy": 0.5627047226886557,
+ "epoch": 2.005354752342704,
+ "grad_norm": 0.7084948420524597,
+ "learning_rate": 0.00021603023417133613,
+ "loss": 0.49937862396240235,
+ "mean_token_accuracy": 0.8557315276126669,
+ "num_tokens": 1768737.0,
+ "step": 750
+ },
+ {
+ "entropy": 0.4591581989824772,
+ "epoch": 2.139223560910308,
+ "grad_norm": 0.6059377789497375,
+ "learning_rate": 0.00021413449073968608,
+ "loss": 0.40025680541992187,
+ "mean_token_accuracy": 0.8771893605589867,
+ "num_tokens": 1888608.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.47255669847130777,
+ "epoch": 2.2730923694779115,
+ "grad_norm": 0.6926851272583008,
+ "learning_rate": 0.00021201503819283564,
+ "loss": 0.4088516616821289,
+ "mean_token_accuracy": 0.8737566202878952,
+ "num_tokens": 2000711.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.4772882568836212,
+ "epoch": 2.4069611780455156,
+ "grad_norm": 0.6667259335517883,
+ "learning_rate": 0.0002096764913682607,
+ "loss": 0.4158624649047852,
+ "mean_token_accuracy": 0.8743814784288406,
+ "num_tokens": 2117216.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.475759482383728,
+ "epoch": 2.540829986613119,
+ "grad_norm": 0.8212043642997742,
+ "learning_rate": 0.0002071239421532701,
+ "loss": 0.41526199340820313,
+ "mean_token_accuracy": 0.8735348290205002,
+ "num_tokens": 2233392.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.4801133926212788,
+ "epoch": 2.674698795180723,
+ "grad_norm": 0.7110676765441895,
+ "learning_rate": 0.00020436294839807075,
+ "loss": 0.419442253112793,
+ "mean_token_accuracy": 0.8728689536452293,
+ "num_tokens": 2349527.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.4800094600021839,
+ "epoch": 2.8085676037483265,
+ "grad_norm": 0.7389078140258789,
+ "learning_rate": 0.00020139952181425823,
+ "loss": 0.4172813415527344,
+ "mean_token_accuracy": 0.8737373587489128,
+ "num_tokens": 2473970.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.48250251173973085,
+ "epoch": 2.9424364123159306,
+ "grad_norm": 0.6837686896324158,
+ "learning_rate": 0.0001982401148850816,
+ "loss": 0.42303115844726563,
+ "mean_token_accuracy": 0.8718802312016487,
+ "num_tokens": 2593727.0,
+ "step": 1100
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.5167068694531918,
+ "eval_loss": 0.6138148307800293,
+ "eval_mean_token_accuracy": 0.8335792717337608,
+ "eval_num_tokens": 2645157.0,
+ "eval_runtime": 98.1828,
+ "eval_samples_per_second": 16.276,
+ "eval_steps_per_second": 2.037,
+ "step": 1122
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.120904934916608e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1496/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1496/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1496/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1496/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1496/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..6bd19552d78207fd6a10c40231faf153510828db
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1496/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.08713363658978694,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1496/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1496/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1496/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1496/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1496/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1496/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1496/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1496/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..1c55a1ae901f4b2c56c1d2328c1122eb613f7b8b
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1496/trainer_state.json
@@ -0,0 +1,368 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 4.0,
+ "eval_steps": 500,
+ "global_step": 1496,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.3166805765032767,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.5876737833023071,
+ "learning_rate": 2.918822371674455e-05,
+ "loss": 1.1408894348144532,
+ "mean_token_accuracy": 0.7406156754493713,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.6964564691483974,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.1449841260910034,
+ "learning_rate": 5.897212546852471e-05,
+ "loss": 0.6171329116821289,
+ "mean_token_accuracy": 0.8284879937767983,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6569836577773094,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.923166811466217,
+ "learning_rate": 8.875602722030486e-05,
+ "loss": 0.5850146102905274,
+ "mean_token_accuracy": 0.8369373443722725,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6257135657966137,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.9985070824623108,
+ "learning_rate": 0.00011853992897208502,
+ "loss": 0.5537805557250977,
+ "mean_token_accuracy": 0.8436850249767304,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6312060031294823,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.8036398887634277,
+ "learning_rate": 0.00014832383072386517,
+ "loss": 0.5614564514160156,
+ "mean_token_accuracy": 0.8422514832019806,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6134552739560604,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8147028088569641,
+ "learning_rate": 0.00017810773247564532,
+ "loss": 0.5460160064697266,
+ "mean_token_accuracy": 0.8461242046952248,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.6121367686986923,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9814819693565369,
+ "learning_rate": 0.0002078916342274255,
+ "loss": 0.5481348037719727,
+ "mean_token_accuracy": 0.845237042605877,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6632432791590691,
+ "eval_loss": 0.6451922655105591,
+ "eval_mean_token_accuracy": 0.8261485403776169,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 99.2666,
+ "eval_samples_per_second": 16.098,
+ "eval_steps_per_second": 2.015,
+ "step": 374
+ },
+ {
+ "entropy": 0.6011885740239211,
+ "epoch": 1.069611780455154,
+ "grad_norm": 0.9249665141105652,
+ "learning_rate": 0.00022275326329957007,
+ "loss": 0.5334951400756835,
+ "mean_token_accuracy": 0.8480202983124088,
+ "num_tokens": 939587.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.5666237922012806,
+ "epoch": 1.2034805890227578,
+ "grad_norm": 0.7326717376708984,
+ "learning_rate": 0.00022251078790688734,
+ "loss": 0.5072500610351562,
+ "mean_token_accuracy": 0.8530087029933929,
+ "num_tokens": 1058209.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.5760806338489055,
+ "epoch": 1.3373493975903614,
+ "grad_norm": 0.7875131964683533,
+ "learning_rate": 0.00022202636508074927,
+ "loss": 0.5134101486206055,
+ "mean_token_accuracy": 0.8527275303006172,
+ "num_tokens": 1178618.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.5638319627940654,
+ "epoch": 1.4712182061579653,
+ "grad_norm": 0.9272266626358032,
+ "learning_rate": 0.00022130104959004678,
+ "loss": 0.5113970184326172,
+ "mean_token_accuracy": 0.8526131376624108,
+ "num_tokens": 1299124.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.5506764651834964,
+ "epoch": 1.605087014725569,
+ "grad_norm": 0.7173363566398621,
+ "learning_rate": 0.00022033642071670964,
+ "loss": 0.5000322723388672,
+ "mean_token_accuracy": 0.8560754117369652,
+ "num_tokens": 1422003.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.5522706513106823,
+ "epoch": 1.7389558232931726,
+ "grad_norm": 0.6712430715560913,
+ "learning_rate": 0.00021913457881702166,
+ "loss": 0.4997842788696289,
+ "mean_token_accuracy": 0.8543951433897018,
+ "num_tokens": 1542283.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.5442864654958248,
+ "epoch": 1.8728246318607764,
+ "grad_norm": 0.7373962998390198,
+ "learning_rate": 0.0002176981407483628,
+ "loss": 0.493228874206543,
+ "mean_token_accuracy": 0.8584304532408714,
+ "num_tokens": 1656994.0,
+ "step": 700
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6364928308129311,
+ "eval_loss": 0.603687584400177,
+ "eval_mean_token_accuracy": 0.8288862618803978,
+ "eval_num_tokens": 1763438.0,
+ "eval_runtime": 98.5299,
+ "eval_samples_per_second": 16.218,
+ "eval_steps_per_second": 2.03,
+ "step": 748
+ },
+ {
+ "entropy": 0.5627047226886557,
+ "epoch": 2.005354752342704,
+ "grad_norm": 0.7084948420524597,
+ "learning_rate": 0.00021603023417133613,
+ "loss": 0.49937862396240235,
+ "mean_token_accuracy": 0.8557315276126669,
+ "num_tokens": 1768737.0,
+ "step": 750
+ },
+ {
+ "entropy": 0.4591581989824772,
+ "epoch": 2.139223560910308,
+ "grad_norm": 0.6059377789497375,
+ "learning_rate": 0.00021413449073968608,
+ "loss": 0.40025680541992187,
+ "mean_token_accuracy": 0.8771893605589867,
+ "num_tokens": 1888608.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.47255669847130777,
+ "epoch": 2.2730923694779115,
+ "grad_norm": 0.6926851272583008,
+ "learning_rate": 0.00021201503819283564,
+ "loss": 0.4088516616821289,
+ "mean_token_accuracy": 0.8737566202878952,
+ "num_tokens": 2000711.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.4772882568836212,
+ "epoch": 2.4069611780455156,
+ "grad_norm": 0.6667259335517883,
+ "learning_rate": 0.0002096764913682607,
+ "loss": 0.4158624649047852,
+ "mean_token_accuracy": 0.8743814784288406,
+ "num_tokens": 2117216.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.475759482383728,
+ "epoch": 2.540829986613119,
+ "grad_norm": 0.8212043642997742,
+ "learning_rate": 0.0002071239421532701,
+ "loss": 0.41526199340820313,
+ "mean_token_accuracy": 0.8735348290205002,
+ "num_tokens": 2233392.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.4801133926212788,
+ "epoch": 2.674698795180723,
+ "grad_norm": 0.7110676765441895,
+ "learning_rate": 0.00020436294839807075,
+ "loss": 0.419442253112793,
+ "mean_token_accuracy": 0.8728689536452293,
+ "num_tokens": 2349527.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.4800094600021839,
+ "epoch": 2.8085676037483265,
+ "grad_norm": 0.7389078140258789,
+ "learning_rate": 0.00020139952181425823,
+ "loss": 0.4172813415527344,
+ "mean_token_accuracy": 0.8737373587489128,
+ "num_tokens": 2473970.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.48250251173973085,
+ "epoch": 2.9424364123159306,
+ "grad_norm": 0.6837686896324158,
+ "learning_rate": 0.0001982401148850816,
+ "loss": 0.42303115844726563,
+ "mean_token_accuracy": 0.8718802312016487,
+ "num_tokens": 2593727.0,
+ "step": 1100
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.5167068694531918,
+ "eval_loss": 0.6138148307800293,
+ "eval_mean_token_accuracy": 0.8335792717337608,
+ "eval_num_tokens": 2645157.0,
+ "eval_runtime": 98.1828,
+ "eval_samples_per_second": 16.276,
+ "eval_steps_per_second": 2.037,
+ "step": 1122
+ },
+ {
+ "entropy": 0.41590692727553724,
+ "epoch": 3.074966532797858,
+ "grad_norm": 0.8835903406143188,
+ "learning_rate": 0.00019489160681598466,
+ "loss": 0.35146915435791015,
+ "mean_token_accuracy": 0.8911325624494841,
+ "num_tokens": 2715241.0,
+ "step": 1150
+ },
+ {
+ "entropy": 0.385117310360074,
+ "epoch": 3.208835341365462,
+ "grad_norm": 0.7755147814750671,
+ "learning_rate": 0.0001913612885560138,
+ "loss": 0.31094758987426757,
+ "mean_token_accuracy": 0.8991402065753937,
+ "num_tokens": 2828619.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.38140676759183406,
+ "epoch": 3.3427041499330654,
+ "grad_norm": 0.7542359828948975,
+ "learning_rate": 0.00018765684692270684,
+ "loss": 0.313527889251709,
+ "mean_token_accuracy": 0.8987139958143234,
+ "num_tokens": 2950688.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.38662350915372373,
+ "epoch": 3.4765729585006695,
+ "grad_norm": 0.6831114888191223,
+ "learning_rate": 0.00018378634786502866,
+ "loss": 0.31917388916015627,
+ "mean_token_accuracy": 0.8976927849650383,
+ "num_tokens": 3071943.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.38961954444646835,
+ "epoch": 3.610441767068273,
+ "grad_norm": 0.7666288018226624,
+ "learning_rate": 0.00017975821890079657,
+ "loss": 0.3270668411254883,
+ "mean_token_accuracy": 0.8966797724366188,
+ "num_tokens": 3187873.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.3998099275678396,
+ "epoch": 3.7443105756358768,
+ "grad_norm": 0.6307014226913452,
+ "learning_rate": 0.00017558123076683558,
+ "loss": 0.3317046356201172,
+ "mean_token_accuracy": 0.8944737592339516,
+ "num_tokens": 3299002.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.4006169463694096,
+ "epoch": 3.878179384203481,
+ "grad_norm": 0.7587270736694336,
+ "learning_rate": 0.0001712644783218182,
+ "loss": 0.33513580322265624,
+ "mean_token_accuracy": 0.8940666383504867,
+ "num_tokens": 3422804.0,
+ "step": 1450
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.4926549245417118,
+ "eval_loss": 0.621785581111908,
+ "eval_mean_token_accuracy": 0.8362715128064155,
+ "eval_num_tokens": 3526876.0,
+ "eval_runtime": 98.2264,
+ "eval_samples_per_second": 16.269,
+ "eval_steps_per_second": 2.036,
+ "step": 1496
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.494594216880988e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1870/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1870/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1870/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1870/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1870/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..6bd19552d78207fd6a10c40231faf153510828db
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1870/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.08713363658978694,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1870/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1870/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1870/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1870/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1870/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1870/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1870/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1870/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..73698d9d7ad6edc2429cbafa94039bed6294098c
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-1870/trainer_state.json
@@ -0,0 +1,459 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 5.0,
+ "eval_steps": 500,
+ "global_step": 1870,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.3166805765032767,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.5876737833023071,
+ "learning_rate": 2.918822371674455e-05,
+ "loss": 1.1408894348144532,
+ "mean_token_accuracy": 0.7406156754493713,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.6964564691483974,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.1449841260910034,
+ "learning_rate": 5.897212546852471e-05,
+ "loss": 0.6171329116821289,
+ "mean_token_accuracy": 0.8284879937767983,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6569836577773094,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.923166811466217,
+ "learning_rate": 8.875602722030486e-05,
+ "loss": 0.5850146102905274,
+ "mean_token_accuracy": 0.8369373443722725,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6257135657966137,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.9985070824623108,
+ "learning_rate": 0.00011853992897208502,
+ "loss": 0.5537805557250977,
+ "mean_token_accuracy": 0.8436850249767304,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6312060031294823,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.8036398887634277,
+ "learning_rate": 0.00014832383072386517,
+ "loss": 0.5614564514160156,
+ "mean_token_accuracy": 0.8422514832019806,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6134552739560604,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8147028088569641,
+ "learning_rate": 0.00017810773247564532,
+ "loss": 0.5460160064697266,
+ "mean_token_accuracy": 0.8461242046952248,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.6121367686986923,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9814819693565369,
+ "learning_rate": 0.0002078916342274255,
+ "loss": 0.5481348037719727,
+ "mean_token_accuracy": 0.845237042605877,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6632432791590691,
+ "eval_loss": 0.6451922655105591,
+ "eval_mean_token_accuracy": 0.8261485403776169,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 99.2666,
+ "eval_samples_per_second": 16.098,
+ "eval_steps_per_second": 2.015,
+ "step": 374
+ },
+ {
+ "entropy": 0.6011885740239211,
+ "epoch": 1.069611780455154,
+ "grad_norm": 0.9249665141105652,
+ "learning_rate": 0.00022275326329957007,
+ "loss": 0.5334951400756835,
+ "mean_token_accuracy": 0.8480202983124088,
+ "num_tokens": 939587.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.5666237922012806,
+ "epoch": 1.2034805890227578,
+ "grad_norm": 0.7326717376708984,
+ "learning_rate": 0.00022251078790688734,
+ "loss": 0.5072500610351562,
+ "mean_token_accuracy": 0.8530087029933929,
+ "num_tokens": 1058209.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.5760806338489055,
+ "epoch": 1.3373493975903614,
+ "grad_norm": 0.7875131964683533,
+ "learning_rate": 0.00022202636508074927,
+ "loss": 0.5134101486206055,
+ "mean_token_accuracy": 0.8527275303006172,
+ "num_tokens": 1178618.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.5638319627940654,
+ "epoch": 1.4712182061579653,
+ "grad_norm": 0.9272266626358032,
+ "learning_rate": 0.00022130104959004678,
+ "loss": 0.5113970184326172,
+ "mean_token_accuracy": 0.8526131376624108,
+ "num_tokens": 1299124.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.5506764651834964,
+ "epoch": 1.605087014725569,
+ "grad_norm": 0.7173363566398621,
+ "learning_rate": 0.00022033642071670964,
+ "loss": 0.5000322723388672,
+ "mean_token_accuracy": 0.8560754117369652,
+ "num_tokens": 1422003.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.5522706513106823,
+ "epoch": 1.7389558232931726,
+ "grad_norm": 0.6712430715560913,
+ "learning_rate": 0.00021913457881702166,
+ "loss": 0.4997842788696289,
+ "mean_token_accuracy": 0.8543951433897018,
+ "num_tokens": 1542283.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.5442864654958248,
+ "epoch": 1.8728246318607764,
+ "grad_norm": 0.7373962998390198,
+ "learning_rate": 0.0002176981407483628,
+ "loss": 0.493228874206543,
+ "mean_token_accuracy": 0.8584304532408714,
+ "num_tokens": 1656994.0,
+ "step": 700
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6364928308129311,
+ "eval_loss": 0.603687584400177,
+ "eval_mean_token_accuracy": 0.8288862618803978,
+ "eval_num_tokens": 1763438.0,
+ "eval_runtime": 98.5299,
+ "eval_samples_per_second": 16.218,
+ "eval_steps_per_second": 2.03,
+ "step": 748
+ },
+ {
+ "entropy": 0.5627047226886557,
+ "epoch": 2.005354752342704,
+ "grad_norm": 0.7084948420524597,
+ "learning_rate": 0.00021603023417133613,
+ "loss": 0.49937862396240235,
+ "mean_token_accuracy": 0.8557315276126669,
+ "num_tokens": 1768737.0,
+ "step": 750
+ },
+ {
+ "entropy": 0.4591581989824772,
+ "epoch": 2.139223560910308,
+ "grad_norm": 0.6059377789497375,
+ "learning_rate": 0.00021413449073968608,
+ "loss": 0.40025680541992187,
+ "mean_token_accuracy": 0.8771893605589867,
+ "num_tokens": 1888608.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.47255669847130777,
+ "epoch": 2.2730923694779115,
+ "grad_norm": 0.6926851272583008,
+ "learning_rate": 0.00021201503819283564,
+ "loss": 0.4088516616821289,
+ "mean_token_accuracy": 0.8737566202878952,
+ "num_tokens": 2000711.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.4772882568836212,
+ "epoch": 2.4069611780455156,
+ "grad_norm": 0.6667259335517883,
+ "learning_rate": 0.0002096764913682607,
+ "loss": 0.4158624649047852,
+ "mean_token_accuracy": 0.8743814784288406,
+ "num_tokens": 2117216.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.475759482383728,
+ "epoch": 2.540829986613119,
+ "grad_norm": 0.8212043642997742,
+ "learning_rate": 0.0002071239421532701,
+ "loss": 0.41526199340820313,
+ "mean_token_accuracy": 0.8735348290205002,
+ "num_tokens": 2233392.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.4801133926212788,
+ "epoch": 2.674698795180723,
+ "grad_norm": 0.7110676765441895,
+ "learning_rate": 0.00020436294839807075,
+ "loss": 0.419442253112793,
+ "mean_token_accuracy": 0.8728689536452293,
+ "num_tokens": 2349527.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.4800094600021839,
+ "epoch": 2.8085676037483265,
+ "grad_norm": 0.7389078140258789,
+ "learning_rate": 0.00020139952181425823,
+ "loss": 0.4172813415527344,
+ "mean_token_accuracy": 0.8737373587489128,
+ "num_tokens": 2473970.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.48250251173973085,
+ "epoch": 2.9424364123159306,
+ "grad_norm": 0.6837686896324158,
+ "learning_rate": 0.0001982401148850816,
+ "loss": 0.42303115844726563,
+ "mean_token_accuracy": 0.8718802312016487,
+ "num_tokens": 2593727.0,
+ "step": 1100
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.5167068694531918,
+ "eval_loss": 0.6138148307800293,
+ "eval_mean_token_accuracy": 0.8335792717337608,
+ "eval_num_tokens": 2645157.0,
+ "eval_runtime": 98.1828,
+ "eval_samples_per_second": 16.276,
+ "eval_steps_per_second": 2.037,
+ "step": 1122
+ },
+ {
+ "entropy": 0.41590692727553724,
+ "epoch": 3.074966532797858,
+ "grad_norm": 0.8835903406143188,
+ "learning_rate": 0.00019489160681598466,
+ "loss": 0.35146915435791015,
+ "mean_token_accuracy": 0.8911325624494841,
+ "num_tokens": 2715241.0,
+ "step": 1150
+ },
+ {
+ "entropy": 0.385117310360074,
+ "epoch": 3.208835341365462,
+ "grad_norm": 0.7755147814750671,
+ "learning_rate": 0.0001913612885560138,
+ "loss": 0.31094758987426757,
+ "mean_token_accuracy": 0.8991402065753937,
+ "num_tokens": 2828619.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.38140676759183406,
+ "epoch": 3.3427041499330654,
+ "grad_norm": 0.7542359828948975,
+ "learning_rate": 0.00018765684692270684,
+ "loss": 0.313527889251709,
+ "mean_token_accuracy": 0.8987139958143234,
+ "num_tokens": 2950688.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.38662350915372373,
+ "epoch": 3.4765729585006695,
+ "grad_norm": 0.6831114888191223,
+ "learning_rate": 0.00018378634786502866,
+ "loss": 0.31917388916015627,
+ "mean_token_accuracy": 0.8976927849650383,
+ "num_tokens": 3071943.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.38961954444646835,
+ "epoch": 3.610441767068273,
+ "grad_norm": 0.7666288018226624,
+ "learning_rate": 0.00017975821890079657,
+ "loss": 0.3270668411254883,
+ "mean_token_accuracy": 0.8966797724366188,
+ "num_tokens": 3187873.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.3998099275678396,
+ "epoch": 3.7443105756358768,
+ "grad_norm": 0.6307014226913452,
+ "learning_rate": 0.00017558123076683558,
+ "loss": 0.3317046356201172,
+ "mean_token_accuracy": 0.8944737592339516,
+ "num_tokens": 3299002.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.4006169463694096,
+ "epoch": 3.878179384203481,
+ "grad_norm": 0.7587270736694336,
+ "learning_rate": 0.0001712644783218182,
+ "loss": 0.33513580322265624,
+ "mean_token_accuracy": 0.8940666383504867,
+ "num_tokens": 3422804.0,
+ "step": 1450
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.4926549245417118,
+ "eval_loss": 0.621785581111908,
+ "eval_mean_token_accuracy": 0.8362715128064155,
+ "eval_num_tokens": 3526876.0,
+ "eval_runtime": 98.2264,
+ "eval_samples_per_second": 16.269,
+ "eval_steps_per_second": 2.036,
+ "step": 1496
+ },
+ {
+ "entropy": 0.39729254293923427,
+ "epoch": 4.010709504685408,
+ "grad_norm": 0.6503797173500061,
+ "learning_rate": 0.0001668173607433701,
+ "loss": 0.3273897171020508,
+ "mean_token_accuracy": 0.8963238362110022,
+ "num_tokens": 3536363.0,
+ "step": 1500
+ },
+ {
+ "entropy": 0.2908159454166889,
+ "epoch": 4.144578313253012,
+ "grad_norm": 0.8419110178947449,
+ "learning_rate": 0.00016224956106256037,
+ "loss": 0.21537052154541014,
+ "mean_token_accuracy": 0.9289350810647011,
+ "num_tokens": 3653560.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.29764041006565095,
+ "epoch": 4.278447121820616,
+ "grad_norm": 0.8312885165214539,
+ "learning_rate": 0.00015757102508033658,
+ "loss": 0.22335006713867187,
+ "mean_token_accuracy": 0.925491844713688,
+ "num_tokens": 3776281.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.29720947936177255,
+ "epoch": 4.412315930388219,
+ "grad_norm": 0.8321210145950317,
+ "learning_rate": 0.00015279193971181267,
+ "loss": 0.225880184173584,
+ "mean_token_accuracy": 0.9243844848871231,
+ "num_tokens": 3896081.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.30463078498840335,
+ "epoch": 4.546184738955823,
+ "grad_norm": 0.9243009686470032,
+ "learning_rate": 0.0001479227108055611,
+ "loss": 0.22961084365844728,
+ "mean_token_accuracy": 0.9234738460183144,
+ "num_tokens": 4010469.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.3072978562861681,
+ "epoch": 4.680053547523427,
+ "grad_norm": 0.6637229919433594,
+ "learning_rate": 0.00014297394048620504,
+ "loss": 0.23345823287963868,
+ "mean_token_accuracy": 0.9230251079797744,
+ "num_tokens": 4125858.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.30938181333243847,
+ "epoch": 4.813922356091031,
+ "grad_norm": 0.8940643072128296,
+ "learning_rate": 0.00013795640406964534,
+ "loss": 0.23422428131103515,
+ "mean_token_accuracy": 0.9217449072003364,
+ "num_tokens": 4243193.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.30204505123198033,
+ "epoch": 4.947791164658635,
+ "grad_norm": 0.7235498428344727,
+ "learning_rate": 0.00013288102660118477,
+ "loss": 0.22840303421020508,
+ "mean_token_accuracy": 0.9235807290673256,
+ "num_tokens": 4365873.0,
+ "step": 1850
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.4090891806781292,
+ "eval_loss": 0.6856206655502319,
+ "eval_mean_token_accuracy": 0.8357787102460861,
+ "eval_num_tokens": 4408595.0,
+ "eval_runtime": 98.466,
+ "eval_samples_per_second": 16.229,
+ "eval_steps_per_second": 2.031,
+ "step": 1870
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.8661573535249203e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2244/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2244/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2244/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2244/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2244/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..6bd19552d78207fd6a10c40231faf153510828db
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2244/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.08713363658978694,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2244/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2244/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2244/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2244/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2244/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2244/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2244/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2244/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..228352be10178a6b60e2834c23f00366a9a18b01
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2244/trainer_state.json
@@ -0,0 +1,540 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 6.0,
+ "eval_steps": 500,
+ "global_step": 2244,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.3166805765032767,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.5876737833023071,
+ "learning_rate": 2.918822371674455e-05,
+ "loss": 1.1408894348144532,
+ "mean_token_accuracy": 0.7406156754493713,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.6964564691483974,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.1449841260910034,
+ "learning_rate": 5.897212546852471e-05,
+ "loss": 0.6171329116821289,
+ "mean_token_accuracy": 0.8284879937767983,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6569836577773094,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.923166811466217,
+ "learning_rate": 8.875602722030486e-05,
+ "loss": 0.5850146102905274,
+ "mean_token_accuracy": 0.8369373443722725,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6257135657966137,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.9985070824623108,
+ "learning_rate": 0.00011853992897208502,
+ "loss": 0.5537805557250977,
+ "mean_token_accuracy": 0.8436850249767304,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6312060031294823,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.8036398887634277,
+ "learning_rate": 0.00014832383072386517,
+ "loss": 0.5614564514160156,
+ "mean_token_accuracy": 0.8422514832019806,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6134552739560604,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8147028088569641,
+ "learning_rate": 0.00017810773247564532,
+ "loss": 0.5460160064697266,
+ "mean_token_accuracy": 0.8461242046952248,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.6121367686986923,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9814819693565369,
+ "learning_rate": 0.0002078916342274255,
+ "loss": 0.5481348037719727,
+ "mean_token_accuracy": 0.845237042605877,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6632432791590691,
+ "eval_loss": 0.6451922655105591,
+ "eval_mean_token_accuracy": 0.8261485403776169,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 99.2666,
+ "eval_samples_per_second": 16.098,
+ "eval_steps_per_second": 2.015,
+ "step": 374
+ },
+ {
+ "entropy": 0.6011885740239211,
+ "epoch": 1.069611780455154,
+ "grad_norm": 0.9249665141105652,
+ "learning_rate": 0.00022275326329957007,
+ "loss": 0.5334951400756835,
+ "mean_token_accuracy": 0.8480202983124088,
+ "num_tokens": 939587.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.5666237922012806,
+ "epoch": 1.2034805890227578,
+ "grad_norm": 0.7326717376708984,
+ "learning_rate": 0.00022251078790688734,
+ "loss": 0.5072500610351562,
+ "mean_token_accuracy": 0.8530087029933929,
+ "num_tokens": 1058209.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.5760806338489055,
+ "epoch": 1.3373493975903614,
+ "grad_norm": 0.7875131964683533,
+ "learning_rate": 0.00022202636508074927,
+ "loss": 0.5134101486206055,
+ "mean_token_accuracy": 0.8527275303006172,
+ "num_tokens": 1178618.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.5638319627940654,
+ "epoch": 1.4712182061579653,
+ "grad_norm": 0.9272266626358032,
+ "learning_rate": 0.00022130104959004678,
+ "loss": 0.5113970184326172,
+ "mean_token_accuracy": 0.8526131376624108,
+ "num_tokens": 1299124.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.5506764651834964,
+ "epoch": 1.605087014725569,
+ "grad_norm": 0.7173363566398621,
+ "learning_rate": 0.00022033642071670964,
+ "loss": 0.5000322723388672,
+ "mean_token_accuracy": 0.8560754117369652,
+ "num_tokens": 1422003.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.5522706513106823,
+ "epoch": 1.7389558232931726,
+ "grad_norm": 0.6712430715560913,
+ "learning_rate": 0.00021913457881702166,
+ "loss": 0.4997842788696289,
+ "mean_token_accuracy": 0.8543951433897018,
+ "num_tokens": 1542283.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.5442864654958248,
+ "epoch": 1.8728246318607764,
+ "grad_norm": 0.7373962998390198,
+ "learning_rate": 0.0002176981407483628,
+ "loss": 0.493228874206543,
+ "mean_token_accuracy": 0.8584304532408714,
+ "num_tokens": 1656994.0,
+ "step": 700
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6364928308129311,
+ "eval_loss": 0.603687584400177,
+ "eval_mean_token_accuracy": 0.8288862618803978,
+ "eval_num_tokens": 1763438.0,
+ "eval_runtime": 98.5299,
+ "eval_samples_per_second": 16.218,
+ "eval_steps_per_second": 2.03,
+ "step": 748
+ },
+ {
+ "entropy": 0.5627047226886557,
+ "epoch": 2.005354752342704,
+ "grad_norm": 0.7084948420524597,
+ "learning_rate": 0.00021603023417133613,
+ "loss": 0.49937862396240235,
+ "mean_token_accuracy": 0.8557315276126669,
+ "num_tokens": 1768737.0,
+ "step": 750
+ },
+ {
+ "entropy": 0.4591581989824772,
+ "epoch": 2.139223560910308,
+ "grad_norm": 0.6059377789497375,
+ "learning_rate": 0.00021413449073968608,
+ "loss": 0.40025680541992187,
+ "mean_token_accuracy": 0.8771893605589867,
+ "num_tokens": 1888608.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.47255669847130777,
+ "epoch": 2.2730923694779115,
+ "grad_norm": 0.6926851272583008,
+ "learning_rate": 0.00021201503819283564,
+ "loss": 0.4088516616821289,
+ "mean_token_accuracy": 0.8737566202878952,
+ "num_tokens": 2000711.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.4772882568836212,
+ "epoch": 2.4069611780455156,
+ "grad_norm": 0.6667259335517883,
+ "learning_rate": 0.0002096764913682607,
+ "loss": 0.4158624649047852,
+ "mean_token_accuracy": 0.8743814784288406,
+ "num_tokens": 2117216.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.475759482383728,
+ "epoch": 2.540829986613119,
+ "grad_norm": 0.8212043642997742,
+ "learning_rate": 0.0002071239421532701,
+ "loss": 0.41526199340820313,
+ "mean_token_accuracy": 0.8735348290205002,
+ "num_tokens": 2233392.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.4801133926212788,
+ "epoch": 2.674698795180723,
+ "grad_norm": 0.7110676765441895,
+ "learning_rate": 0.00020436294839807075,
+ "loss": 0.419442253112793,
+ "mean_token_accuracy": 0.8728689536452293,
+ "num_tokens": 2349527.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.4800094600021839,
+ "epoch": 2.8085676037483265,
+ "grad_norm": 0.7389078140258789,
+ "learning_rate": 0.00020139952181425823,
+ "loss": 0.4172813415527344,
+ "mean_token_accuracy": 0.8737373587489128,
+ "num_tokens": 2473970.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.48250251173973085,
+ "epoch": 2.9424364123159306,
+ "grad_norm": 0.6837686896324158,
+ "learning_rate": 0.0001982401148850816,
+ "loss": 0.42303115844726563,
+ "mean_token_accuracy": 0.8718802312016487,
+ "num_tokens": 2593727.0,
+ "step": 1100
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.5167068694531918,
+ "eval_loss": 0.6138148307800293,
+ "eval_mean_token_accuracy": 0.8335792717337608,
+ "eval_num_tokens": 2645157.0,
+ "eval_runtime": 98.1828,
+ "eval_samples_per_second": 16.276,
+ "eval_steps_per_second": 2.037,
+ "step": 1122
+ },
+ {
+ "entropy": 0.41590692727553724,
+ "epoch": 3.074966532797858,
+ "grad_norm": 0.8835903406143188,
+ "learning_rate": 0.00019489160681598466,
+ "loss": 0.35146915435791015,
+ "mean_token_accuracy": 0.8911325624494841,
+ "num_tokens": 2715241.0,
+ "step": 1150
+ },
+ {
+ "entropy": 0.385117310360074,
+ "epoch": 3.208835341365462,
+ "grad_norm": 0.7755147814750671,
+ "learning_rate": 0.0001913612885560138,
+ "loss": 0.31094758987426757,
+ "mean_token_accuracy": 0.8991402065753937,
+ "num_tokens": 2828619.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.38140676759183406,
+ "epoch": 3.3427041499330654,
+ "grad_norm": 0.7542359828948975,
+ "learning_rate": 0.00018765684692270684,
+ "loss": 0.313527889251709,
+ "mean_token_accuracy": 0.8987139958143234,
+ "num_tokens": 2950688.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.38662350915372373,
+ "epoch": 3.4765729585006695,
+ "grad_norm": 0.6831114888191223,
+ "learning_rate": 0.00018378634786502866,
+ "loss": 0.31917388916015627,
+ "mean_token_accuracy": 0.8976927849650383,
+ "num_tokens": 3071943.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.38961954444646835,
+ "epoch": 3.610441767068273,
+ "grad_norm": 0.7666288018226624,
+ "learning_rate": 0.00017975821890079657,
+ "loss": 0.3270668411254883,
+ "mean_token_accuracy": 0.8966797724366188,
+ "num_tokens": 3187873.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.3998099275678396,
+ "epoch": 3.7443105756358768,
+ "grad_norm": 0.6307014226913452,
+ "learning_rate": 0.00017558123076683558,
+ "loss": 0.3317046356201172,
+ "mean_token_accuracy": 0.8944737592339516,
+ "num_tokens": 3299002.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.4006169463694096,
+ "epoch": 3.878179384203481,
+ "grad_norm": 0.7587270736694336,
+ "learning_rate": 0.0001712644783218182,
+ "loss": 0.33513580322265624,
+ "mean_token_accuracy": 0.8940666383504867,
+ "num_tokens": 3422804.0,
+ "step": 1450
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.4926549245417118,
+ "eval_loss": 0.621785581111908,
+ "eval_mean_token_accuracy": 0.8362715128064155,
+ "eval_num_tokens": 3526876.0,
+ "eval_runtime": 98.2264,
+ "eval_samples_per_second": 16.269,
+ "eval_steps_per_second": 2.036,
+ "step": 1496
+ },
+ {
+ "entropy": 0.39729254293923427,
+ "epoch": 4.010709504685408,
+ "grad_norm": 0.6503797173500061,
+ "learning_rate": 0.0001668173607433701,
+ "loss": 0.3273897171020508,
+ "mean_token_accuracy": 0.8963238362110022,
+ "num_tokens": 3536363.0,
+ "step": 1500
+ },
+ {
+ "entropy": 0.2908159454166889,
+ "epoch": 4.144578313253012,
+ "grad_norm": 0.8419110178947449,
+ "learning_rate": 0.00016224956106256037,
+ "loss": 0.21537052154541014,
+ "mean_token_accuracy": 0.9289350810647011,
+ "num_tokens": 3653560.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.29764041006565095,
+ "epoch": 4.278447121820616,
+ "grad_norm": 0.8312885165214539,
+ "learning_rate": 0.00015757102508033658,
+ "loss": 0.22335006713867187,
+ "mean_token_accuracy": 0.925491844713688,
+ "num_tokens": 3776281.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.29720947936177255,
+ "epoch": 4.412315930388219,
+ "grad_norm": 0.8321210145950317,
+ "learning_rate": 0.00015279193971181267,
+ "loss": 0.225880184173584,
+ "mean_token_accuracy": 0.9243844848871231,
+ "num_tokens": 3896081.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.30463078498840335,
+ "epoch": 4.546184738955823,
+ "grad_norm": 0.9243009686470032,
+ "learning_rate": 0.0001479227108055611,
+ "loss": 0.22961084365844728,
+ "mean_token_accuracy": 0.9234738460183144,
+ "num_tokens": 4010469.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.3072978562861681,
+ "epoch": 4.680053547523427,
+ "grad_norm": 0.6637229919433594,
+ "learning_rate": 0.00014297394048620504,
+ "loss": 0.23345823287963868,
+ "mean_token_accuracy": 0.9230251079797744,
+ "num_tokens": 4125858.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.30938181333243847,
+ "epoch": 4.813922356091031,
+ "grad_norm": 0.8940643072128296,
+ "learning_rate": 0.00013795640406964534,
+ "loss": 0.23422428131103515,
+ "mean_token_accuracy": 0.9217449072003364,
+ "num_tokens": 4243193.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.30204505123198033,
+ "epoch": 4.947791164658635,
+ "grad_norm": 0.7235498428344727,
+ "learning_rate": 0.00013288102660118477,
+ "loss": 0.22840303421020508,
+ "mean_token_accuracy": 0.9235807290673256,
+ "num_tokens": 4365873.0,
+ "step": 1850
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.4090891806781292,
+ "eval_loss": 0.6856206655502319,
+ "eval_mean_token_accuracy": 0.8357787102460861,
+ "eval_num_tokens": 4408595.0,
+ "eval_runtime": 98.466,
+ "eval_samples_per_second": 16.229,
+ "eval_steps_per_second": 2.031,
+ "step": 1870
+ },
+ {
+ "entropy": 0.25283067240709006,
+ "epoch": 5.080321285140562,
+ "grad_norm": 0.569952130317688,
+ "learning_rate": 0.00012775885906763602,
+ "loss": 0.17759113311767577,
+ "mean_token_accuracy": 0.9404868823711319,
+ "num_tokens": 4482540.0,
+ "step": 1900
+ },
+ {
+ "entropy": 0.2179089618474245,
+ "epoch": 5.214190093708166,
+ "grad_norm": 0.6869003772735596,
+ "learning_rate": 0.00012260105433520803,
+ "loss": 0.14341320037841798,
+ "mean_token_accuracy": 0.952086115181446,
+ "num_tokens": 4598191.0,
+ "step": 1950
+ },
+ {
+ "entropy": 0.2195182240009308,
+ "epoch": 5.34805890227577,
+ "grad_norm": 0.8023325800895691,
+ "learning_rate": 0.00011741884286556301,
+ "loss": 0.14362534523010254,
+ "mean_token_accuracy": 0.9521202966570854,
+ "num_tokens": 4716333.0,
+ "step": 2000
+ },
+ {
+ "entropy": 0.22079841319471596,
+ "epoch": 5.481927710843373,
+ "grad_norm": 0.7231829166412354,
+ "learning_rate": 0.00011222350826291902,
+ "loss": 0.14616458892822265,
+ "mean_token_accuracy": 0.9513887700438499,
+ "num_tokens": 4830262.0,
+ "step": 2050
+ },
+ {
+ "entropy": 0.2166673281788826,
+ "epoch": 5.615796519410977,
+ "grad_norm": 0.790205180644989,
+ "learning_rate": 0.0001070263627054418,
+ "loss": 0.14604829788208007,
+ "mean_token_accuracy": 0.9511423835158348,
+ "num_tokens": 4951007.0,
+ "step": 2100
+ },
+ {
+ "entropy": 0.22030118089169265,
+ "epoch": 5.749665327978581,
+ "grad_norm": 0.7226701974868774,
+ "learning_rate": 0.00010183872231442008,
+ "loss": 0.14739423751831054,
+ "mean_token_accuracy": 0.950632050037384,
+ "num_tokens": 5067535.0,
+ "step": 2150
+ },
+ {
+ "entropy": 0.2082419555261731,
+ "epoch": 5.883534136546185,
+ "grad_norm": 0.7288137674331665,
+ "learning_rate": 9.667188251485557e-05,
+ "loss": 0.14179450988769532,
+ "mean_token_accuracy": 0.9521883800625801,
+ "num_tokens": 5185751.0,
+ "step": 2200
+ },
+ {
+ "epoch": 6.0,
+ "eval_entropy": 0.33322301916778085,
+ "eval_loss": 0.7674635648727417,
+ "eval_mean_token_accuracy": 0.8357109513878822,
+ "eval_num_tokens": 5290314.0,
+ "eval_runtime": 98.6056,
+ "eval_samples_per_second": 16.206,
+ "eval_steps_per_second": 2.028,
+ "step": 2244
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.2368361483527168e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2618/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2618/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2618/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2618/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2618/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..6bd19552d78207fd6a10c40231faf153510828db
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2618/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.08713363658978694,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2618/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2618/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2618/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2618/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2618/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2618/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2618/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2618/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..a96d2320fb91de2b936f9b9ea961a0df363bb682
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2618/trainer_state.json
@@ -0,0 +1,631 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 7.0,
+ "eval_steps": 500,
+ "global_step": 2618,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.3166805765032767,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.5876737833023071,
+ "learning_rate": 2.918822371674455e-05,
+ "loss": 1.1408894348144532,
+ "mean_token_accuracy": 0.7406156754493713,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.6964564691483974,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.1449841260910034,
+ "learning_rate": 5.897212546852471e-05,
+ "loss": 0.6171329116821289,
+ "mean_token_accuracy": 0.8284879937767983,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6569836577773094,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.923166811466217,
+ "learning_rate": 8.875602722030486e-05,
+ "loss": 0.5850146102905274,
+ "mean_token_accuracy": 0.8369373443722725,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6257135657966137,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.9985070824623108,
+ "learning_rate": 0.00011853992897208502,
+ "loss": 0.5537805557250977,
+ "mean_token_accuracy": 0.8436850249767304,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6312060031294823,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.8036398887634277,
+ "learning_rate": 0.00014832383072386517,
+ "loss": 0.5614564514160156,
+ "mean_token_accuracy": 0.8422514832019806,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6134552739560604,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8147028088569641,
+ "learning_rate": 0.00017810773247564532,
+ "loss": 0.5460160064697266,
+ "mean_token_accuracy": 0.8461242046952248,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.6121367686986923,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9814819693565369,
+ "learning_rate": 0.0002078916342274255,
+ "loss": 0.5481348037719727,
+ "mean_token_accuracy": 0.845237042605877,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6632432791590691,
+ "eval_loss": 0.6451922655105591,
+ "eval_mean_token_accuracy": 0.8261485403776169,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 99.2666,
+ "eval_samples_per_second": 16.098,
+ "eval_steps_per_second": 2.015,
+ "step": 374
+ },
+ {
+ "entropy": 0.6011885740239211,
+ "epoch": 1.069611780455154,
+ "grad_norm": 0.9249665141105652,
+ "learning_rate": 0.00022275326329957007,
+ "loss": 0.5334951400756835,
+ "mean_token_accuracy": 0.8480202983124088,
+ "num_tokens": 939587.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.5666237922012806,
+ "epoch": 1.2034805890227578,
+ "grad_norm": 0.7326717376708984,
+ "learning_rate": 0.00022251078790688734,
+ "loss": 0.5072500610351562,
+ "mean_token_accuracy": 0.8530087029933929,
+ "num_tokens": 1058209.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.5760806338489055,
+ "epoch": 1.3373493975903614,
+ "grad_norm": 0.7875131964683533,
+ "learning_rate": 0.00022202636508074927,
+ "loss": 0.5134101486206055,
+ "mean_token_accuracy": 0.8527275303006172,
+ "num_tokens": 1178618.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.5638319627940654,
+ "epoch": 1.4712182061579653,
+ "grad_norm": 0.9272266626358032,
+ "learning_rate": 0.00022130104959004678,
+ "loss": 0.5113970184326172,
+ "mean_token_accuracy": 0.8526131376624108,
+ "num_tokens": 1299124.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.5506764651834964,
+ "epoch": 1.605087014725569,
+ "grad_norm": 0.7173363566398621,
+ "learning_rate": 0.00022033642071670964,
+ "loss": 0.5000322723388672,
+ "mean_token_accuracy": 0.8560754117369652,
+ "num_tokens": 1422003.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.5522706513106823,
+ "epoch": 1.7389558232931726,
+ "grad_norm": 0.6712430715560913,
+ "learning_rate": 0.00021913457881702166,
+ "loss": 0.4997842788696289,
+ "mean_token_accuracy": 0.8543951433897018,
+ "num_tokens": 1542283.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.5442864654958248,
+ "epoch": 1.8728246318607764,
+ "grad_norm": 0.7373962998390198,
+ "learning_rate": 0.0002176981407483628,
+ "loss": 0.493228874206543,
+ "mean_token_accuracy": 0.8584304532408714,
+ "num_tokens": 1656994.0,
+ "step": 700
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6364928308129311,
+ "eval_loss": 0.603687584400177,
+ "eval_mean_token_accuracy": 0.8288862618803978,
+ "eval_num_tokens": 1763438.0,
+ "eval_runtime": 98.5299,
+ "eval_samples_per_second": 16.218,
+ "eval_steps_per_second": 2.03,
+ "step": 748
+ },
+ {
+ "entropy": 0.5627047226886557,
+ "epoch": 2.005354752342704,
+ "grad_norm": 0.7084948420524597,
+ "learning_rate": 0.00021603023417133613,
+ "loss": 0.49937862396240235,
+ "mean_token_accuracy": 0.8557315276126669,
+ "num_tokens": 1768737.0,
+ "step": 750
+ },
+ {
+ "entropy": 0.4591581989824772,
+ "epoch": 2.139223560910308,
+ "grad_norm": 0.6059377789497375,
+ "learning_rate": 0.00021413449073968608,
+ "loss": 0.40025680541992187,
+ "mean_token_accuracy": 0.8771893605589867,
+ "num_tokens": 1888608.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.47255669847130777,
+ "epoch": 2.2730923694779115,
+ "grad_norm": 0.6926851272583008,
+ "learning_rate": 0.00021201503819283564,
+ "loss": 0.4088516616821289,
+ "mean_token_accuracy": 0.8737566202878952,
+ "num_tokens": 2000711.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.4772882568836212,
+ "epoch": 2.4069611780455156,
+ "grad_norm": 0.6667259335517883,
+ "learning_rate": 0.0002096764913682607,
+ "loss": 0.4158624649047852,
+ "mean_token_accuracy": 0.8743814784288406,
+ "num_tokens": 2117216.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.475759482383728,
+ "epoch": 2.540829986613119,
+ "grad_norm": 0.8212043642997742,
+ "learning_rate": 0.0002071239421532701,
+ "loss": 0.41526199340820313,
+ "mean_token_accuracy": 0.8735348290205002,
+ "num_tokens": 2233392.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.4801133926212788,
+ "epoch": 2.674698795180723,
+ "grad_norm": 0.7110676765441895,
+ "learning_rate": 0.00020436294839807075,
+ "loss": 0.419442253112793,
+ "mean_token_accuracy": 0.8728689536452293,
+ "num_tokens": 2349527.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.4800094600021839,
+ "epoch": 2.8085676037483265,
+ "grad_norm": 0.7389078140258789,
+ "learning_rate": 0.00020139952181425823,
+ "loss": 0.4172813415527344,
+ "mean_token_accuracy": 0.8737373587489128,
+ "num_tokens": 2473970.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.48250251173973085,
+ "epoch": 2.9424364123159306,
+ "grad_norm": 0.6837686896324158,
+ "learning_rate": 0.0001982401148850816,
+ "loss": 0.42303115844726563,
+ "mean_token_accuracy": 0.8718802312016487,
+ "num_tokens": 2593727.0,
+ "step": 1100
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.5167068694531918,
+ "eval_loss": 0.6138148307800293,
+ "eval_mean_token_accuracy": 0.8335792717337608,
+ "eval_num_tokens": 2645157.0,
+ "eval_runtime": 98.1828,
+ "eval_samples_per_second": 16.276,
+ "eval_steps_per_second": 2.037,
+ "step": 1122
+ },
+ {
+ "entropy": 0.41590692727553724,
+ "epoch": 3.074966532797858,
+ "grad_norm": 0.8835903406143188,
+ "learning_rate": 0.00019489160681598466,
+ "loss": 0.35146915435791015,
+ "mean_token_accuracy": 0.8911325624494841,
+ "num_tokens": 2715241.0,
+ "step": 1150
+ },
+ {
+ "entropy": 0.385117310360074,
+ "epoch": 3.208835341365462,
+ "grad_norm": 0.7755147814750671,
+ "learning_rate": 0.0001913612885560138,
+ "loss": 0.31094758987426757,
+ "mean_token_accuracy": 0.8991402065753937,
+ "num_tokens": 2828619.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.38140676759183406,
+ "epoch": 3.3427041499330654,
+ "grad_norm": 0.7542359828948975,
+ "learning_rate": 0.00018765684692270684,
+ "loss": 0.313527889251709,
+ "mean_token_accuracy": 0.8987139958143234,
+ "num_tokens": 2950688.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.38662350915372373,
+ "epoch": 3.4765729585006695,
+ "grad_norm": 0.6831114888191223,
+ "learning_rate": 0.00018378634786502866,
+ "loss": 0.31917388916015627,
+ "mean_token_accuracy": 0.8976927849650383,
+ "num_tokens": 3071943.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.38961954444646835,
+ "epoch": 3.610441767068273,
+ "grad_norm": 0.7666288018226624,
+ "learning_rate": 0.00017975821890079657,
+ "loss": 0.3270668411254883,
+ "mean_token_accuracy": 0.8966797724366188,
+ "num_tokens": 3187873.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.3998099275678396,
+ "epoch": 3.7443105756358768,
+ "grad_norm": 0.6307014226913452,
+ "learning_rate": 0.00017558123076683558,
+ "loss": 0.3317046356201172,
+ "mean_token_accuracy": 0.8944737592339516,
+ "num_tokens": 3299002.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.4006169463694096,
+ "epoch": 3.878179384203481,
+ "grad_norm": 0.7587270736694336,
+ "learning_rate": 0.0001712644783218182,
+ "loss": 0.33513580322265624,
+ "mean_token_accuracy": 0.8940666383504867,
+ "num_tokens": 3422804.0,
+ "step": 1450
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.4926549245417118,
+ "eval_loss": 0.621785581111908,
+ "eval_mean_token_accuracy": 0.8362715128064155,
+ "eval_num_tokens": 3526876.0,
+ "eval_runtime": 98.2264,
+ "eval_samples_per_second": 16.269,
+ "eval_steps_per_second": 2.036,
+ "step": 1496
+ },
+ {
+ "entropy": 0.39729254293923427,
+ "epoch": 4.010709504685408,
+ "grad_norm": 0.6503797173500061,
+ "learning_rate": 0.0001668173607433701,
+ "loss": 0.3273897171020508,
+ "mean_token_accuracy": 0.8963238362110022,
+ "num_tokens": 3536363.0,
+ "step": 1500
+ },
+ {
+ "entropy": 0.2908159454166889,
+ "epoch": 4.144578313253012,
+ "grad_norm": 0.8419110178947449,
+ "learning_rate": 0.00016224956106256037,
+ "loss": 0.21537052154541014,
+ "mean_token_accuracy": 0.9289350810647011,
+ "num_tokens": 3653560.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.29764041006565095,
+ "epoch": 4.278447121820616,
+ "grad_norm": 0.8312885165214539,
+ "learning_rate": 0.00015757102508033658,
+ "loss": 0.22335006713867187,
+ "mean_token_accuracy": 0.925491844713688,
+ "num_tokens": 3776281.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.29720947936177255,
+ "epoch": 4.412315930388219,
+ "grad_norm": 0.8321210145950317,
+ "learning_rate": 0.00015279193971181267,
+ "loss": 0.225880184173584,
+ "mean_token_accuracy": 0.9243844848871231,
+ "num_tokens": 3896081.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.30463078498840335,
+ "epoch": 4.546184738955823,
+ "grad_norm": 0.9243009686470032,
+ "learning_rate": 0.0001479227108055611,
+ "loss": 0.22961084365844728,
+ "mean_token_accuracy": 0.9234738460183144,
+ "num_tokens": 4010469.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.3072978562861681,
+ "epoch": 4.680053547523427,
+ "grad_norm": 0.6637229919433594,
+ "learning_rate": 0.00014297394048620504,
+ "loss": 0.23345823287963868,
+ "mean_token_accuracy": 0.9230251079797744,
+ "num_tokens": 4125858.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.30938181333243847,
+ "epoch": 4.813922356091031,
+ "grad_norm": 0.8940643072128296,
+ "learning_rate": 0.00013795640406964534,
+ "loss": 0.23422428131103515,
+ "mean_token_accuracy": 0.9217449072003364,
+ "num_tokens": 4243193.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.30204505123198033,
+ "epoch": 4.947791164658635,
+ "grad_norm": 0.7235498428344727,
+ "learning_rate": 0.00013288102660118477,
+ "loss": 0.22840303421020508,
+ "mean_token_accuracy": 0.9235807290673256,
+ "num_tokens": 4365873.0,
+ "step": 1850
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.4090891806781292,
+ "eval_loss": 0.6856206655502319,
+ "eval_mean_token_accuracy": 0.8357787102460861,
+ "eval_num_tokens": 4408595.0,
+ "eval_runtime": 98.466,
+ "eval_samples_per_second": 16.229,
+ "eval_steps_per_second": 2.031,
+ "step": 1870
+ },
+ {
+ "entropy": 0.25283067240709006,
+ "epoch": 5.080321285140562,
+ "grad_norm": 0.569952130317688,
+ "learning_rate": 0.00012775885906763602,
+ "loss": 0.17759113311767577,
+ "mean_token_accuracy": 0.9404868823711319,
+ "num_tokens": 4482540.0,
+ "step": 1900
+ },
+ {
+ "entropy": 0.2179089618474245,
+ "epoch": 5.214190093708166,
+ "grad_norm": 0.6869003772735596,
+ "learning_rate": 0.00012260105433520803,
+ "loss": 0.14341320037841798,
+ "mean_token_accuracy": 0.952086115181446,
+ "num_tokens": 4598191.0,
+ "step": 1950
+ },
+ {
+ "entropy": 0.2195182240009308,
+ "epoch": 5.34805890227577,
+ "grad_norm": 0.8023325800895691,
+ "learning_rate": 0.00011741884286556301,
+ "loss": 0.14362534523010254,
+ "mean_token_accuracy": 0.9521202966570854,
+ "num_tokens": 4716333.0,
+ "step": 2000
+ },
+ {
+ "entropy": 0.22079841319471596,
+ "epoch": 5.481927710843373,
+ "grad_norm": 0.7231829166412354,
+ "learning_rate": 0.00011222350826291902,
+ "loss": 0.14616458892822265,
+ "mean_token_accuracy": 0.9513887700438499,
+ "num_tokens": 4830262.0,
+ "step": 2050
+ },
+ {
+ "entropy": 0.2166673281788826,
+ "epoch": 5.615796519410977,
+ "grad_norm": 0.790205180644989,
+ "learning_rate": 0.0001070263627054418,
+ "loss": 0.14604829788208007,
+ "mean_token_accuracy": 0.9511423835158348,
+ "num_tokens": 4951007.0,
+ "step": 2100
+ },
+ {
+ "entropy": 0.22030118089169265,
+ "epoch": 5.749665327978581,
+ "grad_norm": 0.7226701974868774,
+ "learning_rate": 0.00010183872231442008,
+ "loss": 0.14739423751831054,
+ "mean_token_accuracy": 0.950632050037384,
+ "num_tokens": 5067535.0,
+ "step": 2150
+ },
+ {
+ "entropy": 0.2082419555261731,
+ "epoch": 5.883534136546185,
+ "grad_norm": 0.7288137674331665,
+ "learning_rate": 9.667188251485557e-05,
+ "loss": 0.14179450988769532,
+ "mean_token_accuracy": 0.9521883800625801,
+ "num_tokens": 5185751.0,
+ "step": 2200
+ },
+ {
+ "epoch": 6.0,
+ "eval_entropy": 0.33322301916778085,
+ "eval_loss": 0.7674635648727417,
+ "eval_mean_token_accuracy": 0.8357109513878822,
+ "eval_num_tokens": 5290314.0,
+ "eval_runtime": 98.6056,
+ "eval_samples_per_second": 16.206,
+ "eval_steps_per_second": 2.028,
+ "step": 2244
+ },
+ {
+ "entropy": 0.20694272517405374,
+ "epoch": 6.016064257028113,
+ "grad_norm": 0.48393547534942627,
+ "learning_rate": 9.15370934411162e-05,
+ "loss": 0.1373324966430664,
+ "mean_token_accuracy": 0.9537640566175635,
+ "num_tokens": 5303628.0,
+ "step": 2250
+ },
+ {
+ "entropy": 0.154647543951869,
+ "epoch": 6.149933065595716,
+ "grad_norm": 0.5554987192153931,
+ "learning_rate": 8.64455354412042e-05,
+ "loss": 0.09246989250183106,
+ "mean_token_accuracy": 0.9694107329845428,
+ "num_tokens": 5424681.0,
+ "step": 2300
+ },
+ {
+ "entropy": 0.14666760206222534,
+ "epoch": 6.28380187416332,
+ "grad_norm": 0.6542149186134338,
+ "learning_rate": 8.140829473297485e-05,
+ "loss": 0.09034320831298828,
+ "mean_token_accuracy": 0.9701463675498962,
+ "num_tokens": 5546681.0,
+ "step": 2350
+ },
+ {
+ "entropy": 0.16469813615083695,
+ "epoch": 6.417670682730924,
+ "grad_norm": 0.7045366764068604,
+ "learning_rate": 7.643633926531171e-05,
+ "loss": 0.09508686065673828,
+ "mean_token_accuracy": 0.9681181335449218,
+ "num_tokens": 5660189.0,
+ "step": 2400
+ },
+ {
+ "entropy": 0.15761504411697388,
+ "epoch": 6.551539491298527,
+ "grad_norm": 0.469072550535202,
+ "learning_rate": 7.154049483681754e-05,
+ "loss": 0.09064769744873047,
+ "mean_token_accuracy": 0.9688174912333488,
+ "num_tokens": 5780047.0,
+ "step": 2450
+ },
+ {
+ "entropy": 0.15717160735279323,
+ "epoch": 6.685408299866131,
+ "grad_norm": 0.8090623021125793,
+ "learning_rate": 6.67314215240192e-05,
+ "loss": 0.09295336723327637,
+ "mean_token_accuracy": 0.9688407072424888,
+ "num_tokens": 5897170.0,
+ "step": 2500
+ },
+ {
+ "entropy": 0.1579529893398285,
+ "epoch": 6.8192771084337345,
+ "grad_norm": 0.5264602899551392,
+ "learning_rate": 6.201959047041119e-05,
+ "loss": 0.09421208381652832,
+ "mean_token_accuracy": 0.9683848118782044,
+ "num_tokens": 6012704.0,
+ "step": 2550
+ },
+ {
+ "entropy": 0.157079309374094,
+ "epoch": 6.953145917001339,
+ "grad_norm": 0.38391751050949097,
+ "learning_rate": 5.74152610868768e-05,
+ "loss": 0.09274236679077148,
+ "mean_token_accuracy": 0.9683843395113945,
+ "num_tokens": 6132181.0,
+ "step": 2600
+ },
+ {
+ "epoch": 7.0,
+ "eval_entropy": 0.28141631513834,
+ "eval_loss": 0.8675811290740967,
+ "eval_mean_token_accuracy": 0.8351070094108581,
+ "eval_num_tokens": 6172033.0,
+ "eval_runtime": 98.8812,
+ "eval_samples_per_second": 16.161,
+ "eval_steps_per_second": 2.023,
+ "step": 2618
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.612589418143744e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2992/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2992/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2992/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2992/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2992/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..6bd19552d78207fd6a10c40231faf153510828db
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2992/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.08713363658978694,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2992/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2992/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2992/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2992/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2992/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2992/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2992/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2992/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..2b6963115a5e6c832cd25c106365cbd8139a90cd
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-2992/trainer_state.json
@@ -0,0 +1,712 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 8.0,
+ "eval_steps": 500,
+ "global_step": 2992,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.3166805765032767,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.5876737833023071,
+ "learning_rate": 2.918822371674455e-05,
+ "loss": 1.1408894348144532,
+ "mean_token_accuracy": 0.7406156754493713,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.6964564691483974,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.1449841260910034,
+ "learning_rate": 5.897212546852471e-05,
+ "loss": 0.6171329116821289,
+ "mean_token_accuracy": 0.8284879937767983,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6569836577773094,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.923166811466217,
+ "learning_rate": 8.875602722030486e-05,
+ "loss": 0.5850146102905274,
+ "mean_token_accuracy": 0.8369373443722725,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6257135657966137,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.9985070824623108,
+ "learning_rate": 0.00011853992897208502,
+ "loss": 0.5537805557250977,
+ "mean_token_accuracy": 0.8436850249767304,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6312060031294823,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.8036398887634277,
+ "learning_rate": 0.00014832383072386517,
+ "loss": 0.5614564514160156,
+ "mean_token_accuracy": 0.8422514832019806,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6134552739560604,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8147028088569641,
+ "learning_rate": 0.00017810773247564532,
+ "loss": 0.5460160064697266,
+ "mean_token_accuracy": 0.8461242046952248,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.6121367686986923,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9814819693565369,
+ "learning_rate": 0.0002078916342274255,
+ "loss": 0.5481348037719727,
+ "mean_token_accuracy": 0.845237042605877,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6632432791590691,
+ "eval_loss": 0.6451922655105591,
+ "eval_mean_token_accuracy": 0.8261485403776169,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 99.2666,
+ "eval_samples_per_second": 16.098,
+ "eval_steps_per_second": 2.015,
+ "step": 374
+ },
+ {
+ "entropy": 0.6011885740239211,
+ "epoch": 1.069611780455154,
+ "grad_norm": 0.9249665141105652,
+ "learning_rate": 0.00022275326329957007,
+ "loss": 0.5334951400756835,
+ "mean_token_accuracy": 0.8480202983124088,
+ "num_tokens": 939587.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.5666237922012806,
+ "epoch": 1.2034805890227578,
+ "grad_norm": 0.7326717376708984,
+ "learning_rate": 0.00022251078790688734,
+ "loss": 0.5072500610351562,
+ "mean_token_accuracy": 0.8530087029933929,
+ "num_tokens": 1058209.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.5760806338489055,
+ "epoch": 1.3373493975903614,
+ "grad_norm": 0.7875131964683533,
+ "learning_rate": 0.00022202636508074927,
+ "loss": 0.5134101486206055,
+ "mean_token_accuracy": 0.8527275303006172,
+ "num_tokens": 1178618.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.5638319627940654,
+ "epoch": 1.4712182061579653,
+ "grad_norm": 0.9272266626358032,
+ "learning_rate": 0.00022130104959004678,
+ "loss": 0.5113970184326172,
+ "mean_token_accuracy": 0.8526131376624108,
+ "num_tokens": 1299124.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.5506764651834964,
+ "epoch": 1.605087014725569,
+ "grad_norm": 0.7173363566398621,
+ "learning_rate": 0.00022033642071670964,
+ "loss": 0.5000322723388672,
+ "mean_token_accuracy": 0.8560754117369652,
+ "num_tokens": 1422003.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.5522706513106823,
+ "epoch": 1.7389558232931726,
+ "grad_norm": 0.6712430715560913,
+ "learning_rate": 0.00021913457881702166,
+ "loss": 0.4997842788696289,
+ "mean_token_accuracy": 0.8543951433897018,
+ "num_tokens": 1542283.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.5442864654958248,
+ "epoch": 1.8728246318607764,
+ "grad_norm": 0.7373962998390198,
+ "learning_rate": 0.0002176981407483628,
+ "loss": 0.493228874206543,
+ "mean_token_accuracy": 0.8584304532408714,
+ "num_tokens": 1656994.0,
+ "step": 700
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6364928308129311,
+ "eval_loss": 0.603687584400177,
+ "eval_mean_token_accuracy": 0.8288862618803978,
+ "eval_num_tokens": 1763438.0,
+ "eval_runtime": 98.5299,
+ "eval_samples_per_second": 16.218,
+ "eval_steps_per_second": 2.03,
+ "step": 748
+ },
+ {
+ "entropy": 0.5627047226886557,
+ "epoch": 2.005354752342704,
+ "grad_norm": 0.7084948420524597,
+ "learning_rate": 0.00021603023417133613,
+ "loss": 0.49937862396240235,
+ "mean_token_accuracy": 0.8557315276126669,
+ "num_tokens": 1768737.0,
+ "step": 750
+ },
+ {
+ "entropy": 0.4591581989824772,
+ "epoch": 2.139223560910308,
+ "grad_norm": 0.6059377789497375,
+ "learning_rate": 0.00021413449073968608,
+ "loss": 0.40025680541992187,
+ "mean_token_accuracy": 0.8771893605589867,
+ "num_tokens": 1888608.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.47255669847130777,
+ "epoch": 2.2730923694779115,
+ "grad_norm": 0.6926851272583008,
+ "learning_rate": 0.00021201503819283564,
+ "loss": 0.4088516616821289,
+ "mean_token_accuracy": 0.8737566202878952,
+ "num_tokens": 2000711.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.4772882568836212,
+ "epoch": 2.4069611780455156,
+ "grad_norm": 0.6667259335517883,
+ "learning_rate": 0.0002096764913682607,
+ "loss": 0.4158624649047852,
+ "mean_token_accuracy": 0.8743814784288406,
+ "num_tokens": 2117216.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.475759482383728,
+ "epoch": 2.540829986613119,
+ "grad_norm": 0.8212043642997742,
+ "learning_rate": 0.0002071239421532701,
+ "loss": 0.41526199340820313,
+ "mean_token_accuracy": 0.8735348290205002,
+ "num_tokens": 2233392.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.4801133926212788,
+ "epoch": 2.674698795180723,
+ "grad_norm": 0.7110676765441895,
+ "learning_rate": 0.00020436294839807075,
+ "loss": 0.419442253112793,
+ "mean_token_accuracy": 0.8728689536452293,
+ "num_tokens": 2349527.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.4800094600021839,
+ "epoch": 2.8085676037483265,
+ "grad_norm": 0.7389078140258789,
+ "learning_rate": 0.00020139952181425823,
+ "loss": 0.4172813415527344,
+ "mean_token_accuracy": 0.8737373587489128,
+ "num_tokens": 2473970.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.48250251173973085,
+ "epoch": 2.9424364123159306,
+ "grad_norm": 0.6837686896324158,
+ "learning_rate": 0.0001982401148850816,
+ "loss": 0.42303115844726563,
+ "mean_token_accuracy": 0.8718802312016487,
+ "num_tokens": 2593727.0,
+ "step": 1100
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.5167068694531918,
+ "eval_loss": 0.6138148307800293,
+ "eval_mean_token_accuracy": 0.8335792717337608,
+ "eval_num_tokens": 2645157.0,
+ "eval_runtime": 98.1828,
+ "eval_samples_per_second": 16.276,
+ "eval_steps_per_second": 2.037,
+ "step": 1122
+ },
+ {
+ "entropy": 0.41590692727553724,
+ "epoch": 3.074966532797858,
+ "grad_norm": 0.8835903406143188,
+ "learning_rate": 0.00019489160681598466,
+ "loss": 0.35146915435791015,
+ "mean_token_accuracy": 0.8911325624494841,
+ "num_tokens": 2715241.0,
+ "step": 1150
+ },
+ {
+ "entropy": 0.385117310360074,
+ "epoch": 3.208835341365462,
+ "grad_norm": 0.7755147814750671,
+ "learning_rate": 0.0001913612885560138,
+ "loss": 0.31094758987426757,
+ "mean_token_accuracy": 0.8991402065753937,
+ "num_tokens": 2828619.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.38140676759183406,
+ "epoch": 3.3427041499330654,
+ "grad_norm": 0.7542359828948975,
+ "learning_rate": 0.00018765684692270684,
+ "loss": 0.313527889251709,
+ "mean_token_accuracy": 0.8987139958143234,
+ "num_tokens": 2950688.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.38662350915372373,
+ "epoch": 3.4765729585006695,
+ "grad_norm": 0.6831114888191223,
+ "learning_rate": 0.00018378634786502866,
+ "loss": 0.31917388916015627,
+ "mean_token_accuracy": 0.8976927849650383,
+ "num_tokens": 3071943.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.38961954444646835,
+ "epoch": 3.610441767068273,
+ "grad_norm": 0.7666288018226624,
+ "learning_rate": 0.00017975821890079657,
+ "loss": 0.3270668411254883,
+ "mean_token_accuracy": 0.8966797724366188,
+ "num_tokens": 3187873.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.3998099275678396,
+ "epoch": 3.7443105756358768,
+ "grad_norm": 0.6307014226913452,
+ "learning_rate": 0.00017558123076683558,
+ "loss": 0.3317046356201172,
+ "mean_token_accuracy": 0.8944737592339516,
+ "num_tokens": 3299002.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.4006169463694096,
+ "epoch": 3.878179384203481,
+ "grad_norm": 0.7587270736694336,
+ "learning_rate": 0.0001712644783218182,
+ "loss": 0.33513580322265624,
+ "mean_token_accuracy": 0.8940666383504867,
+ "num_tokens": 3422804.0,
+ "step": 1450
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.4926549245417118,
+ "eval_loss": 0.621785581111908,
+ "eval_mean_token_accuracy": 0.8362715128064155,
+ "eval_num_tokens": 3526876.0,
+ "eval_runtime": 98.2264,
+ "eval_samples_per_second": 16.269,
+ "eval_steps_per_second": 2.036,
+ "step": 1496
+ },
+ {
+ "entropy": 0.39729254293923427,
+ "epoch": 4.010709504685408,
+ "grad_norm": 0.6503797173500061,
+ "learning_rate": 0.0001668173607433701,
+ "loss": 0.3273897171020508,
+ "mean_token_accuracy": 0.8963238362110022,
+ "num_tokens": 3536363.0,
+ "step": 1500
+ },
+ {
+ "entropy": 0.2908159454166889,
+ "epoch": 4.144578313253012,
+ "grad_norm": 0.8419110178947449,
+ "learning_rate": 0.00016224956106256037,
+ "loss": 0.21537052154541014,
+ "mean_token_accuracy": 0.9289350810647011,
+ "num_tokens": 3653560.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.29764041006565095,
+ "epoch": 4.278447121820616,
+ "grad_norm": 0.8312885165214539,
+ "learning_rate": 0.00015757102508033658,
+ "loss": 0.22335006713867187,
+ "mean_token_accuracy": 0.925491844713688,
+ "num_tokens": 3776281.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.29720947936177255,
+ "epoch": 4.412315930388219,
+ "grad_norm": 0.8321210145950317,
+ "learning_rate": 0.00015279193971181267,
+ "loss": 0.225880184173584,
+ "mean_token_accuracy": 0.9243844848871231,
+ "num_tokens": 3896081.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.30463078498840335,
+ "epoch": 4.546184738955823,
+ "grad_norm": 0.9243009686470032,
+ "learning_rate": 0.0001479227108055611,
+ "loss": 0.22961084365844728,
+ "mean_token_accuracy": 0.9234738460183144,
+ "num_tokens": 4010469.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.3072978562861681,
+ "epoch": 4.680053547523427,
+ "grad_norm": 0.6637229919433594,
+ "learning_rate": 0.00014297394048620504,
+ "loss": 0.23345823287963868,
+ "mean_token_accuracy": 0.9230251079797744,
+ "num_tokens": 4125858.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.30938181333243847,
+ "epoch": 4.813922356091031,
+ "grad_norm": 0.8940643072128296,
+ "learning_rate": 0.00013795640406964534,
+ "loss": 0.23422428131103515,
+ "mean_token_accuracy": 0.9217449072003364,
+ "num_tokens": 4243193.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.30204505123198033,
+ "epoch": 4.947791164658635,
+ "grad_norm": 0.7235498428344727,
+ "learning_rate": 0.00013288102660118477,
+ "loss": 0.22840303421020508,
+ "mean_token_accuracy": 0.9235807290673256,
+ "num_tokens": 4365873.0,
+ "step": 1850
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.4090891806781292,
+ "eval_loss": 0.6856206655502319,
+ "eval_mean_token_accuracy": 0.8357787102460861,
+ "eval_num_tokens": 4408595.0,
+ "eval_runtime": 98.466,
+ "eval_samples_per_second": 16.229,
+ "eval_steps_per_second": 2.031,
+ "step": 1870
+ },
+ {
+ "entropy": 0.25283067240709006,
+ "epoch": 5.080321285140562,
+ "grad_norm": 0.569952130317688,
+ "learning_rate": 0.00012775885906763602,
+ "loss": 0.17759113311767577,
+ "mean_token_accuracy": 0.9404868823711319,
+ "num_tokens": 4482540.0,
+ "step": 1900
+ },
+ {
+ "entropy": 0.2179089618474245,
+ "epoch": 5.214190093708166,
+ "grad_norm": 0.6869003772735596,
+ "learning_rate": 0.00012260105433520803,
+ "loss": 0.14341320037841798,
+ "mean_token_accuracy": 0.952086115181446,
+ "num_tokens": 4598191.0,
+ "step": 1950
+ },
+ {
+ "entropy": 0.2195182240009308,
+ "epoch": 5.34805890227577,
+ "grad_norm": 0.8023325800895691,
+ "learning_rate": 0.00011741884286556301,
+ "loss": 0.14362534523010254,
+ "mean_token_accuracy": 0.9521202966570854,
+ "num_tokens": 4716333.0,
+ "step": 2000
+ },
+ {
+ "entropy": 0.22079841319471596,
+ "epoch": 5.481927710843373,
+ "grad_norm": 0.7231829166412354,
+ "learning_rate": 0.00011222350826291902,
+ "loss": 0.14616458892822265,
+ "mean_token_accuracy": 0.9513887700438499,
+ "num_tokens": 4830262.0,
+ "step": 2050
+ },
+ {
+ "entropy": 0.2166673281788826,
+ "epoch": 5.615796519410977,
+ "grad_norm": 0.790205180644989,
+ "learning_rate": 0.0001070263627054418,
+ "loss": 0.14604829788208007,
+ "mean_token_accuracy": 0.9511423835158348,
+ "num_tokens": 4951007.0,
+ "step": 2100
+ },
+ {
+ "entropy": 0.22030118089169265,
+ "epoch": 5.749665327978581,
+ "grad_norm": 0.7226701974868774,
+ "learning_rate": 0.00010183872231442008,
+ "loss": 0.14739423751831054,
+ "mean_token_accuracy": 0.950632050037384,
+ "num_tokens": 5067535.0,
+ "step": 2150
+ },
+ {
+ "entropy": 0.2082419555261731,
+ "epoch": 5.883534136546185,
+ "grad_norm": 0.7288137674331665,
+ "learning_rate": 9.667188251485557e-05,
+ "loss": 0.14179450988769532,
+ "mean_token_accuracy": 0.9521883800625801,
+ "num_tokens": 5185751.0,
+ "step": 2200
+ },
+ {
+ "epoch": 6.0,
+ "eval_entropy": 0.33322301916778085,
+ "eval_loss": 0.7674635648727417,
+ "eval_mean_token_accuracy": 0.8357109513878822,
+ "eval_num_tokens": 5290314.0,
+ "eval_runtime": 98.6056,
+ "eval_samples_per_second": 16.206,
+ "eval_steps_per_second": 2.028,
+ "step": 2244
+ },
+ {
+ "entropy": 0.20694272517405374,
+ "epoch": 6.016064257028113,
+ "grad_norm": 0.48393547534942627,
+ "learning_rate": 9.15370934411162e-05,
+ "loss": 0.1373324966430664,
+ "mean_token_accuracy": 0.9537640566175635,
+ "num_tokens": 5303628.0,
+ "step": 2250
+ },
+ {
+ "entropy": 0.154647543951869,
+ "epoch": 6.149933065595716,
+ "grad_norm": 0.5554987192153931,
+ "learning_rate": 8.64455354412042e-05,
+ "loss": 0.09246989250183106,
+ "mean_token_accuracy": 0.9694107329845428,
+ "num_tokens": 5424681.0,
+ "step": 2300
+ },
+ {
+ "entropy": 0.14666760206222534,
+ "epoch": 6.28380187416332,
+ "grad_norm": 0.6542149186134338,
+ "learning_rate": 8.140829473297485e-05,
+ "loss": 0.09034320831298828,
+ "mean_token_accuracy": 0.9701463675498962,
+ "num_tokens": 5546681.0,
+ "step": 2350
+ },
+ {
+ "entropy": 0.16469813615083695,
+ "epoch": 6.417670682730924,
+ "grad_norm": 0.7045366764068604,
+ "learning_rate": 7.643633926531171e-05,
+ "loss": 0.09508686065673828,
+ "mean_token_accuracy": 0.9681181335449218,
+ "num_tokens": 5660189.0,
+ "step": 2400
+ },
+ {
+ "entropy": 0.15761504411697388,
+ "epoch": 6.551539491298527,
+ "grad_norm": 0.469072550535202,
+ "learning_rate": 7.154049483681754e-05,
+ "loss": 0.09064769744873047,
+ "mean_token_accuracy": 0.9688174912333488,
+ "num_tokens": 5780047.0,
+ "step": 2450
+ },
+ {
+ "entropy": 0.15717160735279323,
+ "epoch": 6.685408299866131,
+ "grad_norm": 0.8090623021125793,
+ "learning_rate": 6.67314215240192e-05,
+ "loss": 0.09295336723327637,
+ "mean_token_accuracy": 0.9688407072424888,
+ "num_tokens": 5897170.0,
+ "step": 2500
+ },
+ {
+ "entropy": 0.1579529893398285,
+ "epoch": 6.8192771084337345,
+ "grad_norm": 0.5264602899551392,
+ "learning_rate": 6.201959047041119e-05,
+ "loss": 0.09421208381652832,
+ "mean_token_accuracy": 0.9683848118782044,
+ "num_tokens": 6012704.0,
+ "step": 2550
+ },
+ {
+ "entropy": 0.157079309374094,
+ "epoch": 6.953145917001339,
+ "grad_norm": 0.38391751050949097,
+ "learning_rate": 5.74152610868768e-05,
+ "loss": 0.09274236679077148,
+ "mean_token_accuracy": 0.9683843395113945,
+ "num_tokens": 6132181.0,
+ "step": 2600
+ },
+ {
+ "epoch": 7.0,
+ "eval_entropy": 0.28141631513834,
+ "eval_loss": 0.8675811290740967,
+ "eval_mean_token_accuracy": 0.8351070094108581,
+ "eval_num_tokens": 6172033.0,
+ "eval_runtime": 98.8812,
+ "eval_samples_per_second": 16.161,
+ "eval_steps_per_second": 2.023,
+ "step": 2618
+ },
+ {
+ "entropy": 0.14476348296033614,
+ "epoch": 7.085676037483267,
+ "grad_norm": 0.2593551278114319,
+ "learning_rate": 5.2928458713130264e-05,
+ "loss": 0.07731590747833252,
+ "mean_token_accuracy": 0.9728465519770227,
+ "num_tokens": 6250634.0,
+ "step": 2650
+ },
+ {
+ "entropy": 0.1265121574141085,
+ "epoch": 7.21954484605087,
+ "grad_norm": 0.3758811950683594,
+ "learning_rate": 4.8568952788818444e-05,
+ "loss": 0.0695980167388916,
+ "mean_token_accuracy": 0.9759974965453148,
+ "num_tokens": 6371882.0,
+ "step": 2700
+ },
+ {
+ "entropy": 0.13779850516468287,
+ "epoch": 7.353413654618474,
+ "grad_norm": 0.4968899190425873,
+ "learning_rate": 4.43462355818128e-05,
+ "loss": 0.0733350658416748,
+ "mean_token_accuracy": 0.9740503445267678,
+ "num_tokens": 6483611.0,
+ "step": 2750
+ },
+ {
+ "entropy": 0.12721603013575078,
+ "epoch": 7.4872824631860775,
+ "grad_norm": 0.2620279788970947,
+ "learning_rate": 4.026950152000728e-05,
+ "loss": 0.07116839408874512,
+ "mean_token_accuracy": 0.9759298366308212,
+ "num_tokens": 6602543.0,
+ "step": 2800
+ },
+ {
+ "entropy": 0.1206374079361558,
+ "epoch": 7.621151271753681,
+ "grad_norm": 0.31671959161758423,
+ "learning_rate": 3.634762717162492e-05,
+ "loss": 0.0680537223815918,
+ "mean_token_accuracy": 0.9765615722537041,
+ "num_tokens": 6727899.0,
+ "step": 2850
+ },
+ {
+ "entropy": 0.12501337694004178,
+ "epoch": 7.755020080321285,
+ "grad_norm": 0.2820671796798706,
+ "learning_rate": 3.2589151917622866e-05,
+ "loss": 0.07058364391326905,
+ "mean_token_accuracy": 0.975590398311615,
+ "num_tokens": 6845842.0,
+ "step": 2900
+ },
+ {
+ "entropy": 0.12397634202614427,
+ "epoch": 7.888888888888889,
+ "grad_norm": 0.3127625286579132,
+ "learning_rate": 2.90022593582803e-05,
+ "loss": 0.07234051704406738,
+ "mean_token_accuracy": 0.9755299004912377,
+ "num_tokens": 6961910.0,
+ "step": 2950
+ },
+ {
+ "epoch": 8.0,
+ "eval_entropy": 0.2577117122709751,
+ "eval_loss": 0.9746259450912476,
+ "eval_mean_token_accuracy": 0.8363272827863694,
+ "eval_num_tokens": 7053752.0,
+ "eval_runtime": 98.6758,
+ "eval_samples_per_second": 16.194,
+ "eval_steps_per_second": 2.027,
+ "step": 2992
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.987999811940086e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3366/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3366/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3366/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3366/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3366/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..6bd19552d78207fd6a10c40231faf153510828db
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3366/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.08713363658978694,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3366/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3366/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3366/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3366/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3366/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3366/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3366/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3366/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..17efe0fca63c49614f800609489f84299e5207c7
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3366/trainer_state.json
@@ -0,0 +1,803 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 9.0,
+ "eval_steps": 500,
+ "global_step": 3366,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.3166805765032767,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.5876737833023071,
+ "learning_rate": 2.918822371674455e-05,
+ "loss": 1.1408894348144532,
+ "mean_token_accuracy": 0.7406156754493713,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.6964564691483974,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.1449841260910034,
+ "learning_rate": 5.897212546852471e-05,
+ "loss": 0.6171329116821289,
+ "mean_token_accuracy": 0.8284879937767983,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6569836577773094,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.923166811466217,
+ "learning_rate": 8.875602722030486e-05,
+ "loss": 0.5850146102905274,
+ "mean_token_accuracy": 0.8369373443722725,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6257135657966137,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.9985070824623108,
+ "learning_rate": 0.00011853992897208502,
+ "loss": 0.5537805557250977,
+ "mean_token_accuracy": 0.8436850249767304,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6312060031294823,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.8036398887634277,
+ "learning_rate": 0.00014832383072386517,
+ "loss": 0.5614564514160156,
+ "mean_token_accuracy": 0.8422514832019806,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6134552739560604,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8147028088569641,
+ "learning_rate": 0.00017810773247564532,
+ "loss": 0.5460160064697266,
+ "mean_token_accuracy": 0.8461242046952248,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.6121367686986923,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9814819693565369,
+ "learning_rate": 0.0002078916342274255,
+ "loss": 0.5481348037719727,
+ "mean_token_accuracy": 0.845237042605877,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6632432791590691,
+ "eval_loss": 0.6451922655105591,
+ "eval_mean_token_accuracy": 0.8261485403776169,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 99.2666,
+ "eval_samples_per_second": 16.098,
+ "eval_steps_per_second": 2.015,
+ "step": 374
+ },
+ {
+ "entropy": 0.6011885740239211,
+ "epoch": 1.069611780455154,
+ "grad_norm": 0.9249665141105652,
+ "learning_rate": 0.00022275326329957007,
+ "loss": 0.5334951400756835,
+ "mean_token_accuracy": 0.8480202983124088,
+ "num_tokens": 939587.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.5666237922012806,
+ "epoch": 1.2034805890227578,
+ "grad_norm": 0.7326717376708984,
+ "learning_rate": 0.00022251078790688734,
+ "loss": 0.5072500610351562,
+ "mean_token_accuracy": 0.8530087029933929,
+ "num_tokens": 1058209.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.5760806338489055,
+ "epoch": 1.3373493975903614,
+ "grad_norm": 0.7875131964683533,
+ "learning_rate": 0.00022202636508074927,
+ "loss": 0.5134101486206055,
+ "mean_token_accuracy": 0.8527275303006172,
+ "num_tokens": 1178618.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.5638319627940654,
+ "epoch": 1.4712182061579653,
+ "grad_norm": 0.9272266626358032,
+ "learning_rate": 0.00022130104959004678,
+ "loss": 0.5113970184326172,
+ "mean_token_accuracy": 0.8526131376624108,
+ "num_tokens": 1299124.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.5506764651834964,
+ "epoch": 1.605087014725569,
+ "grad_norm": 0.7173363566398621,
+ "learning_rate": 0.00022033642071670964,
+ "loss": 0.5000322723388672,
+ "mean_token_accuracy": 0.8560754117369652,
+ "num_tokens": 1422003.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.5522706513106823,
+ "epoch": 1.7389558232931726,
+ "grad_norm": 0.6712430715560913,
+ "learning_rate": 0.00021913457881702166,
+ "loss": 0.4997842788696289,
+ "mean_token_accuracy": 0.8543951433897018,
+ "num_tokens": 1542283.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.5442864654958248,
+ "epoch": 1.8728246318607764,
+ "grad_norm": 0.7373962998390198,
+ "learning_rate": 0.0002176981407483628,
+ "loss": 0.493228874206543,
+ "mean_token_accuracy": 0.8584304532408714,
+ "num_tokens": 1656994.0,
+ "step": 700
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6364928308129311,
+ "eval_loss": 0.603687584400177,
+ "eval_mean_token_accuracy": 0.8288862618803978,
+ "eval_num_tokens": 1763438.0,
+ "eval_runtime": 98.5299,
+ "eval_samples_per_second": 16.218,
+ "eval_steps_per_second": 2.03,
+ "step": 748
+ },
+ {
+ "entropy": 0.5627047226886557,
+ "epoch": 2.005354752342704,
+ "grad_norm": 0.7084948420524597,
+ "learning_rate": 0.00021603023417133613,
+ "loss": 0.49937862396240235,
+ "mean_token_accuracy": 0.8557315276126669,
+ "num_tokens": 1768737.0,
+ "step": 750
+ },
+ {
+ "entropy": 0.4591581989824772,
+ "epoch": 2.139223560910308,
+ "grad_norm": 0.6059377789497375,
+ "learning_rate": 0.00021413449073968608,
+ "loss": 0.40025680541992187,
+ "mean_token_accuracy": 0.8771893605589867,
+ "num_tokens": 1888608.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.47255669847130777,
+ "epoch": 2.2730923694779115,
+ "grad_norm": 0.6926851272583008,
+ "learning_rate": 0.00021201503819283564,
+ "loss": 0.4088516616821289,
+ "mean_token_accuracy": 0.8737566202878952,
+ "num_tokens": 2000711.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.4772882568836212,
+ "epoch": 2.4069611780455156,
+ "grad_norm": 0.6667259335517883,
+ "learning_rate": 0.0002096764913682607,
+ "loss": 0.4158624649047852,
+ "mean_token_accuracy": 0.8743814784288406,
+ "num_tokens": 2117216.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.475759482383728,
+ "epoch": 2.540829986613119,
+ "grad_norm": 0.8212043642997742,
+ "learning_rate": 0.0002071239421532701,
+ "loss": 0.41526199340820313,
+ "mean_token_accuracy": 0.8735348290205002,
+ "num_tokens": 2233392.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.4801133926212788,
+ "epoch": 2.674698795180723,
+ "grad_norm": 0.7110676765441895,
+ "learning_rate": 0.00020436294839807075,
+ "loss": 0.419442253112793,
+ "mean_token_accuracy": 0.8728689536452293,
+ "num_tokens": 2349527.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.4800094600021839,
+ "epoch": 2.8085676037483265,
+ "grad_norm": 0.7389078140258789,
+ "learning_rate": 0.00020139952181425823,
+ "loss": 0.4172813415527344,
+ "mean_token_accuracy": 0.8737373587489128,
+ "num_tokens": 2473970.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.48250251173973085,
+ "epoch": 2.9424364123159306,
+ "grad_norm": 0.6837686896324158,
+ "learning_rate": 0.0001982401148850816,
+ "loss": 0.42303115844726563,
+ "mean_token_accuracy": 0.8718802312016487,
+ "num_tokens": 2593727.0,
+ "step": 1100
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.5167068694531918,
+ "eval_loss": 0.6138148307800293,
+ "eval_mean_token_accuracy": 0.8335792717337608,
+ "eval_num_tokens": 2645157.0,
+ "eval_runtime": 98.1828,
+ "eval_samples_per_second": 16.276,
+ "eval_steps_per_second": 2.037,
+ "step": 1122
+ },
+ {
+ "entropy": 0.41590692727553724,
+ "epoch": 3.074966532797858,
+ "grad_norm": 0.8835903406143188,
+ "learning_rate": 0.00019489160681598466,
+ "loss": 0.35146915435791015,
+ "mean_token_accuracy": 0.8911325624494841,
+ "num_tokens": 2715241.0,
+ "step": 1150
+ },
+ {
+ "entropy": 0.385117310360074,
+ "epoch": 3.208835341365462,
+ "grad_norm": 0.7755147814750671,
+ "learning_rate": 0.0001913612885560138,
+ "loss": 0.31094758987426757,
+ "mean_token_accuracy": 0.8991402065753937,
+ "num_tokens": 2828619.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.38140676759183406,
+ "epoch": 3.3427041499330654,
+ "grad_norm": 0.7542359828948975,
+ "learning_rate": 0.00018765684692270684,
+ "loss": 0.313527889251709,
+ "mean_token_accuracy": 0.8987139958143234,
+ "num_tokens": 2950688.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.38662350915372373,
+ "epoch": 3.4765729585006695,
+ "grad_norm": 0.6831114888191223,
+ "learning_rate": 0.00018378634786502866,
+ "loss": 0.31917388916015627,
+ "mean_token_accuracy": 0.8976927849650383,
+ "num_tokens": 3071943.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.38961954444646835,
+ "epoch": 3.610441767068273,
+ "grad_norm": 0.7666288018226624,
+ "learning_rate": 0.00017975821890079657,
+ "loss": 0.3270668411254883,
+ "mean_token_accuracy": 0.8966797724366188,
+ "num_tokens": 3187873.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.3998099275678396,
+ "epoch": 3.7443105756358768,
+ "grad_norm": 0.6307014226913452,
+ "learning_rate": 0.00017558123076683558,
+ "loss": 0.3317046356201172,
+ "mean_token_accuracy": 0.8944737592339516,
+ "num_tokens": 3299002.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.4006169463694096,
+ "epoch": 3.878179384203481,
+ "grad_norm": 0.7587270736694336,
+ "learning_rate": 0.0001712644783218182,
+ "loss": 0.33513580322265624,
+ "mean_token_accuracy": 0.8940666383504867,
+ "num_tokens": 3422804.0,
+ "step": 1450
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.4926549245417118,
+ "eval_loss": 0.621785581111908,
+ "eval_mean_token_accuracy": 0.8362715128064155,
+ "eval_num_tokens": 3526876.0,
+ "eval_runtime": 98.2264,
+ "eval_samples_per_second": 16.269,
+ "eval_steps_per_second": 2.036,
+ "step": 1496
+ },
+ {
+ "entropy": 0.39729254293923427,
+ "epoch": 4.010709504685408,
+ "grad_norm": 0.6503797173500061,
+ "learning_rate": 0.0001668173607433701,
+ "loss": 0.3273897171020508,
+ "mean_token_accuracy": 0.8963238362110022,
+ "num_tokens": 3536363.0,
+ "step": 1500
+ },
+ {
+ "entropy": 0.2908159454166889,
+ "epoch": 4.144578313253012,
+ "grad_norm": 0.8419110178947449,
+ "learning_rate": 0.00016224956106256037,
+ "loss": 0.21537052154541014,
+ "mean_token_accuracy": 0.9289350810647011,
+ "num_tokens": 3653560.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.29764041006565095,
+ "epoch": 4.278447121820616,
+ "grad_norm": 0.8312885165214539,
+ "learning_rate": 0.00015757102508033658,
+ "loss": 0.22335006713867187,
+ "mean_token_accuracy": 0.925491844713688,
+ "num_tokens": 3776281.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.29720947936177255,
+ "epoch": 4.412315930388219,
+ "grad_norm": 0.8321210145950317,
+ "learning_rate": 0.00015279193971181267,
+ "loss": 0.225880184173584,
+ "mean_token_accuracy": 0.9243844848871231,
+ "num_tokens": 3896081.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.30463078498840335,
+ "epoch": 4.546184738955823,
+ "grad_norm": 0.9243009686470032,
+ "learning_rate": 0.0001479227108055611,
+ "loss": 0.22961084365844728,
+ "mean_token_accuracy": 0.9234738460183144,
+ "num_tokens": 4010469.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.3072978562861681,
+ "epoch": 4.680053547523427,
+ "grad_norm": 0.6637229919433594,
+ "learning_rate": 0.00014297394048620504,
+ "loss": 0.23345823287963868,
+ "mean_token_accuracy": 0.9230251079797744,
+ "num_tokens": 4125858.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.30938181333243847,
+ "epoch": 4.813922356091031,
+ "grad_norm": 0.8940643072128296,
+ "learning_rate": 0.00013795640406964534,
+ "loss": 0.23422428131103515,
+ "mean_token_accuracy": 0.9217449072003364,
+ "num_tokens": 4243193.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.30204505123198033,
+ "epoch": 4.947791164658635,
+ "grad_norm": 0.7235498428344727,
+ "learning_rate": 0.00013288102660118477,
+ "loss": 0.22840303421020508,
+ "mean_token_accuracy": 0.9235807290673256,
+ "num_tokens": 4365873.0,
+ "step": 1850
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.4090891806781292,
+ "eval_loss": 0.6856206655502319,
+ "eval_mean_token_accuracy": 0.8357787102460861,
+ "eval_num_tokens": 4408595.0,
+ "eval_runtime": 98.466,
+ "eval_samples_per_second": 16.229,
+ "eval_steps_per_second": 2.031,
+ "step": 1870
+ },
+ {
+ "entropy": 0.25283067240709006,
+ "epoch": 5.080321285140562,
+ "grad_norm": 0.569952130317688,
+ "learning_rate": 0.00012775885906763602,
+ "loss": 0.17759113311767577,
+ "mean_token_accuracy": 0.9404868823711319,
+ "num_tokens": 4482540.0,
+ "step": 1900
+ },
+ {
+ "entropy": 0.2179089618474245,
+ "epoch": 5.214190093708166,
+ "grad_norm": 0.6869003772735596,
+ "learning_rate": 0.00012260105433520803,
+ "loss": 0.14341320037841798,
+ "mean_token_accuracy": 0.952086115181446,
+ "num_tokens": 4598191.0,
+ "step": 1950
+ },
+ {
+ "entropy": 0.2195182240009308,
+ "epoch": 5.34805890227577,
+ "grad_norm": 0.8023325800895691,
+ "learning_rate": 0.00011741884286556301,
+ "loss": 0.14362534523010254,
+ "mean_token_accuracy": 0.9521202966570854,
+ "num_tokens": 4716333.0,
+ "step": 2000
+ },
+ {
+ "entropy": 0.22079841319471596,
+ "epoch": 5.481927710843373,
+ "grad_norm": 0.7231829166412354,
+ "learning_rate": 0.00011222350826291902,
+ "loss": 0.14616458892822265,
+ "mean_token_accuracy": 0.9513887700438499,
+ "num_tokens": 4830262.0,
+ "step": 2050
+ },
+ {
+ "entropy": 0.2166673281788826,
+ "epoch": 5.615796519410977,
+ "grad_norm": 0.790205180644989,
+ "learning_rate": 0.0001070263627054418,
+ "loss": 0.14604829788208007,
+ "mean_token_accuracy": 0.9511423835158348,
+ "num_tokens": 4951007.0,
+ "step": 2100
+ },
+ {
+ "entropy": 0.22030118089169265,
+ "epoch": 5.749665327978581,
+ "grad_norm": 0.7226701974868774,
+ "learning_rate": 0.00010183872231442008,
+ "loss": 0.14739423751831054,
+ "mean_token_accuracy": 0.950632050037384,
+ "num_tokens": 5067535.0,
+ "step": 2150
+ },
+ {
+ "entropy": 0.2082419555261731,
+ "epoch": 5.883534136546185,
+ "grad_norm": 0.7288137674331665,
+ "learning_rate": 9.667188251485557e-05,
+ "loss": 0.14179450988769532,
+ "mean_token_accuracy": 0.9521883800625801,
+ "num_tokens": 5185751.0,
+ "step": 2200
+ },
+ {
+ "epoch": 6.0,
+ "eval_entropy": 0.33322301916778085,
+ "eval_loss": 0.7674635648727417,
+ "eval_mean_token_accuracy": 0.8357109513878822,
+ "eval_num_tokens": 5290314.0,
+ "eval_runtime": 98.6056,
+ "eval_samples_per_second": 16.206,
+ "eval_steps_per_second": 2.028,
+ "step": 2244
+ },
+ {
+ "entropy": 0.20694272517405374,
+ "epoch": 6.016064257028113,
+ "grad_norm": 0.48393547534942627,
+ "learning_rate": 9.15370934411162e-05,
+ "loss": 0.1373324966430664,
+ "mean_token_accuracy": 0.9537640566175635,
+ "num_tokens": 5303628.0,
+ "step": 2250
+ },
+ {
+ "entropy": 0.154647543951869,
+ "epoch": 6.149933065595716,
+ "grad_norm": 0.5554987192153931,
+ "learning_rate": 8.64455354412042e-05,
+ "loss": 0.09246989250183106,
+ "mean_token_accuracy": 0.9694107329845428,
+ "num_tokens": 5424681.0,
+ "step": 2300
+ },
+ {
+ "entropy": 0.14666760206222534,
+ "epoch": 6.28380187416332,
+ "grad_norm": 0.6542149186134338,
+ "learning_rate": 8.140829473297485e-05,
+ "loss": 0.09034320831298828,
+ "mean_token_accuracy": 0.9701463675498962,
+ "num_tokens": 5546681.0,
+ "step": 2350
+ },
+ {
+ "entropy": 0.16469813615083695,
+ "epoch": 6.417670682730924,
+ "grad_norm": 0.7045366764068604,
+ "learning_rate": 7.643633926531171e-05,
+ "loss": 0.09508686065673828,
+ "mean_token_accuracy": 0.9681181335449218,
+ "num_tokens": 5660189.0,
+ "step": 2400
+ },
+ {
+ "entropy": 0.15761504411697388,
+ "epoch": 6.551539491298527,
+ "grad_norm": 0.469072550535202,
+ "learning_rate": 7.154049483681754e-05,
+ "loss": 0.09064769744873047,
+ "mean_token_accuracy": 0.9688174912333488,
+ "num_tokens": 5780047.0,
+ "step": 2450
+ },
+ {
+ "entropy": 0.15717160735279323,
+ "epoch": 6.685408299866131,
+ "grad_norm": 0.8090623021125793,
+ "learning_rate": 6.67314215240192e-05,
+ "loss": 0.09295336723327637,
+ "mean_token_accuracy": 0.9688407072424888,
+ "num_tokens": 5897170.0,
+ "step": 2500
+ },
+ {
+ "entropy": 0.1579529893398285,
+ "epoch": 6.8192771084337345,
+ "grad_norm": 0.5264602899551392,
+ "learning_rate": 6.201959047041119e-05,
+ "loss": 0.09421208381652832,
+ "mean_token_accuracy": 0.9683848118782044,
+ "num_tokens": 6012704.0,
+ "step": 2550
+ },
+ {
+ "entropy": 0.157079309374094,
+ "epoch": 6.953145917001339,
+ "grad_norm": 0.38391751050949097,
+ "learning_rate": 5.74152610868768e-05,
+ "loss": 0.09274236679077148,
+ "mean_token_accuracy": 0.9683843395113945,
+ "num_tokens": 6132181.0,
+ "step": 2600
+ },
+ {
+ "epoch": 7.0,
+ "eval_entropy": 0.28141631513834,
+ "eval_loss": 0.8675811290740967,
+ "eval_mean_token_accuracy": 0.8351070094108581,
+ "eval_num_tokens": 6172033.0,
+ "eval_runtime": 98.8812,
+ "eval_samples_per_second": 16.161,
+ "eval_steps_per_second": 2.023,
+ "step": 2618
+ },
+ {
+ "entropy": 0.14476348296033614,
+ "epoch": 7.085676037483267,
+ "grad_norm": 0.2593551278114319,
+ "learning_rate": 5.2928458713130264e-05,
+ "loss": 0.07731590747833252,
+ "mean_token_accuracy": 0.9728465519770227,
+ "num_tokens": 6250634.0,
+ "step": 2650
+ },
+ {
+ "entropy": 0.1265121574141085,
+ "epoch": 7.21954484605087,
+ "grad_norm": 0.3758811950683594,
+ "learning_rate": 4.8568952788818444e-05,
+ "loss": 0.0695980167388916,
+ "mean_token_accuracy": 0.9759974965453148,
+ "num_tokens": 6371882.0,
+ "step": 2700
+ },
+ {
+ "entropy": 0.13779850516468287,
+ "epoch": 7.353413654618474,
+ "grad_norm": 0.4968899190425873,
+ "learning_rate": 4.43462355818128e-05,
+ "loss": 0.0733350658416748,
+ "mean_token_accuracy": 0.9740503445267678,
+ "num_tokens": 6483611.0,
+ "step": 2750
+ },
+ {
+ "entropy": 0.12721603013575078,
+ "epoch": 7.4872824631860775,
+ "grad_norm": 0.2620279788970947,
+ "learning_rate": 4.026950152000728e-05,
+ "loss": 0.07116839408874512,
+ "mean_token_accuracy": 0.9759298366308212,
+ "num_tokens": 6602543.0,
+ "step": 2800
+ },
+ {
+ "entropy": 0.1206374079361558,
+ "epoch": 7.621151271753681,
+ "grad_norm": 0.31671959161758423,
+ "learning_rate": 3.634762717162492e-05,
+ "loss": 0.0680537223815918,
+ "mean_token_accuracy": 0.9765615722537041,
+ "num_tokens": 6727899.0,
+ "step": 2850
+ },
+ {
+ "entropy": 0.12501337694004178,
+ "epoch": 7.755020080321285,
+ "grad_norm": 0.2820671796798706,
+ "learning_rate": 3.2589151917622866e-05,
+ "loss": 0.07058364391326905,
+ "mean_token_accuracy": 0.975590398311615,
+ "num_tokens": 6845842.0,
+ "step": 2900
+ },
+ {
+ "entropy": 0.12397634202614427,
+ "epoch": 7.888888888888889,
+ "grad_norm": 0.3127625286579132,
+ "learning_rate": 2.90022593582803e-05,
+ "loss": 0.07234051704406738,
+ "mean_token_accuracy": 0.9755299004912377,
+ "num_tokens": 6961910.0,
+ "step": 2950
+ },
+ {
+ "epoch": 8.0,
+ "eval_entropy": 0.2577117122709751,
+ "eval_loss": 0.9746259450912476,
+ "eval_mean_token_accuracy": 0.8363272827863694,
+ "eval_num_tokens": 7053752.0,
+ "eval_runtime": 98.6758,
+ "eval_samples_per_second": 16.194,
+ "eval_steps_per_second": 2.027,
+ "step": 2992
+ },
+ {
+ "entropy": 0.1300840698032066,
+ "epoch": 8.021419009370817,
+ "grad_norm": 0.3320488929748535,
+ "learning_rate": 2.5594759494453022e-05,
+ "loss": 0.07301389694213867,
+ "mean_token_accuracy": 0.9750576636405907,
+ "num_tokens": 7073570.0,
+ "step": 3000
+ },
+ {
+ "entropy": 0.12560308592393996,
+ "epoch": 8.15528781793842,
+ "grad_norm": 0.21358203887939453,
+ "learning_rate": 2.237407172229299e-05,
+ "loss": 0.06497196197509765,
+ "mean_token_accuracy": 0.977231368124485,
+ "num_tokens": 7186955.0,
+ "step": 3050
+ },
+ {
+ "entropy": 0.1189459022693336,
+ "epoch": 8.289156626506024,
+ "grad_norm": 0.16097290813922882,
+ "learning_rate": 1.934720867846005e-05,
+ "loss": 0.06050713062286377,
+ "mean_token_accuracy": 0.9776808521151543,
+ "num_tokens": 7307543.0,
+ "step": 3100
+ },
+ {
+ "entropy": 0.11440163658931851,
+ "epoch": 8.423025435073628,
+ "grad_norm": 0.20240066945552826,
+ "learning_rate": 1.6520760971000454e-05,
+ "loss": 0.061727256774902345,
+ "mean_token_accuracy": 0.978034191429615,
+ "num_tokens": 7425896.0,
+ "step": 3150
+ },
+ {
+ "entropy": 0.11541362870484591,
+ "epoch": 8.556894243641231,
+ "grad_norm": 0.23941750824451447,
+ "learning_rate": 1.3900882829138724e-05,
+ "loss": 0.06106637477874756,
+ "mean_token_accuracy": 0.9773423126339913,
+ "num_tokens": 7546590.0,
+ "step": 3200
+ },
+ {
+ "entropy": 0.11437087619677186,
+ "epoch": 8.690763052208835,
+ "grad_norm": 0.23929737508296967,
+ "learning_rate": 1.1493278703228827e-05,
+ "loss": 0.061713147163391116,
+ "mean_token_accuracy": 0.9784347105026245,
+ "num_tokens": 7666771.0,
+ "step": 3250
+ },
+ {
+ "entropy": 0.12138450246304273,
+ "epoch": 8.824631860776439,
+ "grad_norm": 0.15623463690280914,
+ "learning_rate": 9.303190844041513e-06,
+ "loss": 0.06417192459106445,
+ "mean_token_accuracy": 0.9768883720040321,
+ "num_tokens": 7782159.0,
+ "step": 3300
+ },
+ {
+ "entropy": 0.12119852557778359,
+ "epoch": 8.958500669344042,
+ "grad_norm": 0.15322408080101013,
+ "learning_rate": 7.33538788843202e-06,
+ "loss": 0.06281635761260987,
+ "mean_token_accuracy": 0.9767992544174194,
+ "num_tokens": 7900010.0,
+ "step": 3350
+ },
+ {
+ "epoch": 9.0,
+ "eval_entropy": 0.24177180588245392,
+ "eval_loss": 1.0658025741577148,
+ "eval_mean_token_accuracy": 0.8357376670837402,
+ "eval_num_tokens": 7935471.0,
+ "eval_runtime": 98.4754,
+ "eval_samples_per_second": 16.227,
+ "eval_steps_per_second": 2.031,
+ "step": 3366
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 3.3602237926968115e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-374/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-374/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-374/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-374/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-374/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..6bd19552d78207fd6a10c40231faf153510828db
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-374/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.08713363658978694,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-374/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-374/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-374/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-374/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-374/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-374/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-374/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-374/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..b82d59837676f0752738d10de4eeb0e20072c3c1
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-374/trainer_state.json
@@ -0,0 +1,115 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 1.0,
+ "eval_steps": 500,
+ "global_step": 374,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.3166805765032767,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.5876737833023071,
+ "learning_rate": 2.918822371674455e-05,
+ "loss": 1.1408894348144532,
+ "mean_token_accuracy": 0.7406156754493713,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.6964564691483974,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.1449841260910034,
+ "learning_rate": 5.897212546852471e-05,
+ "loss": 0.6171329116821289,
+ "mean_token_accuracy": 0.8284879937767983,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6569836577773094,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.923166811466217,
+ "learning_rate": 8.875602722030486e-05,
+ "loss": 0.5850146102905274,
+ "mean_token_accuracy": 0.8369373443722725,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6257135657966137,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.9985070824623108,
+ "learning_rate": 0.00011853992897208502,
+ "loss": 0.5537805557250977,
+ "mean_token_accuracy": 0.8436850249767304,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6312060031294823,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.8036398887634277,
+ "learning_rate": 0.00014832383072386517,
+ "loss": 0.5614564514160156,
+ "mean_token_accuracy": 0.8422514832019806,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6134552739560604,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8147028088569641,
+ "learning_rate": 0.00017810773247564532,
+ "loss": 0.5460160064697266,
+ "mean_token_accuracy": 0.8461242046952248,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.6121367686986923,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9814819693565369,
+ "learning_rate": 0.0002078916342274255,
+ "loss": 0.5481348037719727,
+ "mean_token_accuracy": 0.845237042605877,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6632432791590691,
+ "eval_loss": 0.6451922655105591,
+ "eval_mean_token_accuracy": 0.8261485403776169,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 99.2666,
+ "eval_samples_per_second": 16.098,
+ "eval_steps_per_second": 2.015,
+ "step": 374
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 3.709458252163277e+16,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3740/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3740/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3740/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3740/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3740/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..6bd19552d78207fd6a10c40231faf153510828db
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3740/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.08713363658978694,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3740/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3740/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3740/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3740/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3740/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3740/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3740/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3740/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..766538023376860572e111434da0389740b3113a
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-3740/trainer_state.json
@@ -0,0 +1,884 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 10.0,
+ "eval_steps": 500,
+ "global_step": 3740,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.3166805765032767,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.5876737833023071,
+ "learning_rate": 2.918822371674455e-05,
+ "loss": 1.1408894348144532,
+ "mean_token_accuracy": 0.7406156754493713,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.6964564691483974,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.1449841260910034,
+ "learning_rate": 5.897212546852471e-05,
+ "loss": 0.6171329116821289,
+ "mean_token_accuracy": 0.8284879937767983,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6569836577773094,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.923166811466217,
+ "learning_rate": 8.875602722030486e-05,
+ "loss": 0.5850146102905274,
+ "mean_token_accuracy": 0.8369373443722725,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6257135657966137,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.9985070824623108,
+ "learning_rate": 0.00011853992897208502,
+ "loss": 0.5537805557250977,
+ "mean_token_accuracy": 0.8436850249767304,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6312060031294823,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.8036398887634277,
+ "learning_rate": 0.00014832383072386517,
+ "loss": 0.5614564514160156,
+ "mean_token_accuracy": 0.8422514832019806,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6134552739560604,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8147028088569641,
+ "learning_rate": 0.00017810773247564532,
+ "loss": 0.5460160064697266,
+ "mean_token_accuracy": 0.8461242046952248,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.6121367686986923,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9814819693565369,
+ "learning_rate": 0.0002078916342274255,
+ "loss": 0.5481348037719727,
+ "mean_token_accuracy": 0.845237042605877,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6632432791590691,
+ "eval_loss": 0.6451922655105591,
+ "eval_mean_token_accuracy": 0.8261485403776169,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 99.2666,
+ "eval_samples_per_second": 16.098,
+ "eval_steps_per_second": 2.015,
+ "step": 374
+ },
+ {
+ "entropy": 0.6011885740239211,
+ "epoch": 1.069611780455154,
+ "grad_norm": 0.9249665141105652,
+ "learning_rate": 0.00022275326329957007,
+ "loss": 0.5334951400756835,
+ "mean_token_accuracy": 0.8480202983124088,
+ "num_tokens": 939587.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.5666237922012806,
+ "epoch": 1.2034805890227578,
+ "grad_norm": 0.7326717376708984,
+ "learning_rate": 0.00022251078790688734,
+ "loss": 0.5072500610351562,
+ "mean_token_accuracy": 0.8530087029933929,
+ "num_tokens": 1058209.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.5760806338489055,
+ "epoch": 1.3373493975903614,
+ "grad_norm": 0.7875131964683533,
+ "learning_rate": 0.00022202636508074927,
+ "loss": 0.5134101486206055,
+ "mean_token_accuracy": 0.8527275303006172,
+ "num_tokens": 1178618.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.5638319627940654,
+ "epoch": 1.4712182061579653,
+ "grad_norm": 0.9272266626358032,
+ "learning_rate": 0.00022130104959004678,
+ "loss": 0.5113970184326172,
+ "mean_token_accuracy": 0.8526131376624108,
+ "num_tokens": 1299124.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.5506764651834964,
+ "epoch": 1.605087014725569,
+ "grad_norm": 0.7173363566398621,
+ "learning_rate": 0.00022033642071670964,
+ "loss": 0.5000322723388672,
+ "mean_token_accuracy": 0.8560754117369652,
+ "num_tokens": 1422003.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.5522706513106823,
+ "epoch": 1.7389558232931726,
+ "grad_norm": 0.6712430715560913,
+ "learning_rate": 0.00021913457881702166,
+ "loss": 0.4997842788696289,
+ "mean_token_accuracy": 0.8543951433897018,
+ "num_tokens": 1542283.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.5442864654958248,
+ "epoch": 1.8728246318607764,
+ "grad_norm": 0.7373962998390198,
+ "learning_rate": 0.0002176981407483628,
+ "loss": 0.493228874206543,
+ "mean_token_accuracy": 0.8584304532408714,
+ "num_tokens": 1656994.0,
+ "step": 700
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6364928308129311,
+ "eval_loss": 0.603687584400177,
+ "eval_mean_token_accuracy": 0.8288862618803978,
+ "eval_num_tokens": 1763438.0,
+ "eval_runtime": 98.5299,
+ "eval_samples_per_second": 16.218,
+ "eval_steps_per_second": 2.03,
+ "step": 748
+ },
+ {
+ "entropy": 0.5627047226886557,
+ "epoch": 2.005354752342704,
+ "grad_norm": 0.7084948420524597,
+ "learning_rate": 0.00021603023417133613,
+ "loss": 0.49937862396240235,
+ "mean_token_accuracy": 0.8557315276126669,
+ "num_tokens": 1768737.0,
+ "step": 750
+ },
+ {
+ "entropy": 0.4591581989824772,
+ "epoch": 2.139223560910308,
+ "grad_norm": 0.6059377789497375,
+ "learning_rate": 0.00021413449073968608,
+ "loss": 0.40025680541992187,
+ "mean_token_accuracy": 0.8771893605589867,
+ "num_tokens": 1888608.0,
+ "step": 800
+ },
+ {
+ "entropy": 0.47255669847130777,
+ "epoch": 2.2730923694779115,
+ "grad_norm": 0.6926851272583008,
+ "learning_rate": 0.00021201503819283564,
+ "loss": 0.4088516616821289,
+ "mean_token_accuracy": 0.8737566202878952,
+ "num_tokens": 2000711.0,
+ "step": 850
+ },
+ {
+ "entropy": 0.4772882568836212,
+ "epoch": 2.4069611780455156,
+ "grad_norm": 0.6667259335517883,
+ "learning_rate": 0.0002096764913682607,
+ "loss": 0.4158624649047852,
+ "mean_token_accuracy": 0.8743814784288406,
+ "num_tokens": 2117216.0,
+ "step": 900
+ },
+ {
+ "entropy": 0.475759482383728,
+ "epoch": 2.540829986613119,
+ "grad_norm": 0.8212043642997742,
+ "learning_rate": 0.0002071239421532701,
+ "loss": 0.41526199340820313,
+ "mean_token_accuracy": 0.8735348290205002,
+ "num_tokens": 2233392.0,
+ "step": 950
+ },
+ {
+ "entropy": 0.4801133926212788,
+ "epoch": 2.674698795180723,
+ "grad_norm": 0.7110676765441895,
+ "learning_rate": 0.00020436294839807075,
+ "loss": 0.419442253112793,
+ "mean_token_accuracy": 0.8728689536452293,
+ "num_tokens": 2349527.0,
+ "step": 1000
+ },
+ {
+ "entropy": 0.4800094600021839,
+ "epoch": 2.8085676037483265,
+ "grad_norm": 0.7389078140258789,
+ "learning_rate": 0.00020139952181425823,
+ "loss": 0.4172813415527344,
+ "mean_token_accuracy": 0.8737373587489128,
+ "num_tokens": 2473970.0,
+ "step": 1050
+ },
+ {
+ "entropy": 0.48250251173973085,
+ "epoch": 2.9424364123159306,
+ "grad_norm": 0.6837686896324158,
+ "learning_rate": 0.0001982401148850816,
+ "loss": 0.42303115844726563,
+ "mean_token_accuracy": 0.8718802312016487,
+ "num_tokens": 2593727.0,
+ "step": 1100
+ },
+ {
+ "epoch": 3.0,
+ "eval_entropy": 0.5167068694531918,
+ "eval_loss": 0.6138148307800293,
+ "eval_mean_token_accuracy": 0.8335792717337608,
+ "eval_num_tokens": 2645157.0,
+ "eval_runtime": 98.1828,
+ "eval_samples_per_second": 16.276,
+ "eval_steps_per_second": 2.037,
+ "step": 1122
+ },
+ {
+ "entropy": 0.41590692727553724,
+ "epoch": 3.074966532797858,
+ "grad_norm": 0.8835903406143188,
+ "learning_rate": 0.00019489160681598466,
+ "loss": 0.35146915435791015,
+ "mean_token_accuracy": 0.8911325624494841,
+ "num_tokens": 2715241.0,
+ "step": 1150
+ },
+ {
+ "entropy": 0.385117310360074,
+ "epoch": 3.208835341365462,
+ "grad_norm": 0.7755147814750671,
+ "learning_rate": 0.0001913612885560138,
+ "loss": 0.31094758987426757,
+ "mean_token_accuracy": 0.8991402065753937,
+ "num_tokens": 2828619.0,
+ "step": 1200
+ },
+ {
+ "entropy": 0.38140676759183406,
+ "epoch": 3.3427041499330654,
+ "grad_norm": 0.7542359828948975,
+ "learning_rate": 0.00018765684692270684,
+ "loss": 0.313527889251709,
+ "mean_token_accuracy": 0.8987139958143234,
+ "num_tokens": 2950688.0,
+ "step": 1250
+ },
+ {
+ "entropy": 0.38662350915372373,
+ "epoch": 3.4765729585006695,
+ "grad_norm": 0.6831114888191223,
+ "learning_rate": 0.00018378634786502866,
+ "loss": 0.31917388916015627,
+ "mean_token_accuracy": 0.8976927849650383,
+ "num_tokens": 3071943.0,
+ "step": 1300
+ },
+ {
+ "entropy": 0.38961954444646835,
+ "epoch": 3.610441767068273,
+ "grad_norm": 0.7666288018226624,
+ "learning_rate": 0.00017975821890079657,
+ "loss": 0.3270668411254883,
+ "mean_token_accuracy": 0.8966797724366188,
+ "num_tokens": 3187873.0,
+ "step": 1350
+ },
+ {
+ "entropy": 0.3998099275678396,
+ "epoch": 3.7443105756358768,
+ "grad_norm": 0.6307014226913452,
+ "learning_rate": 0.00017558123076683558,
+ "loss": 0.3317046356201172,
+ "mean_token_accuracy": 0.8944737592339516,
+ "num_tokens": 3299002.0,
+ "step": 1400
+ },
+ {
+ "entropy": 0.4006169463694096,
+ "epoch": 3.878179384203481,
+ "grad_norm": 0.7587270736694336,
+ "learning_rate": 0.0001712644783218182,
+ "loss": 0.33513580322265624,
+ "mean_token_accuracy": 0.8940666383504867,
+ "num_tokens": 3422804.0,
+ "step": 1450
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.4926549245417118,
+ "eval_loss": 0.621785581111908,
+ "eval_mean_token_accuracy": 0.8362715128064155,
+ "eval_num_tokens": 3526876.0,
+ "eval_runtime": 98.2264,
+ "eval_samples_per_second": 16.269,
+ "eval_steps_per_second": 2.036,
+ "step": 1496
+ },
+ {
+ "entropy": 0.39729254293923427,
+ "epoch": 4.010709504685408,
+ "grad_norm": 0.6503797173500061,
+ "learning_rate": 0.0001668173607433701,
+ "loss": 0.3273897171020508,
+ "mean_token_accuracy": 0.8963238362110022,
+ "num_tokens": 3536363.0,
+ "step": 1500
+ },
+ {
+ "entropy": 0.2908159454166889,
+ "epoch": 4.144578313253012,
+ "grad_norm": 0.8419110178947449,
+ "learning_rate": 0.00016224956106256037,
+ "loss": 0.21537052154541014,
+ "mean_token_accuracy": 0.9289350810647011,
+ "num_tokens": 3653560.0,
+ "step": 1550
+ },
+ {
+ "entropy": 0.29764041006565095,
+ "epoch": 4.278447121820616,
+ "grad_norm": 0.8312885165214539,
+ "learning_rate": 0.00015757102508033658,
+ "loss": 0.22335006713867187,
+ "mean_token_accuracy": 0.925491844713688,
+ "num_tokens": 3776281.0,
+ "step": 1600
+ },
+ {
+ "entropy": 0.29720947936177255,
+ "epoch": 4.412315930388219,
+ "grad_norm": 0.8321210145950317,
+ "learning_rate": 0.00015279193971181267,
+ "loss": 0.225880184173584,
+ "mean_token_accuracy": 0.9243844848871231,
+ "num_tokens": 3896081.0,
+ "step": 1650
+ },
+ {
+ "entropy": 0.30463078498840335,
+ "epoch": 4.546184738955823,
+ "grad_norm": 0.9243009686470032,
+ "learning_rate": 0.0001479227108055611,
+ "loss": 0.22961084365844728,
+ "mean_token_accuracy": 0.9234738460183144,
+ "num_tokens": 4010469.0,
+ "step": 1700
+ },
+ {
+ "entropy": 0.3072978562861681,
+ "epoch": 4.680053547523427,
+ "grad_norm": 0.6637229919433594,
+ "learning_rate": 0.00014297394048620504,
+ "loss": 0.23345823287963868,
+ "mean_token_accuracy": 0.9230251079797744,
+ "num_tokens": 4125858.0,
+ "step": 1750
+ },
+ {
+ "entropy": 0.30938181333243847,
+ "epoch": 4.813922356091031,
+ "grad_norm": 0.8940643072128296,
+ "learning_rate": 0.00013795640406964534,
+ "loss": 0.23422428131103515,
+ "mean_token_accuracy": 0.9217449072003364,
+ "num_tokens": 4243193.0,
+ "step": 1800
+ },
+ {
+ "entropy": 0.30204505123198033,
+ "epoch": 4.947791164658635,
+ "grad_norm": 0.7235498428344727,
+ "learning_rate": 0.00013288102660118477,
+ "loss": 0.22840303421020508,
+ "mean_token_accuracy": 0.9235807290673256,
+ "num_tokens": 4365873.0,
+ "step": 1850
+ },
+ {
+ "epoch": 5.0,
+ "eval_entropy": 0.4090891806781292,
+ "eval_loss": 0.6856206655502319,
+ "eval_mean_token_accuracy": 0.8357787102460861,
+ "eval_num_tokens": 4408595.0,
+ "eval_runtime": 98.466,
+ "eval_samples_per_second": 16.229,
+ "eval_steps_per_second": 2.031,
+ "step": 1870
+ },
+ {
+ "entropy": 0.25283067240709006,
+ "epoch": 5.080321285140562,
+ "grad_norm": 0.569952130317688,
+ "learning_rate": 0.00012775885906763602,
+ "loss": 0.17759113311767577,
+ "mean_token_accuracy": 0.9404868823711319,
+ "num_tokens": 4482540.0,
+ "step": 1900
+ },
+ {
+ "entropy": 0.2179089618474245,
+ "epoch": 5.214190093708166,
+ "grad_norm": 0.6869003772735596,
+ "learning_rate": 0.00012260105433520803,
+ "loss": 0.14341320037841798,
+ "mean_token_accuracy": 0.952086115181446,
+ "num_tokens": 4598191.0,
+ "step": 1950
+ },
+ {
+ "entropy": 0.2195182240009308,
+ "epoch": 5.34805890227577,
+ "grad_norm": 0.8023325800895691,
+ "learning_rate": 0.00011741884286556301,
+ "loss": 0.14362534523010254,
+ "mean_token_accuracy": 0.9521202966570854,
+ "num_tokens": 4716333.0,
+ "step": 2000
+ },
+ {
+ "entropy": 0.22079841319471596,
+ "epoch": 5.481927710843373,
+ "grad_norm": 0.7231829166412354,
+ "learning_rate": 0.00011222350826291902,
+ "loss": 0.14616458892822265,
+ "mean_token_accuracy": 0.9513887700438499,
+ "num_tokens": 4830262.0,
+ "step": 2050
+ },
+ {
+ "entropy": 0.2166673281788826,
+ "epoch": 5.615796519410977,
+ "grad_norm": 0.790205180644989,
+ "learning_rate": 0.0001070263627054418,
+ "loss": 0.14604829788208007,
+ "mean_token_accuracy": 0.9511423835158348,
+ "num_tokens": 4951007.0,
+ "step": 2100
+ },
+ {
+ "entropy": 0.22030118089169265,
+ "epoch": 5.749665327978581,
+ "grad_norm": 0.7226701974868774,
+ "learning_rate": 0.00010183872231442008,
+ "loss": 0.14739423751831054,
+ "mean_token_accuracy": 0.950632050037384,
+ "num_tokens": 5067535.0,
+ "step": 2150
+ },
+ {
+ "entropy": 0.2082419555261731,
+ "epoch": 5.883534136546185,
+ "grad_norm": 0.7288137674331665,
+ "learning_rate": 9.667188251485557e-05,
+ "loss": 0.14179450988769532,
+ "mean_token_accuracy": 0.9521883800625801,
+ "num_tokens": 5185751.0,
+ "step": 2200
+ },
+ {
+ "epoch": 6.0,
+ "eval_entropy": 0.33322301916778085,
+ "eval_loss": 0.7674635648727417,
+ "eval_mean_token_accuracy": 0.8357109513878822,
+ "eval_num_tokens": 5290314.0,
+ "eval_runtime": 98.6056,
+ "eval_samples_per_second": 16.206,
+ "eval_steps_per_second": 2.028,
+ "step": 2244
+ },
+ {
+ "entropy": 0.20694272517405374,
+ "epoch": 6.016064257028113,
+ "grad_norm": 0.48393547534942627,
+ "learning_rate": 9.15370934411162e-05,
+ "loss": 0.1373324966430664,
+ "mean_token_accuracy": 0.9537640566175635,
+ "num_tokens": 5303628.0,
+ "step": 2250
+ },
+ {
+ "entropy": 0.154647543951869,
+ "epoch": 6.149933065595716,
+ "grad_norm": 0.5554987192153931,
+ "learning_rate": 8.64455354412042e-05,
+ "loss": 0.09246989250183106,
+ "mean_token_accuracy": 0.9694107329845428,
+ "num_tokens": 5424681.0,
+ "step": 2300
+ },
+ {
+ "entropy": 0.14666760206222534,
+ "epoch": 6.28380187416332,
+ "grad_norm": 0.6542149186134338,
+ "learning_rate": 8.140829473297485e-05,
+ "loss": 0.09034320831298828,
+ "mean_token_accuracy": 0.9701463675498962,
+ "num_tokens": 5546681.0,
+ "step": 2350
+ },
+ {
+ "entropy": 0.16469813615083695,
+ "epoch": 6.417670682730924,
+ "grad_norm": 0.7045366764068604,
+ "learning_rate": 7.643633926531171e-05,
+ "loss": 0.09508686065673828,
+ "mean_token_accuracy": 0.9681181335449218,
+ "num_tokens": 5660189.0,
+ "step": 2400
+ },
+ {
+ "entropy": 0.15761504411697388,
+ "epoch": 6.551539491298527,
+ "grad_norm": 0.469072550535202,
+ "learning_rate": 7.154049483681754e-05,
+ "loss": 0.09064769744873047,
+ "mean_token_accuracy": 0.9688174912333488,
+ "num_tokens": 5780047.0,
+ "step": 2450
+ },
+ {
+ "entropy": 0.15717160735279323,
+ "epoch": 6.685408299866131,
+ "grad_norm": 0.8090623021125793,
+ "learning_rate": 6.67314215240192e-05,
+ "loss": 0.09295336723327637,
+ "mean_token_accuracy": 0.9688407072424888,
+ "num_tokens": 5897170.0,
+ "step": 2500
+ },
+ {
+ "entropy": 0.1579529893398285,
+ "epoch": 6.8192771084337345,
+ "grad_norm": 0.5264602899551392,
+ "learning_rate": 6.201959047041119e-05,
+ "loss": 0.09421208381652832,
+ "mean_token_accuracy": 0.9683848118782044,
+ "num_tokens": 6012704.0,
+ "step": 2550
+ },
+ {
+ "entropy": 0.157079309374094,
+ "epoch": 6.953145917001339,
+ "grad_norm": 0.38391751050949097,
+ "learning_rate": 5.74152610868768e-05,
+ "loss": 0.09274236679077148,
+ "mean_token_accuracy": 0.9683843395113945,
+ "num_tokens": 6132181.0,
+ "step": 2600
+ },
+ {
+ "epoch": 7.0,
+ "eval_entropy": 0.28141631513834,
+ "eval_loss": 0.8675811290740967,
+ "eval_mean_token_accuracy": 0.8351070094108581,
+ "eval_num_tokens": 6172033.0,
+ "eval_runtime": 98.8812,
+ "eval_samples_per_second": 16.161,
+ "eval_steps_per_second": 2.023,
+ "step": 2618
+ },
+ {
+ "entropy": 0.14476348296033614,
+ "epoch": 7.085676037483267,
+ "grad_norm": 0.2593551278114319,
+ "learning_rate": 5.2928458713130264e-05,
+ "loss": 0.07731590747833252,
+ "mean_token_accuracy": 0.9728465519770227,
+ "num_tokens": 6250634.0,
+ "step": 2650
+ },
+ {
+ "entropy": 0.1265121574141085,
+ "epoch": 7.21954484605087,
+ "grad_norm": 0.3758811950683594,
+ "learning_rate": 4.8568952788818444e-05,
+ "loss": 0.0695980167388916,
+ "mean_token_accuracy": 0.9759974965453148,
+ "num_tokens": 6371882.0,
+ "step": 2700
+ },
+ {
+ "entropy": 0.13779850516468287,
+ "epoch": 7.353413654618474,
+ "grad_norm": 0.4968899190425873,
+ "learning_rate": 4.43462355818128e-05,
+ "loss": 0.0733350658416748,
+ "mean_token_accuracy": 0.9740503445267678,
+ "num_tokens": 6483611.0,
+ "step": 2750
+ },
+ {
+ "entropy": 0.12721603013575078,
+ "epoch": 7.4872824631860775,
+ "grad_norm": 0.2620279788970947,
+ "learning_rate": 4.026950152000728e-05,
+ "loss": 0.07116839408874512,
+ "mean_token_accuracy": 0.9759298366308212,
+ "num_tokens": 6602543.0,
+ "step": 2800
+ },
+ {
+ "entropy": 0.1206374079361558,
+ "epoch": 7.621151271753681,
+ "grad_norm": 0.31671959161758423,
+ "learning_rate": 3.634762717162492e-05,
+ "loss": 0.0680537223815918,
+ "mean_token_accuracy": 0.9765615722537041,
+ "num_tokens": 6727899.0,
+ "step": 2850
+ },
+ {
+ "entropy": 0.12501337694004178,
+ "epoch": 7.755020080321285,
+ "grad_norm": 0.2820671796798706,
+ "learning_rate": 3.2589151917622866e-05,
+ "loss": 0.07058364391326905,
+ "mean_token_accuracy": 0.975590398311615,
+ "num_tokens": 6845842.0,
+ "step": 2900
+ },
+ {
+ "entropy": 0.12397634202614427,
+ "epoch": 7.888888888888889,
+ "grad_norm": 0.3127625286579132,
+ "learning_rate": 2.90022593582803e-05,
+ "loss": 0.07234051704406738,
+ "mean_token_accuracy": 0.9755299004912377,
+ "num_tokens": 6961910.0,
+ "step": 2950
+ },
+ {
+ "epoch": 8.0,
+ "eval_entropy": 0.2577117122709751,
+ "eval_loss": 0.9746259450912476,
+ "eval_mean_token_accuracy": 0.8363272827863694,
+ "eval_num_tokens": 7053752.0,
+ "eval_runtime": 98.6758,
+ "eval_samples_per_second": 16.194,
+ "eval_steps_per_second": 2.027,
+ "step": 2992
+ },
+ {
+ "entropy": 0.1300840698032066,
+ "epoch": 8.021419009370817,
+ "grad_norm": 0.3320488929748535,
+ "learning_rate": 2.5594759494453022e-05,
+ "loss": 0.07301389694213867,
+ "mean_token_accuracy": 0.9750576636405907,
+ "num_tokens": 7073570.0,
+ "step": 3000
+ },
+ {
+ "entropy": 0.12560308592393996,
+ "epoch": 8.15528781793842,
+ "grad_norm": 0.21358203887939453,
+ "learning_rate": 2.237407172229299e-05,
+ "loss": 0.06497196197509765,
+ "mean_token_accuracy": 0.977231368124485,
+ "num_tokens": 7186955.0,
+ "step": 3050
+ },
+ {
+ "entropy": 0.1189459022693336,
+ "epoch": 8.289156626506024,
+ "grad_norm": 0.16097290813922882,
+ "learning_rate": 1.934720867846005e-05,
+ "loss": 0.06050713062286377,
+ "mean_token_accuracy": 0.9776808521151543,
+ "num_tokens": 7307543.0,
+ "step": 3100
+ },
+ {
+ "entropy": 0.11440163658931851,
+ "epoch": 8.423025435073628,
+ "grad_norm": 0.20240066945552826,
+ "learning_rate": 1.6520760971000454e-05,
+ "loss": 0.061727256774902345,
+ "mean_token_accuracy": 0.978034191429615,
+ "num_tokens": 7425896.0,
+ "step": 3150
+ },
+ {
+ "entropy": 0.11541362870484591,
+ "epoch": 8.556894243641231,
+ "grad_norm": 0.23941750824451447,
+ "learning_rate": 1.3900882829138724e-05,
+ "loss": 0.06106637477874756,
+ "mean_token_accuracy": 0.9773423126339913,
+ "num_tokens": 7546590.0,
+ "step": 3200
+ },
+ {
+ "entropy": 0.11437087619677186,
+ "epoch": 8.690763052208835,
+ "grad_norm": 0.23929737508296967,
+ "learning_rate": 1.1493278703228827e-05,
+ "loss": 0.061713147163391116,
+ "mean_token_accuracy": 0.9784347105026245,
+ "num_tokens": 7666771.0,
+ "step": 3250
+ },
+ {
+ "entropy": 0.12138450246304273,
+ "epoch": 8.824631860776439,
+ "grad_norm": 0.15623463690280914,
+ "learning_rate": 9.303190844041513e-06,
+ "loss": 0.06417192459106445,
+ "mean_token_accuracy": 0.9768883720040321,
+ "num_tokens": 7782159.0,
+ "step": 3300
+ },
+ {
+ "entropy": 0.12119852557778359,
+ "epoch": 8.958500669344042,
+ "grad_norm": 0.15322408080101013,
+ "learning_rate": 7.33538788843202e-06,
+ "loss": 0.06281635761260987,
+ "mean_token_accuracy": 0.9767992544174194,
+ "num_tokens": 7900010.0,
+ "step": 3350
+ },
+ {
+ "epoch": 9.0,
+ "eval_entropy": 0.24177180588245392,
+ "eval_loss": 1.0658025741577148,
+ "eval_mean_token_accuracy": 0.8357376670837402,
+ "eval_num_tokens": 7935471.0,
+ "eval_runtime": 98.4754,
+ "eval_samples_per_second": 16.227,
+ "eval_steps_per_second": 2.031,
+ "step": 3366
+ },
+ {
+ "entropy": 0.12643026225645132,
+ "epoch": 9.09103078982597,
+ "grad_norm": 0.20806892216205597,
+ "learning_rate": 5.594154476241933e-06,
+ "loss": 0.06305371761322022,
+ "mean_token_accuracy": 0.9766921066876614,
+ "num_tokens": 8010103.0,
+ "step": 3400
+ },
+ {
+ "entropy": 0.1165965672209859,
+ "epoch": 9.224899598393574,
+ "grad_norm": 0.18308691680431366,
+ "learning_rate": 4.083281921042599e-06,
+ "loss": 0.06130090713500977,
+ "mean_token_accuracy": 0.9778196820616722,
+ "num_tokens": 8122846.0,
+ "step": 3450
+ },
+ {
+ "entropy": 0.11787618484348059,
+ "epoch": 9.358768406961179,
+ "grad_norm": 0.19463828206062317,
+ "learning_rate": 2.806059955033672e-06,
+ "loss": 0.05993570327758789,
+ "mean_token_accuracy": 0.9778310161828995,
+ "num_tokens": 8239477.0,
+ "step": 3500
+ },
+ {
+ "entropy": 0.10821869468316436,
+ "epoch": 9.492637215528783,
+ "grad_norm": 0.14283855259418488,
+ "learning_rate": 1.7652695660708882e-06,
+ "loss": 0.05648116111755371,
+ "mean_token_accuracy": 0.9792589458823204,
+ "num_tokens": 8363547.0,
+ "step": 3550
+ },
+ {
+ "entropy": 0.11550636570900678,
+ "epoch": 9.626506024096386,
+ "grad_norm": 0.2377292960882187,
+ "learning_rate": 9.63176942419982e-07,
+ "loss": 0.05863102436065674,
+ "mean_token_accuracy": 0.978409596979618,
+ "num_tokens": 8482651.0,
+ "step": 3600
+ },
+ {
+ "entropy": 0.11191969187930226,
+ "epoch": 9.76037483266399,
+ "grad_norm": 0.2096184492111206,
+ "learning_rate": 4.015285384208993e-07,
+ "loss": 0.05814181327819824,
+ "mean_token_accuracy": 0.9789791169762612,
+ "num_tokens": 8601806.0,
+ "step": 3650
+ },
+ {
+ "entropy": 0.1046653332747519,
+ "epoch": 9.894243641231594,
+ "grad_norm": 0.2754163444042206,
+ "learning_rate": 8.154727180626226e-08,
+ "loss": 0.05543069839477539,
+ "mean_token_accuracy": 0.9802288645505906,
+ "num_tokens": 8728357.0,
+ "step": 3700
+ },
+ {
+ "epoch": 10.0,
+ "eval_entropy": 0.23624430641531943,
+ "eval_loss": 1.1085858345031738,
+ "eval_mean_token_accuracy": 0.8358771416544915,
+ "eval_num_tokens": 8817190.0,
+ "eval_runtime": 98.8123,
+ "eval_samples_per_second": 16.172,
+ "eval_steps_per_second": 2.024,
+ "step": 3740
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": true
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 3.7345927679778816e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-748/README.md b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-748/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-748/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-748/adapter_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-748/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..6bd19552d78207fd6a10c40231faf153510828db
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-748/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.08713363658978694,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "o_proj",
+ "gate_proj",
+ "up_proj",
+ "k_proj",
+ "v_proj",
+ "down_proj",
+ "q_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-748/chat_template.jinja b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-748/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-748/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-748/tokenizer_config.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-748/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-748/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-748/trainer_state.json b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-748/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..a23b4b7063e2707b1724c6d3956a751f5b6d883f
--- /dev/null
+++ b/DBCA_original_Swedish/Qwen3.5-4B-Base_original_features_structural_train_original_features_structural_test2/checkpoint-748/trainer_state.json
@@ -0,0 +1,196 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 2.0,
+ "eval_steps": 500,
+ "global_step": 748,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 1.3166805765032767,
+ "epoch": 0.13386880856760375,
+ "grad_norm": 1.5876737833023071,
+ "learning_rate": 2.918822371674455e-05,
+ "loss": 1.1408894348144532,
+ "mean_token_accuracy": 0.7406156754493713,
+ "num_tokens": 117463.0,
+ "step": 50
+ },
+ {
+ "entropy": 0.6964564691483974,
+ "epoch": 0.2677376171352075,
+ "grad_norm": 1.1449841260910034,
+ "learning_rate": 5.897212546852471e-05,
+ "loss": 0.6171329116821289,
+ "mean_token_accuracy": 0.8284879937767983,
+ "num_tokens": 235397.0,
+ "step": 100
+ },
+ {
+ "entropy": 0.6569836577773094,
+ "epoch": 0.40160642570281124,
+ "grad_norm": 0.923166811466217,
+ "learning_rate": 8.875602722030486e-05,
+ "loss": 0.5850146102905274,
+ "mean_token_accuracy": 0.8369373443722725,
+ "num_tokens": 356166.0,
+ "step": 150
+ },
+ {
+ "entropy": 0.6257135657966137,
+ "epoch": 0.535475234270415,
+ "grad_norm": 0.9985070824623108,
+ "learning_rate": 0.00011853992897208502,
+ "loss": 0.5537805557250977,
+ "mean_token_accuracy": 0.8436850249767304,
+ "num_tokens": 479742.0,
+ "step": 200
+ },
+ {
+ "entropy": 0.6312060031294823,
+ "epoch": 0.6693440428380187,
+ "grad_norm": 0.8036398887634277,
+ "learning_rate": 0.00014832383072386517,
+ "loss": 0.5614564514160156,
+ "mean_token_accuracy": 0.8422514832019806,
+ "num_tokens": 595619.0,
+ "step": 250
+ },
+ {
+ "entropy": 0.6134552739560604,
+ "epoch": 0.8032128514056225,
+ "grad_norm": 0.8147028088569641,
+ "learning_rate": 0.00017810773247564532,
+ "loss": 0.5460160064697266,
+ "mean_token_accuracy": 0.8461242046952248,
+ "num_tokens": 714771.0,
+ "step": 300
+ },
+ {
+ "entropy": 0.6121367686986923,
+ "epoch": 0.9370816599732262,
+ "grad_norm": 0.9814819693565369,
+ "learning_rate": 0.0002078916342274255,
+ "loss": 0.5481348037719727,
+ "mean_token_accuracy": 0.845237042605877,
+ "num_tokens": 831791.0,
+ "step": 350
+ },
+ {
+ "epoch": 1.0,
+ "eval_entropy": 0.6632432791590691,
+ "eval_loss": 0.6451922655105591,
+ "eval_mean_token_accuracy": 0.8261485403776169,
+ "eval_num_tokens": 881719.0,
+ "eval_runtime": 99.2666,
+ "eval_samples_per_second": 16.098,
+ "eval_steps_per_second": 2.015,
+ "step": 374
+ },
+ {
+ "entropy": 0.6011885740239211,
+ "epoch": 1.069611780455154,
+ "grad_norm": 0.9249665141105652,
+ "learning_rate": 0.00022275326329957007,
+ "loss": 0.5334951400756835,
+ "mean_token_accuracy": 0.8480202983124088,
+ "num_tokens": 939587.0,
+ "step": 400
+ },
+ {
+ "entropy": 0.5666237922012806,
+ "epoch": 1.2034805890227578,
+ "grad_norm": 0.7326717376708984,
+ "learning_rate": 0.00022251078790688734,
+ "loss": 0.5072500610351562,
+ "mean_token_accuracy": 0.8530087029933929,
+ "num_tokens": 1058209.0,
+ "step": 450
+ },
+ {
+ "entropy": 0.5760806338489055,
+ "epoch": 1.3373493975903614,
+ "grad_norm": 0.7875131964683533,
+ "learning_rate": 0.00022202636508074927,
+ "loss": 0.5134101486206055,
+ "mean_token_accuracy": 0.8527275303006172,
+ "num_tokens": 1178618.0,
+ "step": 500
+ },
+ {
+ "entropy": 0.5638319627940654,
+ "epoch": 1.4712182061579653,
+ "grad_norm": 0.9272266626358032,
+ "learning_rate": 0.00022130104959004678,
+ "loss": 0.5113970184326172,
+ "mean_token_accuracy": 0.8526131376624108,
+ "num_tokens": 1299124.0,
+ "step": 550
+ },
+ {
+ "entropy": 0.5506764651834964,
+ "epoch": 1.605087014725569,
+ "grad_norm": 0.7173363566398621,
+ "learning_rate": 0.00022033642071670964,
+ "loss": 0.5000322723388672,
+ "mean_token_accuracy": 0.8560754117369652,
+ "num_tokens": 1422003.0,
+ "step": 600
+ },
+ {
+ "entropy": 0.5522706513106823,
+ "epoch": 1.7389558232931726,
+ "grad_norm": 0.6712430715560913,
+ "learning_rate": 0.00021913457881702166,
+ "loss": 0.4997842788696289,
+ "mean_token_accuracy": 0.8543951433897018,
+ "num_tokens": 1542283.0,
+ "step": 650
+ },
+ {
+ "entropy": 0.5442864654958248,
+ "epoch": 1.8728246318607764,
+ "grad_norm": 0.7373962998390198,
+ "learning_rate": 0.0002176981407483628,
+ "loss": 0.493228874206543,
+ "mean_token_accuracy": 0.8584304532408714,
+ "num_tokens": 1656994.0,
+ "step": 700
+ },
+ {
+ "epoch": 2.0,
+ "eval_entropy": 0.6364928308129311,
+ "eval_loss": 0.603687584400177,
+ "eval_mean_token_accuracy": 0.8288862618803978,
+ "eval_num_tokens": 1763438.0,
+ "eval_runtime": 98.5299,
+ "eval_samples_per_second": 16.218,
+ "eval_steps_per_second": 2.03,
+ "step": 748
+ }
+ ],
+ "logging_steps": 50,
+ "max_steps": 3740,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 500,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 7.45087263678935e+16,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1820/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1820/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1820/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1820/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1820/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1820/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1820/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1820/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1820/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1820/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1820/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1820/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1820/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1820/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..60baca57750668752150904d8f67c39de6483cff
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1820/trainer_state.json
@@ -0,0 +1,1945 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 4.386240193120097,
+ "eval_steps": 20,
+ "global_step": 1820,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.8229614856578048e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1840/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1840/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1840/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1840/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1840/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1840/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1840/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1840/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1840/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1840/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1840/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1840/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1840/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1840/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..4f702cff36ef984d8ce1f02dbb18a8f0b3a239eb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1840/trainer_state.json
@@ -0,0 +1,1966 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 4.434520217260109,
+ "eval_steps": 20,
+ "global_step": 1840,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.8423384672632218e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1860/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1860/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1860/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1860/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1860/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1860/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1860/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1860/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1860/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1860/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1860/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1860/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1860/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1860/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..db28671dd80b048e561f06b3c215b0ee20e95ac2
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1860/trainer_state.json
@@ -0,0 +1,1987 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 4.4828002414001205,
+ "eval_steps": 20,
+ "global_step": 1860,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.8636111402319258e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1880/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1880/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1880/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1880/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1880/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1880/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1880/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1880/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1880/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1880/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1880/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1880/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1880/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1880/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..f0c6f09137df96fad16441a118820a06dd1eda16
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1880/trainer_state.json
@@ -0,0 +1,2008 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 4.531080265540133,
+ "eval_steps": 20,
+ "global_step": 1880,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.884319234664878e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1900/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1900/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1900/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1900/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1900/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1900/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1900/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1900/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1900/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1900/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1900/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1900/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1900/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1900/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..edb5a2a788f0a3d93979aeaf45c3bf778a4a234a
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1900/trainer_state.json
@@ -0,0 +1,2029 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 4.579360289680145,
+ "eval_steps": 20,
+ "global_step": 1900,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ },
+ {
+ "entropy": 0.4528769593685865,
+ "epoch": 4.579360289680145,
+ "grad_norm": 0.7440020442008972,
+ "learning_rate": 0.0002401706127645863,
+ "loss": 0.43678741455078124,
+ "mean_token_accuracy": 0.8619599856436253,
+ "num_tokens": 4401934.0,
+ "step": 1900
+ },
+ {
+ "epoch": 4.579360289680145,
+ "eval_entropy": 0.5085621437664782,
+ "eval_loss": 0.6766273379325867,
+ "eval_mean_token_accuracy": 0.8227389695939054,
+ "eval_num_tokens": 4401934.0,
+ "eval_runtime": 89.4661,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 1900
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.9031154812740608e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1920/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1920/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1920/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1920/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1920/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1920/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1920/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1920/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1920/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1920/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1920/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1920/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1920/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1920/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..ded7332eb2077605e94ea0ae49268f5a5ca71904
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1920/trainer_state.json
@@ -0,0 +1,2050 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 4.627640313820157,
+ "eval_steps": 20,
+ "global_step": 1920,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ },
+ {
+ "entropy": 0.4528769593685865,
+ "epoch": 4.579360289680145,
+ "grad_norm": 0.7440020442008972,
+ "learning_rate": 0.0002401706127645863,
+ "loss": 0.43678741455078124,
+ "mean_token_accuracy": 0.8619599856436253,
+ "num_tokens": 4401934.0,
+ "step": 1900
+ },
+ {
+ "epoch": 4.579360289680145,
+ "eval_entropy": 0.5085621437664782,
+ "eval_loss": 0.6766273379325867,
+ "eval_mean_token_accuracy": 0.8227389695939054,
+ "eval_num_tokens": 4401934.0,
+ "eval_runtime": 89.4661,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 1900
+ },
+ {
+ "entropy": 0.4356316838413477,
+ "epoch": 4.627640313820157,
+ "grad_norm": 1.3870011568069458,
+ "learning_rate": 0.0002372530330023795,
+ "loss": 0.41338543891906737,
+ "mean_token_accuracy": 0.8642870113253593,
+ "num_tokens": 4449745.0,
+ "step": 1920
+ },
+ {
+ "epoch": 4.627640313820157,
+ "eval_entropy": 0.4954608427674583,
+ "eval_loss": 0.6783625483512878,
+ "eval_mean_token_accuracy": 0.8214637864841504,
+ "eval_num_tokens": 4449745.0,
+ "eval_runtime": 89.2291,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 1920
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.9229771566939546e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1940/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1940/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1940/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1940/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1940/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1940/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1940/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1940/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1940/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1940/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1940/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1940/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1940/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1940/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..2f9ca7cee6c4a69858d8d05e732d33c3fdae932c
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1940/trainer_state.json
@@ -0,0 +1,2071 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 4.675920337960169,
+ "eval_steps": 20,
+ "global_step": 1940,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ },
+ {
+ "entropy": 0.4528769593685865,
+ "epoch": 4.579360289680145,
+ "grad_norm": 0.7440020442008972,
+ "learning_rate": 0.0002401706127645863,
+ "loss": 0.43678741455078124,
+ "mean_token_accuracy": 0.8619599856436253,
+ "num_tokens": 4401934.0,
+ "step": 1900
+ },
+ {
+ "epoch": 4.579360289680145,
+ "eval_entropy": 0.5085621437664782,
+ "eval_loss": 0.6766273379325867,
+ "eval_mean_token_accuracy": 0.8227389695939054,
+ "eval_num_tokens": 4401934.0,
+ "eval_runtime": 89.4661,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 1900
+ },
+ {
+ "entropy": 0.4356316838413477,
+ "epoch": 4.627640313820157,
+ "grad_norm": 1.3870011568069458,
+ "learning_rate": 0.0002372530330023795,
+ "loss": 0.41338543891906737,
+ "mean_token_accuracy": 0.8642870113253593,
+ "num_tokens": 4449745.0,
+ "step": 1920
+ },
+ {
+ "epoch": 4.627640313820157,
+ "eval_entropy": 0.4954608427674583,
+ "eval_loss": 0.6783625483512878,
+ "eval_mean_token_accuracy": 0.8214637864841504,
+ "eval_num_tokens": 4449745.0,
+ "eval_runtime": 89.2291,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 1920
+ },
+ {
+ "entropy": 0.4468349639326334,
+ "epoch": 4.675920337960169,
+ "grad_norm": 0.8028655648231506,
+ "learning_rate": 0.00023431991948050538,
+ "loss": 0.4324653625488281,
+ "mean_token_accuracy": 0.8607208795845509,
+ "num_tokens": 4495074.0,
+ "step": 1940
+ },
+ {
+ "epoch": 4.675920337960169,
+ "eval_entropy": 0.5259480836351266,
+ "eval_loss": 0.6809864044189453,
+ "eval_mean_token_accuracy": 0.8200277019752545,
+ "eval_num_tokens": 4495074.0,
+ "eval_runtime": 89.8247,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 1940
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.9425677626101965e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1960/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1960/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1960/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1960/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1960/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1960/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1960/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1960/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1960/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1960/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1960/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1960/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1960/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1960/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..8dc77ca944be4fb1440b6eee2236f9ad52dce45e
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1960/trainer_state.json
@@ -0,0 +1,2092 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 4.724200362100181,
+ "eval_steps": 20,
+ "global_step": 1960,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ },
+ {
+ "entropy": 0.4528769593685865,
+ "epoch": 4.579360289680145,
+ "grad_norm": 0.7440020442008972,
+ "learning_rate": 0.0002401706127645863,
+ "loss": 0.43678741455078124,
+ "mean_token_accuracy": 0.8619599856436253,
+ "num_tokens": 4401934.0,
+ "step": 1900
+ },
+ {
+ "epoch": 4.579360289680145,
+ "eval_entropy": 0.5085621437664782,
+ "eval_loss": 0.6766273379325867,
+ "eval_mean_token_accuracy": 0.8227389695939054,
+ "eval_num_tokens": 4401934.0,
+ "eval_runtime": 89.4661,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 1900
+ },
+ {
+ "entropy": 0.4356316838413477,
+ "epoch": 4.627640313820157,
+ "grad_norm": 1.3870011568069458,
+ "learning_rate": 0.0002372530330023795,
+ "loss": 0.41338543891906737,
+ "mean_token_accuracy": 0.8642870113253593,
+ "num_tokens": 4449745.0,
+ "step": 1920
+ },
+ {
+ "epoch": 4.627640313820157,
+ "eval_entropy": 0.4954608427674583,
+ "eval_loss": 0.6783625483512878,
+ "eval_mean_token_accuracy": 0.8214637864841504,
+ "eval_num_tokens": 4449745.0,
+ "eval_runtime": 89.2291,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 1920
+ },
+ {
+ "entropy": 0.4468349639326334,
+ "epoch": 4.675920337960169,
+ "grad_norm": 0.8028655648231506,
+ "learning_rate": 0.00023431991948050538,
+ "loss": 0.4324653625488281,
+ "mean_token_accuracy": 0.8607208795845509,
+ "num_tokens": 4495074.0,
+ "step": 1940
+ },
+ {
+ "epoch": 4.675920337960169,
+ "eval_entropy": 0.5259480836351266,
+ "eval_loss": 0.6809864044189453,
+ "eval_mean_token_accuracy": 0.8200277019752545,
+ "eval_num_tokens": 4495074.0,
+ "eval_runtime": 89.8247,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 1940
+ },
+ {
+ "entropy": 0.4610355503857136,
+ "epoch": 4.724200362100181,
+ "grad_norm": 0.8928298950195312,
+ "learning_rate": 0.0002313721022353953,
+ "loss": 0.4314274787902832,
+ "mean_token_accuracy": 0.8559959702193737,
+ "num_tokens": 4541053.0,
+ "step": 1960
+ },
+ {
+ "epoch": 4.724200362100181,
+ "eval_entropy": 0.501018373651451,
+ "eval_loss": 0.6834176182746887,
+ "eval_mean_token_accuracy": 0.8220113566082515,
+ "eval_num_tokens": 4541053.0,
+ "eval_runtime": 90.4088,
+ "eval_samples_per_second": 15.706,
+ "eval_steps_per_second": 1.969,
+ "step": 1960
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.9632246949182874e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1980/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1980/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1980/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1980/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1980/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1980/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1980/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1980/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1980/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1980/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1980/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1980/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1980/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1980/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..f2ae77b19eb9fb30511407c4acd0354e5d510a45
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-1980/trainer_state.json
@@ -0,0 +1,2113 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 4.772480386240193,
+ "eval_steps": 20,
+ "global_step": 1980,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ },
+ {
+ "entropy": 0.4528769593685865,
+ "epoch": 4.579360289680145,
+ "grad_norm": 0.7440020442008972,
+ "learning_rate": 0.0002401706127645863,
+ "loss": 0.43678741455078124,
+ "mean_token_accuracy": 0.8619599856436253,
+ "num_tokens": 4401934.0,
+ "step": 1900
+ },
+ {
+ "epoch": 4.579360289680145,
+ "eval_entropy": 0.5085621437664782,
+ "eval_loss": 0.6766273379325867,
+ "eval_mean_token_accuracy": 0.8227389695939054,
+ "eval_num_tokens": 4401934.0,
+ "eval_runtime": 89.4661,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 1900
+ },
+ {
+ "entropy": 0.4356316838413477,
+ "epoch": 4.627640313820157,
+ "grad_norm": 1.3870011568069458,
+ "learning_rate": 0.0002372530330023795,
+ "loss": 0.41338543891906737,
+ "mean_token_accuracy": 0.8642870113253593,
+ "num_tokens": 4449745.0,
+ "step": 1920
+ },
+ {
+ "epoch": 4.627640313820157,
+ "eval_entropy": 0.4954608427674583,
+ "eval_loss": 0.6783625483512878,
+ "eval_mean_token_accuracy": 0.8214637864841504,
+ "eval_num_tokens": 4449745.0,
+ "eval_runtime": 89.2291,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 1920
+ },
+ {
+ "entropy": 0.4468349639326334,
+ "epoch": 4.675920337960169,
+ "grad_norm": 0.8028655648231506,
+ "learning_rate": 0.00023431991948050538,
+ "loss": 0.4324653625488281,
+ "mean_token_accuracy": 0.8607208795845509,
+ "num_tokens": 4495074.0,
+ "step": 1940
+ },
+ {
+ "epoch": 4.675920337960169,
+ "eval_entropy": 0.5259480836351266,
+ "eval_loss": 0.6809864044189453,
+ "eval_mean_token_accuracy": 0.8200277019752545,
+ "eval_num_tokens": 4495074.0,
+ "eval_runtime": 89.8247,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 1940
+ },
+ {
+ "entropy": 0.4610355503857136,
+ "epoch": 4.724200362100181,
+ "grad_norm": 0.8928298950195312,
+ "learning_rate": 0.0002313721022353953,
+ "loss": 0.4314274787902832,
+ "mean_token_accuracy": 0.8559959702193737,
+ "num_tokens": 4541053.0,
+ "step": 1960
+ },
+ {
+ "epoch": 4.724200362100181,
+ "eval_entropy": 0.501018373651451,
+ "eval_loss": 0.6834176182746887,
+ "eval_mean_token_accuracy": 0.8220113566082515,
+ "eval_num_tokens": 4541053.0,
+ "eval_runtime": 90.4088,
+ "eval_samples_per_second": 15.706,
+ "eval_steps_per_second": 1.969,
+ "step": 1960
+ },
+ {
+ "entropy": 0.4568953149020672,
+ "epoch": 4.772480386240193,
+ "grad_norm": 1.018485188484192,
+ "learning_rate": 0.00022841041546446032,
+ "loss": 0.4306424617767334,
+ "mean_token_accuracy": 0.8598173558712006,
+ "num_tokens": 4584210.0,
+ "step": 1980
+ },
+ {
+ "epoch": 4.772480386240193,
+ "eval_entropy": 0.47633055371514865,
+ "eval_loss": 0.6886085867881775,
+ "eval_mean_token_accuracy": 0.823554899585381,
+ "eval_num_tokens": 4584210.0,
+ "eval_runtime": 89.8185,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 1980
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 1.9816260058267853e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-20/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-20/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-20/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-20/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-20/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-20/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-20/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-20/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-20/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-20/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-20/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-20/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-20/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-20/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..a4dcf7860ff3e7b3219b930f1adcd5b38c0c19fc
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-20/trainer_state.json
@@ -0,0 +1,55 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 0.04828002414001207,
+ "eval_steps": 20,
+ "global_step": 20,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2026648451309568.0,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-200/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-200/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-200/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-200/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-200/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-200/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-200/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-200/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-200/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-200/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-200/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-200/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-200/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-200/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..933ef7c29891e90fef38445e9003662bd3921481
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-200/trainer_state.json
@@ -0,0 +1,244 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 0.4828002414001207,
+ "eval_steps": 20,
+ "global_step": 200,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.049482915458253e+16,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2000/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2000/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2000/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2000/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2000/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2000/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2000/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2000/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2000/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2000/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2000/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2000/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2000/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2000/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..23dd9bbbfcafe9de9ce7802bae25ea8c257d382f
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2000/trainer_state.json
@@ -0,0 +1,2134 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 4.820760410380205,
+ "eval_steps": 20,
+ "global_step": 2000,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ },
+ {
+ "entropy": 0.4528769593685865,
+ "epoch": 4.579360289680145,
+ "grad_norm": 0.7440020442008972,
+ "learning_rate": 0.0002401706127645863,
+ "loss": 0.43678741455078124,
+ "mean_token_accuracy": 0.8619599856436253,
+ "num_tokens": 4401934.0,
+ "step": 1900
+ },
+ {
+ "epoch": 4.579360289680145,
+ "eval_entropy": 0.5085621437664782,
+ "eval_loss": 0.6766273379325867,
+ "eval_mean_token_accuracy": 0.8227389695939054,
+ "eval_num_tokens": 4401934.0,
+ "eval_runtime": 89.4661,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 1900
+ },
+ {
+ "entropy": 0.4356316838413477,
+ "epoch": 4.627640313820157,
+ "grad_norm": 1.3870011568069458,
+ "learning_rate": 0.0002372530330023795,
+ "loss": 0.41338543891906737,
+ "mean_token_accuracy": 0.8642870113253593,
+ "num_tokens": 4449745.0,
+ "step": 1920
+ },
+ {
+ "epoch": 4.627640313820157,
+ "eval_entropy": 0.4954608427674583,
+ "eval_loss": 0.6783625483512878,
+ "eval_mean_token_accuracy": 0.8214637864841504,
+ "eval_num_tokens": 4449745.0,
+ "eval_runtime": 89.2291,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 1920
+ },
+ {
+ "entropy": 0.4468349639326334,
+ "epoch": 4.675920337960169,
+ "grad_norm": 0.8028655648231506,
+ "learning_rate": 0.00023431991948050538,
+ "loss": 0.4324653625488281,
+ "mean_token_accuracy": 0.8607208795845509,
+ "num_tokens": 4495074.0,
+ "step": 1940
+ },
+ {
+ "epoch": 4.675920337960169,
+ "eval_entropy": 0.5259480836351266,
+ "eval_loss": 0.6809864044189453,
+ "eval_mean_token_accuracy": 0.8200277019752545,
+ "eval_num_tokens": 4495074.0,
+ "eval_runtime": 89.8247,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 1940
+ },
+ {
+ "entropy": 0.4610355503857136,
+ "epoch": 4.724200362100181,
+ "grad_norm": 0.8928298950195312,
+ "learning_rate": 0.0002313721022353953,
+ "loss": 0.4314274787902832,
+ "mean_token_accuracy": 0.8559959702193737,
+ "num_tokens": 4541053.0,
+ "step": 1960
+ },
+ {
+ "epoch": 4.724200362100181,
+ "eval_entropy": 0.501018373651451,
+ "eval_loss": 0.6834176182746887,
+ "eval_mean_token_accuracy": 0.8220113566082515,
+ "eval_num_tokens": 4541053.0,
+ "eval_runtime": 90.4088,
+ "eval_samples_per_second": 15.706,
+ "eval_steps_per_second": 1.969,
+ "step": 1960
+ },
+ {
+ "entropy": 0.4568953149020672,
+ "epoch": 4.772480386240193,
+ "grad_norm": 1.018485188484192,
+ "learning_rate": 0.00022841041546446032,
+ "loss": 0.4306424617767334,
+ "mean_token_accuracy": 0.8598173558712006,
+ "num_tokens": 4584210.0,
+ "step": 1980
+ },
+ {
+ "epoch": 4.772480386240193,
+ "eval_entropy": 0.47633055371514865,
+ "eval_loss": 0.6886085867881775,
+ "eval_mean_token_accuracy": 0.823554899585381,
+ "eval_num_tokens": 4584210.0,
+ "eval_runtime": 89.8185,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 1980
+ },
+ {
+ "entropy": 0.4718513362109661,
+ "epoch": 4.820760410380205,
+ "grad_norm": 0.7674988508224487,
+ "learning_rate": 0.00022543569729002318,
+ "loss": 0.44841318130493163,
+ "mean_token_accuracy": 0.8566767282783985,
+ "num_tokens": 4630254.0,
+ "step": 2000
+ },
+ {
+ "epoch": 4.820760410380205,
+ "eval_entropy": 0.49708309284086977,
+ "eval_loss": 0.6878687739372253,
+ "eval_mean_token_accuracy": 0.8217844698536262,
+ "eval_num_tokens": 4630254.0,
+ "eval_runtime": 89.7497,
+ "eval_samples_per_second": 15.822,
+ "eval_steps_per_second": 1.983,
+ "step": 2000
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.0020307178351206e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2020/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2020/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2020/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2020/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2020/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2020/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2020/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2020/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2020/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2020/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2020/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2020/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2020/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2020/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..1ceb0a86764192e0427c9d11f9b02d3384718531
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2020/trainer_state.json
@@ -0,0 +1,2155 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 4.869040434520217,
+ "eval_steps": 20,
+ "global_step": 2020,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ },
+ {
+ "entropy": 0.4528769593685865,
+ "epoch": 4.579360289680145,
+ "grad_norm": 0.7440020442008972,
+ "learning_rate": 0.0002401706127645863,
+ "loss": 0.43678741455078124,
+ "mean_token_accuracy": 0.8619599856436253,
+ "num_tokens": 4401934.0,
+ "step": 1900
+ },
+ {
+ "epoch": 4.579360289680145,
+ "eval_entropy": 0.5085621437664782,
+ "eval_loss": 0.6766273379325867,
+ "eval_mean_token_accuracy": 0.8227389695939054,
+ "eval_num_tokens": 4401934.0,
+ "eval_runtime": 89.4661,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 1900
+ },
+ {
+ "entropy": 0.4356316838413477,
+ "epoch": 4.627640313820157,
+ "grad_norm": 1.3870011568069458,
+ "learning_rate": 0.0002372530330023795,
+ "loss": 0.41338543891906737,
+ "mean_token_accuracy": 0.8642870113253593,
+ "num_tokens": 4449745.0,
+ "step": 1920
+ },
+ {
+ "epoch": 4.627640313820157,
+ "eval_entropy": 0.4954608427674583,
+ "eval_loss": 0.6783625483512878,
+ "eval_mean_token_accuracy": 0.8214637864841504,
+ "eval_num_tokens": 4449745.0,
+ "eval_runtime": 89.2291,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 1920
+ },
+ {
+ "entropy": 0.4468349639326334,
+ "epoch": 4.675920337960169,
+ "grad_norm": 0.8028655648231506,
+ "learning_rate": 0.00023431991948050538,
+ "loss": 0.4324653625488281,
+ "mean_token_accuracy": 0.8607208795845509,
+ "num_tokens": 4495074.0,
+ "step": 1940
+ },
+ {
+ "epoch": 4.675920337960169,
+ "eval_entropy": 0.5259480836351266,
+ "eval_loss": 0.6809864044189453,
+ "eval_mean_token_accuracy": 0.8200277019752545,
+ "eval_num_tokens": 4495074.0,
+ "eval_runtime": 89.8247,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 1940
+ },
+ {
+ "entropy": 0.4610355503857136,
+ "epoch": 4.724200362100181,
+ "grad_norm": 0.8928298950195312,
+ "learning_rate": 0.0002313721022353953,
+ "loss": 0.4314274787902832,
+ "mean_token_accuracy": 0.8559959702193737,
+ "num_tokens": 4541053.0,
+ "step": 1960
+ },
+ {
+ "epoch": 4.724200362100181,
+ "eval_entropy": 0.501018373651451,
+ "eval_loss": 0.6834176182746887,
+ "eval_mean_token_accuracy": 0.8220113566082515,
+ "eval_num_tokens": 4541053.0,
+ "eval_runtime": 90.4088,
+ "eval_samples_per_second": 15.706,
+ "eval_steps_per_second": 1.969,
+ "step": 1960
+ },
+ {
+ "entropy": 0.4568953149020672,
+ "epoch": 4.772480386240193,
+ "grad_norm": 1.018485188484192,
+ "learning_rate": 0.00022841041546446032,
+ "loss": 0.4306424617767334,
+ "mean_token_accuracy": 0.8598173558712006,
+ "num_tokens": 4584210.0,
+ "step": 1980
+ },
+ {
+ "epoch": 4.772480386240193,
+ "eval_entropy": 0.47633055371514865,
+ "eval_loss": 0.6886085867881775,
+ "eval_mean_token_accuracy": 0.823554899585381,
+ "eval_num_tokens": 4584210.0,
+ "eval_runtime": 89.8185,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 1980
+ },
+ {
+ "entropy": 0.4718513362109661,
+ "epoch": 4.820760410380205,
+ "grad_norm": 0.7674988508224487,
+ "learning_rate": 0.00022543569729002318,
+ "loss": 0.44841318130493163,
+ "mean_token_accuracy": 0.8566767282783985,
+ "num_tokens": 4630254.0,
+ "step": 2000
+ },
+ {
+ "epoch": 4.820760410380205,
+ "eval_entropy": 0.49708309284086977,
+ "eval_loss": 0.6878687739372253,
+ "eval_mean_token_accuracy": 0.8217844698536262,
+ "eval_num_tokens": 4630254.0,
+ "eval_runtime": 89.7497,
+ "eval_samples_per_second": 15.822,
+ "eval_steps_per_second": 1.983,
+ "step": 2000
+ },
+ {
+ "entropy": 0.44017268233001233,
+ "epoch": 4.869040434520217,
+ "grad_norm": 0.6380876898765564,
+ "learning_rate": 0.00022244878952213976,
+ "loss": 0.42589750289916994,
+ "mean_token_accuracy": 0.8615814067423344,
+ "num_tokens": 4679265.0,
+ "step": 2020
+ },
+ {
+ "epoch": 4.869040434520217,
+ "eval_entropy": 0.5042208635740066,
+ "eval_loss": 0.6718228459358215,
+ "eval_mean_token_accuracy": 0.8232053056191863,
+ "eval_num_tokens": 4679265.0,
+ "eval_runtime": 89.7331,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 2020
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.023522400601477e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2040/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2040/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2040/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2040/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2040/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2040/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2040/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2040/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2040/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2040/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2040/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2040/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2040/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2040/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..011cbad1f2ac6288e743e21ee5614f71879d2494
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2040/trainer_state.json
@@ -0,0 +1,2176 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 4.917320458660229,
+ "eval_steps": 20,
+ "global_step": 2040,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ },
+ {
+ "entropy": 0.4528769593685865,
+ "epoch": 4.579360289680145,
+ "grad_norm": 0.7440020442008972,
+ "learning_rate": 0.0002401706127645863,
+ "loss": 0.43678741455078124,
+ "mean_token_accuracy": 0.8619599856436253,
+ "num_tokens": 4401934.0,
+ "step": 1900
+ },
+ {
+ "epoch": 4.579360289680145,
+ "eval_entropy": 0.5085621437664782,
+ "eval_loss": 0.6766273379325867,
+ "eval_mean_token_accuracy": 0.8227389695939054,
+ "eval_num_tokens": 4401934.0,
+ "eval_runtime": 89.4661,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 1900
+ },
+ {
+ "entropy": 0.4356316838413477,
+ "epoch": 4.627640313820157,
+ "grad_norm": 1.3870011568069458,
+ "learning_rate": 0.0002372530330023795,
+ "loss": 0.41338543891906737,
+ "mean_token_accuracy": 0.8642870113253593,
+ "num_tokens": 4449745.0,
+ "step": 1920
+ },
+ {
+ "epoch": 4.627640313820157,
+ "eval_entropy": 0.4954608427674583,
+ "eval_loss": 0.6783625483512878,
+ "eval_mean_token_accuracy": 0.8214637864841504,
+ "eval_num_tokens": 4449745.0,
+ "eval_runtime": 89.2291,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 1920
+ },
+ {
+ "entropy": 0.4468349639326334,
+ "epoch": 4.675920337960169,
+ "grad_norm": 0.8028655648231506,
+ "learning_rate": 0.00023431991948050538,
+ "loss": 0.4324653625488281,
+ "mean_token_accuracy": 0.8607208795845509,
+ "num_tokens": 4495074.0,
+ "step": 1940
+ },
+ {
+ "epoch": 4.675920337960169,
+ "eval_entropy": 0.5259480836351266,
+ "eval_loss": 0.6809864044189453,
+ "eval_mean_token_accuracy": 0.8200277019752545,
+ "eval_num_tokens": 4495074.0,
+ "eval_runtime": 89.8247,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 1940
+ },
+ {
+ "entropy": 0.4610355503857136,
+ "epoch": 4.724200362100181,
+ "grad_norm": 0.8928298950195312,
+ "learning_rate": 0.0002313721022353953,
+ "loss": 0.4314274787902832,
+ "mean_token_accuracy": 0.8559959702193737,
+ "num_tokens": 4541053.0,
+ "step": 1960
+ },
+ {
+ "epoch": 4.724200362100181,
+ "eval_entropy": 0.501018373651451,
+ "eval_loss": 0.6834176182746887,
+ "eval_mean_token_accuracy": 0.8220113566082515,
+ "eval_num_tokens": 4541053.0,
+ "eval_runtime": 90.4088,
+ "eval_samples_per_second": 15.706,
+ "eval_steps_per_second": 1.969,
+ "step": 1960
+ },
+ {
+ "entropy": 0.4568953149020672,
+ "epoch": 4.772480386240193,
+ "grad_norm": 1.018485188484192,
+ "learning_rate": 0.00022841041546446032,
+ "loss": 0.4306424617767334,
+ "mean_token_accuracy": 0.8598173558712006,
+ "num_tokens": 4584210.0,
+ "step": 1980
+ },
+ {
+ "epoch": 4.772480386240193,
+ "eval_entropy": 0.47633055371514865,
+ "eval_loss": 0.6886085867881775,
+ "eval_mean_token_accuracy": 0.823554899585381,
+ "eval_num_tokens": 4584210.0,
+ "eval_runtime": 89.8185,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 1980
+ },
+ {
+ "entropy": 0.4718513362109661,
+ "epoch": 4.820760410380205,
+ "grad_norm": 0.7674988508224487,
+ "learning_rate": 0.00022543569729002318,
+ "loss": 0.44841318130493163,
+ "mean_token_accuracy": 0.8566767282783985,
+ "num_tokens": 4630254.0,
+ "step": 2000
+ },
+ {
+ "epoch": 4.820760410380205,
+ "eval_entropy": 0.49708309284086977,
+ "eval_loss": 0.6878687739372253,
+ "eval_mean_token_accuracy": 0.8217844698536262,
+ "eval_num_tokens": 4630254.0,
+ "eval_runtime": 89.7497,
+ "eval_samples_per_second": 15.822,
+ "eval_steps_per_second": 1.983,
+ "step": 2000
+ },
+ {
+ "entropy": 0.44017268233001233,
+ "epoch": 4.869040434520217,
+ "grad_norm": 0.6380876898765564,
+ "learning_rate": 0.00022244878952213976,
+ "loss": 0.42589750289916994,
+ "mean_token_accuracy": 0.8615814067423344,
+ "num_tokens": 4679265.0,
+ "step": 2020
+ },
+ {
+ "epoch": 4.869040434520217,
+ "eval_entropy": 0.5042208635740066,
+ "eval_loss": 0.6718228459358215,
+ "eval_mean_token_accuracy": 0.8232053056191863,
+ "eval_num_tokens": 4679265.0,
+ "eval_runtime": 89.7331,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 2020
+ },
+ {
+ "entropy": 0.4654298175126314,
+ "epoch": 4.917320458660229,
+ "grad_norm": 0.8285323977470398,
+ "learning_rate": 0.0002194505374203766,
+ "loss": 0.44148526191711424,
+ "mean_token_accuracy": 0.8545936144888401,
+ "num_tokens": 4723592.0,
+ "step": 2040
+ },
+ {
+ "epoch": 4.917320458660229,
+ "eval_entropy": 0.4957777807551823,
+ "eval_loss": 0.6694385409355164,
+ "eval_mean_token_accuracy": 0.8243840380331103,
+ "eval_num_tokens": 4723592.0,
+ "eval_runtime": 89.5293,
+ "eval_samples_per_second": 15.861,
+ "eval_steps_per_second": 1.988,
+ "step": 2040
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.042472133585244e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2060/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2060/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2060/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2060/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2060/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2060/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2060/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2060/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2060/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2060/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2060/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2060/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2060/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2060/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..b382720b33fded35c45d9fc378acafe2293a1cbe
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2060/trainer_state.json
@@ -0,0 +1,2197 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 4.965600482800241,
+ "eval_steps": 20,
+ "global_step": 2060,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ },
+ {
+ "entropy": 0.4528769593685865,
+ "epoch": 4.579360289680145,
+ "grad_norm": 0.7440020442008972,
+ "learning_rate": 0.0002401706127645863,
+ "loss": 0.43678741455078124,
+ "mean_token_accuracy": 0.8619599856436253,
+ "num_tokens": 4401934.0,
+ "step": 1900
+ },
+ {
+ "epoch": 4.579360289680145,
+ "eval_entropy": 0.5085621437664782,
+ "eval_loss": 0.6766273379325867,
+ "eval_mean_token_accuracy": 0.8227389695939054,
+ "eval_num_tokens": 4401934.0,
+ "eval_runtime": 89.4661,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 1900
+ },
+ {
+ "entropy": 0.4356316838413477,
+ "epoch": 4.627640313820157,
+ "grad_norm": 1.3870011568069458,
+ "learning_rate": 0.0002372530330023795,
+ "loss": 0.41338543891906737,
+ "mean_token_accuracy": 0.8642870113253593,
+ "num_tokens": 4449745.0,
+ "step": 1920
+ },
+ {
+ "epoch": 4.627640313820157,
+ "eval_entropy": 0.4954608427674583,
+ "eval_loss": 0.6783625483512878,
+ "eval_mean_token_accuracy": 0.8214637864841504,
+ "eval_num_tokens": 4449745.0,
+ "eval_runtime": 89.2291,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 1920
+ },
+ {
+ "entropy": 0.4468349639326334,
+ "epoch": 4.675920337960169,
+ "grad_norm": 0.8028655648231506,
+ "learning_rate": 0.00023431991948050538,
+ "loss": 0.4324653625488281,
+ "mean_token_accuracy": 0.8607208795845509,
+ "num_tokens": 4495074.0,
+ "step": 1940
+ },
+ {
+ "epoch": 4.675920337960169,
+ "eval_entropy": 0.5259480836351266,
+ "eval_loss": 0.6809864044189453,
+ "eval_mean_token_accuracy": 0.8200277019752545,
+ "eval_num_tokens": 4495074.0,
+ "eval_runtime": 89.8247,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 1940
+ },
+ {
+ "entropy": 0.4610355503857136,
+ "epoch": 4.724200362100181,
+ "grad_norm": 0.8928298950195312,
+ "learning_rate": 0.0002313721022353953,
+ "loss": 0.4314274787902832,
+ "mean_token_accuracy": 0.8559959702193737,
+ "num_tokens": 4541053.0,
+ "step": 1960
+ },
+ {
+ "epoch": 4.724200362100181,
+ "eval_entropy": 0.501018373651451,
+ "eval_loss": 0.6834176182746887,
+ "eval_mean_token_accuracy": 0.8220113566082515,
+ "eval_num_tokens": 4541053.0,
+ "eval_runtime": 90.4088,
+ "eval_samples_per_second": 15.706,
+ "eval_steps_per_second": 1.969,
+ "step": 1960
+ },
+ {
+ "entropy": 0.4568953149020672,
+ "epoch": 4.772480386240193,
+ "grad_norm": 1.018485188484192,
+ "learning_rate": 0.00022841041546446032,
+ "loss": 0.4306424617767334,
+ "mean_token_accuracy": 0.8598173558712006,
+ "num_tokens": 4584210.0,
+ "step": 1980
+ },
+ {
+ "epoch": 4.772480386240193,
+ "eval_entropy": 0.47633055371514865,
+ "eval_loss": 0.6886085867881775,
+ "eval_mean_token_accuracy": 0.823554899585381,
+ "eval_num_tokens": 4584210.0,
+ "eval_runtime": 89.8185,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 1980
+ },
+ {
+ "entropy": 0.4718513362109661,
+ "epoch": 4.820760410380205,
+ "grad_norm": 0.7674988508224487,
+ "learning_rate": 0.00022543569729002318,
+ "loss": 0.44841318130493163,
+ "mean_token_accuracy": 0.8566767282783985,
+ "num_tokens": 4630254.0,
+ "step": 2000
+ },
+ {
+ "epoch": 4.820760410380205,
+ "eval_entropy": 0.49708309284086977,
+ "eval_loss": 0.6878687739372253,
+ "eval_mean_token_accuracy": 0.8217844698536262,
+ "eval_num_tokens": 4630254.0,
+ "eval_runtime": 89.7497,
+ "eval_samples_per_second": 15.822,
+ "eval_steps_per_second": 1.983,
+ "step": 2000
+ },
+ {
+ "entropy": 0.44017268233001233,
+ "epoch": 4.869040434520217,
+ "grad_norm": 0.6380876898765564,
+ "learning_rate": 0.00022244878952213976,
+ "loss": 0.42589750289916994,
+ "mean_token_accuracy": 0.8615814067423344,
+ "num_tokens": 4679265.0,
+ "step": 2020
+ },
+ {
+ "epoch": 4.869040434520217,
+ "eval_entropy": 0.5042208635740066,
+ "eval_loss": 0.6718228459358215,
+ "eval_mean_token_accuracy": 0.8232053056191863,
+ "eval_num_tokens": 4679265.0,
+ "eval_runtime": 89.7331,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 2020
+ },
+ {
+ "entropy": 0.4654298175126314,
+ "epoch": 4.917320458660229,
+ "grad_norm": 0.8285323977470398,
+ "learning_rate": 0.0002194505374203766,
+ "loss": 0.44148526191711424,
+ "mean_token_accuracy": 0.8545936144888401,
+ "num_tokens": 4723592.0,
+ "step": 2040
+ },
+ {
+ "epoch": 4.917320458660229,
+ "eval_entropy": 0.4957777807551823,
+ "eval_loss": 0.6694385409355164,
+ "eval_mean_token_accuracy": 0.8243840380331103,
+ "eval_num_tokens": 4723592.0,
+ "eval_runtime": 89.5293,
+ "eval_samples_per_second": 15.861,
+ "eval_steps_per_second": 1.988,
+ "step": 2040
+ },
+ {
+ "entropy": 0.46651641055941584,
+ "epoch": 4.965600482800241,
+ "grad_norm": 0.946260929107666,
+ "learning_rate": 0.00021644178945461268,
+ "loss": 0.43698992729187014,
+ "mean_token_accuracy": 0.8586322009563446,
+ "num_tokens": 4767253.0,
+ "step": 2060
+ },
+ {
+ "epoch": 4.965600482800241,
+ "eval_entropy": 0.5063281496254246,
+ "eval_loss": 0.6729032397270203,
+ "eval_mean_token_accuracy": 0.824299869912394,
+ "eval_num_tokens": 4767253.0,
+ "eval_runtime": 89.6207,
+ "eval_samples_per_second": 15.845,
+ "eval_steps_per_second": 1.986,
+ "step": 2060
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.060940763079086e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2080/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2080/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2080/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2080/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2080/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2080/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2080/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2080/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2080/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2080/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2080/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2080/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2080/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2080/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..f828798ca4f6eb545c0b8ddf1c0acbfbf5e8f607
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2080/trainer_state.json
@@ -0,0 +1,2218 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 5.012070006035003,
+ "eval_steps": 20,
+ "global_step": 2080,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ },
+ {
+ "entropy": 0.4528769593685865,
+ "epoch": 4.579360289680145,
+ "grad_norm": 0.7440020442008972,
+ "learning_rate": 0.0002401706127645863,
+ "loss": 0.43678741455078124,
+ "mean_token_accuracy": 0.8619599856436253,
+ "num_tokens": 4401934.0,
+ "step": 1900
+ },
+ {
+ "epoch": 4.579360289680145,
+ "eval_entropy": 0.5085621437664782,
+ "eval_loss": 0.6766273379325867,
+ "eval_mean_token_accuracy": 0.8227389695939054,
+ "eval_num_tokens": 4401934.0,
+ "eval_runtime": 89.4661,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 1900
+ },
+ {
+ "entropy": 0.4356316838413477,
+ "epoch": 4.627640313820157,
+ "grad_norm": 1.3870011568069458,
+ "learning_rate": 0.0002372530330023795,
+ "loss": 0.41338543891906737,
+ "mean_token_accuracy": 0.8642870113253593,
+ "num_tokens": 4449745.0,
+ "step": 1920
+ },
+ {
+ "epoch": 4.627640313820157,
+ "eval_entropy": 0.4954608427674583,
+ "eval_loss": 0.6783625483512878,
+ "eval_mean_token_accuracy": 0.8214637864841504,
+ "eval_num_tokens": 4449745.0,
+ "eval_runtime": 89.2291,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 1920
+ },
+ {
+ "entropy": 0.4468349639326334,
+ "epoch": 4.675920337960169,
+ "grad_norm": 0.8028655648231506,
+ "learning_rate": 0.00023431991948050538,
+ "loss": 0.4324653625488281,
+ "mean_token_accuracy": 0.8607208795845509,
+ "num_tokens": 4495074.0,
+ "step": 1940
+ },
+ {
+ "epoch": 4.675920337960169,
+ "eval_entropy": 0.5259480836351266,
+ "eval_loss": 0.6809864044189453,
+ "eval_mean_token_accuracy": 0.8200277019752545,
+ "eval_num_tokens": 4495074.0,
+ "eval_runtime": 89.8247,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 1940
+ },
+ {
+ "entropy": 0.4610355503857136,
+ "epoch": 4.724200362100181,
+ "grad_norm": 0.8928298950195312,
+ "learning_rate": 0.0002313721022353953,
+ "loss": 0.4314274787902832,
+ "mean_token_accuracy": 0.8559959702193737,
+ "num_tokens": 4541053.0,
+ "step": 1960
+ },
+ {
+ "epoch": 4.724200362100181,
+ "eval_entropy": 0.501018373651451,
+ "eval_loss": 0.6834176182746887,
+ "eval_mean_token_accuracy": 0.8220113566082515,
+ "eval_num_tokens": 4541053.0,
+ "eval_runtime": 90.4088,
+ "eval_samples_per_second": 15.706,
+ "eval_steps_per_second": 1.969,
+ "step": 1960
+ },
+ {
+ "entropy": 0.4568953149020672,
+ "epoch": 4.772480386240193,
+ "grad_norm": 1.018485188484192,
+ "learning_rate": 0.00022841041546446032,
+ "loss": 0.4306424617767334,
+ "mean_token_accuracy": 0.8598173558712006,
+ "num_tokens": 4584210.0,
+ "step": 1980
+ },
+ {
+ "epoch": 4.772480386240193,
+ "eval_entropy": 0.47633055371514865,
+ "eval_loss": 0.6886085867881775,
+ "eval_mean_token_accuracy": 0.823554899585381,
+ "eval_num_tokens": 4584210.0,
+ "eval_runtime": 89.8185,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 1980
+ },
+ {
+ "entropy": 0.4718513362109661,
+ "epoch": 4.820760410380205,
+ "grad_norm": 0.7674988508224487,
+ "learning_rate": 0.00022543569729002318,
+ "loss": 0.44841318130493163,
+ "mean_token_accuracy": 0.8566767282783985,
+ "num_tokens": 4630254.0,
+ "step": 2000
+ },
+ {
+ "epoch": 4.820760410380205,
+ "eval_entropy": 0.49708309284086977,
+ "eval_loss": 0.6878687739372253,
+ "eval_mean_token_accuracy": 0.8217844698536262,
+ "eval_num_tokens": 4630254.0,
+ "eval_runtime": 89.7497,
+ "eval_samples_per_second": 15.822,
+ "eval_steps_per_second": 1.983,
+ "step": 2000
+ },
+ {
+ "entropy": 0.44017268233001233,
+ "epoch": 4.869040434520217,
+ "grad_norm": 0.6380876898765564,
+ "learning_rate": 0.00022244878952213976,
+ "loss": 0.42589750289916994,
+ "mean_token_accuracy": 0.8615814067423344,
+ "num_tokens": 4679265.0,
+ "step": 2020
+ },
+ {
+ "epoch": 4.869040434520217,
+ "eval_entropy": 0.5042208635740066,
+ "eval_loss": 0.6718228459358215,
+ "eval_mean_token_accuracy": 0.8232053056191863,
+ "eval_num_tokens": 4679265.0,
+ "eval_runtime": 89.7331,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 2020
+ },
+ {
+ "entropy": 0.4654298175126314,
+ "epoch": 4.917320458660229,
+ "grad_norm": 0.8285323977470398,
+ "learning_rate": 0.0002194505374203766,
+ "loss": 0.44148526191711424,
+ "mean_token_accuracy": 0.8545936144888401,
+ "num_tokens": 4723592.0,
+ "step": 2040
+ },
+ {
+ "epoch": 4.917320458660229,
+ "eval_entropy": 0.4957777807551823,
+ "eval_loss": 0.6694385409355164,
+ "eval_mean_token_accuracy": 0.8243840380331103,
+ "eval_num_tokens": 4723592.0,
+ "eval_runtime": 89.5293,
+ "eval_samples_per_second": 15.861,
+ "eval_steps_per_second": 1.988,
+ "step": 2040
+ },
+ {
+ "entropy": 0.46651641055941584,
+ "epoch": 4.965600482800241,
+ "grad_norm": 0.946260929107666,
+ "learning_rate": 0.00021644178945461268,
+ "loss": 0.43698992729187014,
+ "mean_token_accuracy": 0.8586322009563446,
+ "num_tokens": 4767253.0,
+ "step": 2060
+ },
+ {
+ "epoch": 4.965600482800241,
+ "eval_entropy": 0.5063281496254246,
+ "eval_loss": 0.6729032397270203,
+ "eval_mean_token_accuracy": 0.824299869912394,
+ "eval_num_tokens": 4767253.0,
+ "eval_runtime": 89.6207,
+ "eval_samples_per_second": 15.845,
+ "eval_steps_per_second": 1.986,
+ "step": 2060
+ },
+ {
+ "entropy": 0.4135968596130222,
+ "epoch": 5.012070006035003,
+ "grad_norm": 0.6853868961334229,
+ "learning_rate": 0.0002134233970649322,
+ "loss": 0.3800280332565308,
+ "mean_token_accuracy": 0.8756770677380747,
+ "num_tokens": 4816587.0,
+ "step": 2080
+ },
+ {
+ "epoch": 5.012070006035003,
+ "eval_entropy": 0.4107752423942759,
+ "eval_loss": 0.7204955220222473,
+ "eval_mean_token_accuracy": 0.8226918417416261,
+ "eval_num_tokens": 4816587.0,
+ "eval_runtime": 89.5817,
+ "eval_samples_per_second": 15.851,
+ "eval_steps_per_second": 1.987,
+ "step": 2080
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.0814343356200755e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2100/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2100/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2100/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2100/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2100/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2100/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2100/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2100/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2100/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2100/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2100/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2100/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2100/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2100/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..d40a8f4af1929ce838d83cc407423592cdab5b49
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2100/trainer_state.json
@@ -0,0 +1,2239 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 5.060350030175015,
+ "eval_steps": 20,
+ "global_step": 2100,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ },
+ {
+ "entropy": 0.4528769593685865,
+ "epoch": 4.579360289680145,
+ "grad_norm": 0.7440020442008972,
+ "learning_rate": 0.0002401706127645863,
+ "loss": 0.43678741455078124,
+ "mean_token_accuracy": 0.8619599856436253,
+ "num_tokens": 4401934.0,
+ "step": 1900
+ },
+ {
+ "epoch": 4.579360289680145,
+ "eval_entropy": 0.5085621437664782,
+ "eval_loss": 0.6766273379325867,
+ "eval_mean_token_accuracy": 0.8227389695939054,
+ "eval_num_tokens": 4401934.0,
+ "eval_runtime": 89.4661,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 1900
+ },
+ {
+ "entropy": 0.4356316838413477,
+ "epoch": 4.627640313820157,
+ "grad_norm": 1.3870011568069458,
+ "learning_rate": 0.0002372530330023795,
+ "loss": 0.41338543891906737,
+ "mean_token_accuracy": 0.8642870113253593,
+ "num_tokens": 4449745.0,
+ "step": 1920
+ },
+ {
+ "epoch": 4.627640313820157,
+ "eval_entropy": 0.4954608427674583,
+ "eval_loss": 0.6783625483512878,
+ "eval_mean_token_accuracy": 0.8214637864841504,
+ "eval_num_tokens": 4449745.0,
+ "eval_runtime": 89.2291,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 1920
+ },
+ {
+ "entropy": 0.4468349639326334,
+ "epoch": 4.675920337960169,
+ "grad_norm": 0.8028655648231506,
+ "learning_rate": 0.00023431991948050538,
+ "loss": 0.4324653625488281,
+ "mean_token_accuracy": 0.8607208795845509,
+ "num_tokens": 4495074.0,
+ "step": 1940
+ },
+ {
+ "epoch": 4.675920337960169,
+ "eval_entropy": 0.5259480836351266,
+ "eval_loss": 0.6809864044189453,
+ "eval_mean_token_accuracy": 0.8200277019752545,
+ "eval_num_tokens": 4495074.0,
+ "eval_runtime": 89.8247,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 1940
+ },
+ {
+ "entropy": 0.4610355503857136,
+ "epoch": 4.724200362100181,
+ "grad_norm": 0.8928298950195312,
+ "learning_rate": 0.0002313721022353953,
+ "loss": 0.4314274787902832,
+ "mean_token_accuracy": 0.8559959702193737,
+ "num_tokens": 4541053.0,
+ "step": 1960
+ },
+ {
+ "epoch": 4.724200362100181,
+ "eval_entropy": 0.501018373651451,
+ "eval_loss": 0.6834176182746887,
+ "eval_mean_token_accuracy": 0.8220113566082515,
+ "eval_num_tokens": 4541053.0,
+ "eval_runtime": 90.4088,
+ "eval_samples_per_second": 15.706,
+ "eval_steps_per_second": 1.969,
+ "step": 1960
+ },
+ {
+ "entropy": 0.4568953149020672,
+ "epoch": 4.772480386240193,
+ "grad_norm": 1.018485188484192,
+ "learning_rate": 0.00022841041546446032,
+ "loss": 0.4306424617767334,
+ "mean_token_accuracy": 0.8598173558712006,
+ "num_tokens": 4584210.0,
+ "step": 1980
+ },
+ {
+ "epoch": 4.772480386240193,
+ "eval_entropy": 0.47633055371514865,
+ "eval_loss": 0.6886085867881775,
+ "eval_mean_token_accuracy": 0.823554899585381,
+ "eval_num_tokens": 4584210.0,
+ "eval_runtime": 89.8185,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 1980
+ },
+ {
+ "entropy": 0.4718513362109661,
+ "epoch": 4.820760410380205,
+ "grad_norm": 0.7674988508224487,
+ "learning_rate": 0.00022543569729002318,
+ "loss": 0.44841318130493163,
+ "mean_token_accuracy": 0.8566767282783985,
+ "num_tokens": 4630254.0,
+ "step": 2000
+ },
+ {
+ "epoch": 4.820760410380205,
+ "eval_entropy": 0.49708309284086977,
+ "eval_loss": 0.6878687739372253,
+ "eval_mean_token_accuracy": 0.8217844698536262,
+ "eval_num_tokens": 4630254.0,
+ "eval_runtime": 89.7497,
+ "eval_samples_per_second": 15.822,
+ "eval_steps_per_second": 1.983,
+ "step": 2000
+ },
+ {
+ "entropy": 0.44017268233001233,
+ "epoch": 4.869040434520217,
+ "grad_norm": 0.6380876898765564,
+ "learning_rate": 0.00022244878952213976,
+ "loss": 0.42589750289916994,
+ "mean_token_accuracy": 0.8615814067423344,
+ "num_tokens": 4679265.0,
+ "step": 2020
+ },
+ {
+ "epoch": 4.869040434520217,
+ "eval_entropy": 0.5042208635740066,
+ "eval_loss": 0.6718228459358215,
+ "eval_mean_token_accuracy": 0.8232053056191863,
+ "eval_num_tokens": 4679265.0,
+ "eval_runtime": 89.7331,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 2020
+ },
+ {
+ "entropy": 0.4654298175126314,
+ "epoch": 4.917320458660229,
+ "grad_norm": 0.8285323977470398,
+ "learning_rate": 0.0002194505374203766,
+ "loss": 0.44148526191711424,
+ "mean_token_accuracy": 0.8545936144888401,
+ "num_tokens": 4723592.0,
+ "step": 2040
+ },
+ {
+ "epoch": 4.917320458660229,
+ "eval_entropy": 0.4957777807551823,
+ "eval_loss": 0.6694385409355164,
+ "eval_mean_token_accuracy": 0.8243840380331103,
+ "eval_num_tokens": 4723592.0,
+ "eval_runtime": 89.5293,
+ "eval_samples_per_second": 15.861,
+ "eval_steps_per_second": 1.988,
+ "step": 2040
+ },
+ {
+ "entropy": 0.46651641055941584,
+ "epoch": 4.965600482800241,
+ "grad_norm": 0.946260929107666,
+ "learning_rate": 0.00021644178945461268,
+ "loss": 0.43698992729187014,
+ "mean_token_accuracy": 0.8586322009563446,
+ "num_tokens": 4767253.0,
+ "step": 2060
+ },
+ {
+ "epoch": 4.965600482800241,
+ "eval_entropy": 0.5063281496254246,
+ "eval_loss": 0.6729032397270203,
+ "eval_mean_token_accuracy": 0.824299869912394,
+ "eval_num_tokens": 4767253.0,
+ "eval_runtime": 89.6207,
+ "eval_samples_per_second": 15.845,
+ "eval_steps_per_second": 1.986,
+ "step": 2060
+ },
+ {
+ "entropy": 0.4135968596130222,
+ "epoch": 5.012070006035003,
+ "grad_norm": 0.6853868961334229,
+ "learning_rate": 0.0002134233970649322,
+ "loss": 0.3800280332565308,
+ "mean_token_accuracy": 0.8756770677380747,
+ "num_tokens": 4816587.0,
+ "step": 2080
+ },
+ {
+ "epoch": 5.012070006035003,
+ "eval_entropy": 0.4107752423942759,
+ "eval_loss": 0.7204955220222473,
+ "eval_mean_token_accuracy": 0.8226918417416261,
+ "eval_num_tokens": 4816587.0,
+ "eval_runtime": 89.5817,
+ "eval_samples_per_second": 15.851,
+ "eval_steps_per_second": 1.987,
+ "step": 2080
+ },
+ {
+ "entropy": 0.316746598854661,
+ "epoch": 5.060350030175015,
+ "grad_norm": 0.6939145922660828,
+ "learning_rate": 0.00021039621442067732,
+ "loss": 0.2942554473876953,
+ "mean_token_accuracy": 0.90015934035182,
+ "num_tokens": 4861071.0,
+ "step": 2100
+ },
+ {
+ "epoch": 5.060350030175015,
+ "eval_entropy": 0.4214997919422857,
+ "eval_loss": 0.7268305420875549,
+ "eval_mean_token_accuracy": 0.8216965369294199,
+ "eval_num_tokens": 4861071.0,
+ "eval_runtime": 89.8077,
+ "eval_samples_per_second": 15.812,
+ "eval_steps_per_second": 1.982,
+ "step": 2100
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.1007556671949414e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2120/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2120/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2120/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2120/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2120/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2120/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2120/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2120/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2120/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2120/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2120/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2120/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2120/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2120/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..30dec68a7e7c4b051664080848b28ee5ff5a4f7d
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2120/trainer_state.json
@@ -0,0 +1,2260 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 5.108630054315027,
+ "eval_steps": 20,
+ "global_step": 2120,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ },
+ {
+ "entropy": 0.4528769593685865,
+ "epoch": 4.579360289680145,
+ "grad_norm": 0.7440020442008972,
+ "learning_rate": 0.0002401706127645863,
+ "loss": 0.43678741455078124,
+ "mean_token_accuracy": 0.8619599856436253,
+ "num_tokens": 4401934.0,
+ "step": 1900
+ },
+ {
+ "epoch": 4.579360289680145,
+ "eval_entropy": 0.5085621437664782,
+ "eval_loss": 0.6766273379325867,
+ "eval_mean_token_accuracy": 0.8227389695939054,
+ "eval_num_tokens": 4401934.0,
+ "eval_runtime": 89.4661,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 1900
+ },
+ {
+ "entropy": 0.4356316838413477,
+ "epoch": 4.627640313820157,
+ "grad_norm": 1.3870011568069458,
+ "learning_rate": 0.0002372530330023795,
+ "loss": 0.41338543891906737,
+ "mean_token_accuracy": 0.8642870113253593,
+ "num_tokens": 4449745.0,
+ "step": 1920
+ },
+ {
+ "epoch": 4.627640313820157,
+ "eval_entropy": 0.4954608427674583,
+ "eval_loss": 0.6783625483512878,
+ "eval_mean_token_accuracy": 0.8214637864841504,
+ "eval_num_tokens": 4449745.0,
+ "eval_runtime": 89.2291,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 1920
+ },
+ {
+ "entropy": 0.4468349639326334,
+ "epoch": 4.675920337960169,
+ "grad_norm": 0.8028655648231506,
+ "learning_rate": 0.00023431991948050538,
+ "loss": 0.4324653625488281,
+ "mean_token_accuracy": 0.8607208795845509,
+ "num_tokens": 4495074.0,
+ "step": 1940
+ },
+ {
+ "epoch": 4.675920337960169,
+ "eval_entropy": 0.5259480836351266,
+ "eval_loss": 0.6809864044189453,
+ "eval_mean_token_accuracy": 0.8200277019752545,
+ "eval_num_tokens": 4495074.0,
+ "eval_runtime": 89.8247,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 1940
+ },
+ {
+ "entropy": 0.4610355503857136,
+ "epoch": 4.724200362100181,
+ "grad_norm": 0.8928298950195312,
+ "learning_rate": 0.0002313721022353953,
+ "loss": 0.4314274787902832,
+ "mean_token_accuracy": 0.8559959702193737,
+ "num_tokens": 4541053.0,
+ "step": 1960
+ },
+ {
+ "epoch": 4.724200362100181,
+ "eval_entropy": 0.501018373651451,
+ "eval_loss": 0.6834176182746887,
+ "eval_mean_token_accuracy": 0.8220113566082515,
+ "eval_num_tokens": 4541053.0,
+ "eval_runtime": 90.4088,
+ "eval_samples_per_second": 15.706,
+ "eval_steps_per_second": 1.969,
+ "step": 1960
+ },
+ {
+ "entropy": 0.4568953149020672,
+ "epoch": 4.772480386240193,
+ "grad_norm": 1.018485188484192,
+ "learning_rate": 0.00022841041546446032,
+ "loss": 0.4306424617767334,
+ "mean_token_accuracy": 0.8598173558712006,
+ "num_tokens": 4584210.0,
+ "step": 1980
+ },
+ {
+ "epoch": 4.772480386240193,
+ "eval_entropy": 0.47633055371514865,
+ "eval_loss": 0.6886085867881775,
+ "eval_mean_token_accuracy": 0.823554899585381,
+ "eval_num_tokens": 4584210.0,
+ "eval_runtime": 89.8185,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 1980
+ },
+ {
+ "entropy": 0.4718513362109661,
+ "epoch": 4.820760410380205,
+ "grad_norm": 0.7674988508224487,
+ "learning_rate": 0.00022543569729002318,
+ "loss": 0.44841318130493163,
+ "mean_token_accuracy": 0.8566767282783985,
+ "num_tokens": 4630254.0,
+ "step": 2000
+ },
+ {
+ "epoch": 4.820760410380205,
+ "eval_entropy": 0.49708309284086977,
+ "eval_loss": 0.6878687739372253,
+ "eval_mean_token_accuracy": 0.8217844698536262,
+ "eval_num_tokens": 4630254.0,
+ "eval_runtime": 89.7497,
+ "eval_samples_per_second": 15.822,
+ "eval_steps_per_second": 1.983,
+ "step": 2000
+ },
+ {
+ "entropy": 0.44017268233001233,
+ "epoch": 4.869040434520217,
+ "grad_norm": 0.6380876898765564,
+ "learning_rate": 0.00022244878952213976,
+ "loss": 0.42589750289916994,
+ "mean_token_accuracy": 0.8615814067423344,
+ "num_tokens": 4679265.0,
+ "step": 2020
+ },
+ {
+ "epoch": 4.869040434520217,
+ "eval_entropy": 0.5042208635740066,
+ "eval_loss": 0.6718228459358215,
+ "eval_mean_token_accuracy": 0.8232053056191863,
+ "eval_num_tokens": 4679265.0,
+ "eval_runtime": 89.7331,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 2020
+ },
+ {
+ "entropy": 0.4654298175126314,
+ "epoch": 4.917320458660229,
+ "grad_norm": 0.8285323977470398,
+ "learning_rate": 0.0002194505374203766,
+ "loss": 0.44148526191711424,
+ "mean_token_accuracy": 0.8545936144888401,
+ "num_tokens": 4723592.0,
+ "step": 2040
+ },
+ {
+ "epoch": 4.917320458660229,
+ "eval_entropy": 0.4957777807551823,
+ "eval_loss": 0.6694385409355164,
+ "eval_mean_token_accuracy": 0.8243840380331103,
+ "eval_num_tokens": 4723592.0,
+ "eval_runtime": 89.5293,
+ "eval_samples_per_second": 15.861,
+ "eval_steps_per_second": 1.988,
+ "step": 2040
+ },
+ {
+ "entropy": 0.46651641055941584,
+ "epoch": 4.965600482800241,
+ "grad_norm": 0.946260929107666,
+ "learning_rate": 0.00021644178945461268,
+ "loss": 0.43698992729187014,
+ "mean_token_accuracy": 0.8586322009563446,
+ "num_tokens": 4767253.0,
+ "step": 2060
+ },
+ {
+ "epoch": 4.965600482800241,
+ "eval_entropy": 0.5063281496254246,
+ "eval_loss": 0.6729032397270203,
+ "eval_mean_token_accuracy": 0.824299869912394,
+ "eval_num_tokens": 4767253.0,
+ "eval_runtime": 89.6207,
+ "eval_samples_per_second": 15.845,
+ "eval_steps_per_second": 1.986,
+ "step": 2060
+ },
+ {
+ "entropy": 0.4135968596130222,
+ "epoch": 5.012070006035003,
+ "grad_norm": 0.6853868961334229,
+ "learning_rate": 0.0002134233970649322,
+ "loss": 0.3800280332565308,
+ "mean_token_accuracy": 0.8756770677380747,
+ "num_tokens": 4816587.0,
+ "step": 2080
+ },
+ {
+ "epoch": 5.012070006035003,
+ "eval_entropy": 0.4107752423942759,
+ "eval_loss": 0.7204955220222473,
+ "eval_mean_token_accuracy": 0.8226918417416261,
+ "eval_num_tokens": 4816587.0,
+ "eval_runtime": 89.5817,
+ "eval_samples_per_second": 15.851,
+ "eval_steps_per_second": 1.987,
+ "step": 2080
+ },
+ {
+ "entropy": 0.316746598854661,
+ "epoch": 5.060350030175015,
+ "grad_norm": 0.6939145922660828,
+ "learning_rate": 0.00021039621442067732,
+ "loss": 0.2942554473876953,
+ "mean_token_accuracy": 0.90015934035182,
+ "num_tokens": 4861071.0,
+ "step": 2100
+ },
+ {
+ "epoch": 5.060350030175015,
+ "eval_entropy": 0.4214997919422857,
+ "eval_loss": 0.7268305420875549,
+ "eval_mean_token_accuracy": 0.8216965369294199,
+ "eval_num_tokens": 4861071.0,
+ "eval_runtime": 89.8077,
+ "eval_samples_per_second": 15.812,
+ "eval_steps_per_second": 1.982,
+ "step": 2100
+ },
+ {
+ "entropy": 0.3187096672132611,
+ "epoch": 5.108630054315027,
+ "grad_norm": 0.9134758114814758,
+ "learning_rate": 0.00020736109817872823,
+ "loss": 0.2839747190475464,
+ "mean_token_accuracy": 0.9006688073277473,
+ "num_tokens": 4904742.0,
+ "step": 2120
+ },
+ {
+ "epoch": 5.108630054315027,
+ "eval_entropy": 0.4375539443800958,
+ "eval_loss": 0.7280111908912659,
+ "eval_mean_token_accuracy": 0.8186122694712007,
+ "eval_num_tokens": 4904742.0,
+ "eval_runtime": 89.5868,
+ "eval_samples_per_second": 15.851,
+ "eval_steps_per_second": 1.987,
+ "step": 2120
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.1198535010664653e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2140/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2140/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2140/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2140/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2140/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2140/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2140/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2140/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2140/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2140/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2140/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2140/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2140/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2140/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..db7a424e0b6ebdcd2813c65c3c18491f016c739b
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2140/trainer_state.json
@@ -0,0 +1,2281 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 5.15691007845504,
+ "eval_steps": 20,
+ "global_step": 2140,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ },
+ {
+ "entropy": 0.4528769593685865,
+ "epoch": 4.579360289680145,
+ "grad_norm": 0.7440020442008972,
+ "learning_rate": 0.0002401706127645863,
+ "loss": 0.43678741455078124,
+ "mean_token_accuracy": 0.8619599856436253,
+ "num_tokens": 4401934.0,
+ "step": 1900
+ },
+ {
+ "epoch": 4.579360289680145,
+ "eval_entropy": 0.5085621437664782,
+ "eval_loss": 0.6766273379325867,
+ "eval_mean_token_accuracy": 0.8227389695939054,
+ "eval_num_tokens": 4401934.0,
+ "eval_runtime": 89.4661,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 1900
+ },
+ {
+ "entropy": 0.4356316838413477,
+ "epoch": 4.627640313820157,
+ "grad_norm": 1.3870011568069458,
+ "learning_rate": 0.0002372530330023795,
+ "loss": 0.41338543891906737,
+ "mean_token_accuracy": 0.8642870113253593,
+ "num_tokens": 4449745.0,
+ "step": 1920
+ },
+ {
+ "epoch": 4.627640313820157,
+ "eval_entropy": 0.4954608427674583,
+ "eval_loss": 0.6783625483512878,
+ "eval_mean_token_accuracy": 0.8214637864841504,
+ "eval_num_tokens": 4449745.0,
+ "eval_runtime": 89.2291,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 1920
+ },
+ {
+ "entropy": 0.4468349639326334,
+ "epoch": 4.675920337960169,
+ "grad_norm": 0.8028655648231506,
+ "learning_rate": 0.00023431991948050538,
+ "loss": 0.4324653625488281,
+ "mean_token_accuracy": 0.8607208795845509,
+ "num_tokens": 4495074.0,
+ "step": 1940
+ },
+ {
+ "epoch": 4.675920337960169,
+ "eval_entropy": 0.5259480836351266,
+ "eval_loss": 0.6809864044189453,
+ "eval_mean_token_accuracy": 0.8200277019752545,
+ "eval_num_tokens": 4495074.0,
+ "eval_runtime": 89.8247,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 1940
+ },
+ {
+ "entropy": 0.4610355503857136,
+ "epoch": 4.724200362100181,
+ "grad_norm": 0.8928298950195312,
+ "learning_rate": 0.0002313721022353953,
+ "loss": 0.4314274787902832,
+ "mean_token_accuracy": 0.8559959702193737,
+ "num_tokens": 4541053.0,
+ "step": 1960
+ },
+ {
+ "epoch": 4.724200362100181,
+ "eval_entropy": 0.501018373651451,
+ "eval_loss": 0.6834176182746887,
+ "eval_mean_token_accuracy": 0.8220113566082515,
+ "eval_num_tokens": 4541053.0,
+ "eval_runtime": 90.4088,
+ "eval_samples_per_second": 15.706,
+ "eval_steps_per_second": 1.969,
+ "step": 1960
+ },
+ {
+ "entropy": 0.4568953149020672,
+ "epoch": 4.772480386240193,
+ "grad_norm": 1.018485188484192,
+ "learning_rate": 0.00022841041546446032,
+ "loss": 0.4306424617767334,
+ "mean_token_accuracy": 0.8598173558712006,
+ "num_tokens": 4584210.0,
+ "step": 1980
+ },
+ {
+ "epoch": 4.772480386240193,
+ "eval_entropy": 0.47633055371514865,
+ "eval_loss": 0.6886085867881775,
+ "eval_mean_token_accuracy": 0.823554899585381,
+ "eval_num_tokens": 4584210.0,
+ "eval_runtime": 89.8185,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 1980
+ },
+ {
+ "entropy": 0.4718513362109661,
+ "epoch": 4.820760410380205,
+ "grad_norm": 0.7674988508224487,
+ "learning_rate": 0.00022543569729002318,
+ "loss": 0.44841318130493163,
+ "mean_token_accuracy": 0.8566767282783985,
+ "num_tokens": 4630254.0,
+ "step": 2000
+ },
+ {
+ "epoch": 4.820760410380205,
+ "eval_entropy": 0.49708309284086977,
+ "eval_loss": 0.6878687739372253,
+ "eval_mean_token_accuracy": 0.8217844698536262,
+ "eval_num_tokens": 4630254.0,
+ "eval_runtime": 89.7497,
+ "eval_samples_per_second": 15.822,
+ "eval_steps_per_second": 1.983,
+ "step": 2000
+ },
+ {
+ "entropy": 0.44017268233001233,
+ "epoch": 4.869040434520217,
+ "grad_norm": 0.6380876898765564,
+ "learning_rate": 0.00022244878952213976,
+ "loss": 0.42589750289916994,
+ "mean_token_accuracy": 0.8615814067423344,
+ "num_tokens": 4679265.0,
+ "step": 2020
+ },
+ {
+ "epoch": 4.869040434520217,
+ "eval_entropy": 0.5042208635740066,
+ "eval_loss": 0.6718228459358215,
+ "eval_mean_token_accuracy": 0.8232053056191863,
+ "eval_num_tokens": 4679265.0,
+ "eval_runtime": 89.7331,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 2020
+ },
+ {
+ "entropy": 0.4654298175126314,
+ "epoch": 4.917320458660229,
+ "grad_norm": 0.8285323977470398,
+ "learning_rate": 0.0002194505374203766,
+ "loss": 0.44148526191711424,
+ "mean_token_accuracy": 0.8545936144888401,
+ "num_tokens": 4723592.0,
+ "step": 2040
+ },
+ {
+ "epoch": 4.917320458660229,
+ "eval_entropy": 0.4957777807551823,
+ "eval_loss": 0.6694385409355164,
+ "eval_mean_token_accuracy": 0.8243840380331103,
+ "eval_num_tokens": 4723592.0,
+ "eval_runtime": 89.5293,
+ "eval_samples_per_second": 15.861,
+ "eval_steps_per_second": 1.988,
+ "step": 2040
+ },
+ {
+ "entropy": 0.46651641055941584,
+ "epoch": 4.965600482800241,
+ "grad_norm": 0.946260929107666,
+ "learning_rate": 0.00021644178945461268,
+ "loss": 0.43698992729187014,
+ "mean_token_accuracy": 0.8586322009563446,
+ "num_tokens": 4767253.0,
+ "step": 2060
+ },
+ {
+ "epoch": 4.965600482800241,
+ "eval_entropy": 0.5063281496254246,
+ "eval_loss": 0.6729032397270203,
+ "eval_mean_token_accuracy": 0.824299869912394,
+ "eval_num_tokens": 4767253.0,
+ "eval_runtime": 89.6207,
+ "eval_samples_per_second": 15.845,
+ "eval_steps_per_second": 1.986,
+ "step": 2060
+ },
+ {
+ "entropy": 0.4135968596130222,
+ "epoch": 5.012070006035003,
+ "grad_norm": 0.6853868961334229,
+ "learning_rate": 0.0002134233970649322,
+ "loss": 0.3800280332565308,
+ "mean_token_accuracy": 0.8756770677380747,
+ "num_tokens": 4816587.0,
+ "step": 2080
+ },
+ {
+ "epoch": 5.012070006035003,
+ "eval_entropy": 0.4107752423942759,
+ "eval_loss": 0.7204955220222473,
+ "eval_mean_token_accuracy": 0.8226918417416261,
+ "eval_num_tokens": 4816587.0,
+ "eval_runtime": 89.5817,
+ "eval_samples_per_second": 15.851,
+ "eval_steps_per_second": 1.987,
+ "step": 2080
+ },
+ {
+ "entropy": 0.316746598854661,
+ "epoch": 5.060350030175015,
+ "grad_norm": 0.6939145922660828,
+ "learning_rate": 0.00021039621442067732,
+ "loss": 0.2942554473876953,
+ "mean_token_accuracy": 0.90015934035182,
+ "num_tokens": 4861071.0,
+ "step": 2100
+ },
+ {
+ "epoch": 5.060350030175015,
+ "eval_entropy": 0.4214997919422857,
+ "eval_loss": 0.7268305420875549,
+ "eval_mean_token_accuracy": 0.8216965369294199,
+ "eval_num_tokens": 4861071.0,
+ "eval_runtime": 89.8077,
+ "eval_samples_per_second": 15.812,
+ "eval_steps_per_second": 1.982,
+ "step": 2100
+ },
+ {
+ "entropy": 0.3187096672132611,
+ "epoch": 5.108630054315027,
+ "grad_norm": 0.9134758114814758,
+ "learning_rate": 0.00020736109817872823,
+ "loss": 0.2839747190475464,
+ "mean_token_accuracy": 0.9006688073277473,
+ "num_tokens": 4904742.0,
+ "step": 2120
+ },
+ {
+ "epoch": 5.108630054315027,
+ "eval_entropy": 0.4375539443800958,
+ "eval_loss": 0.7280111908912659,
+ "eval_mean_token_accuracy": 0.8186122694712007,
+ "eval_num_tokens": 4904742.0,
+ "eval_runtime": 89.5868,
+ "eval_samples_per_second": 15.851,
+ "eval_steps_per_second": 1.987,
+ "step": 2120
+ },
+ {
+ "entropy": 0.3137735094875097,
+ "epoch": 5.15691007845504,
+ "grad_norm": 0.7324768304824829,
+ "learning_rate": 0.00020431890724107943,
+ "loss": 0.28816254138946534,
+ "mean_token_accuracy": 0.9023615352809429,
+ "num_tokens": 4950795.0,
+ "step": 2140
+ },
+ {
+ "epoch": 5.15691007845504,
+ "eval_entropy": 0.42255786312430094,
+ "eval_loss": 0.7335684895515442,
+ "eval_mean_token_accuracy": 0.8213735473959634,
+ "eval_num_tokens": 4950795.0,
+ "eval_runtime": 89.5711,
+ "eval_samples_per_second": 15.853,
+ "eval_steps_per_second": 1.987,
+ "step": 2140
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.1394198722919834e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2160/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2160/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2160/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2160/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2160/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2160/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2160/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2160/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2160/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2160/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2160/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2160/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2160/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2160/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..3bd0ca432c3b3712f1d32f102181b93da14ddebe
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2160/trainer_state.json
@@ -0,0 +1,2302 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 5.2051901025950515,
+ "eval_steps": 20,
+ "global_step": 2160,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ },
+ {
+ "entropy": 0.4528769593685865,
+ "epoch": 4.579360289680145,
+ "grad_norm": 0.7440020442008972,
+ "learning_rate": 0.0002401706127645863,
+ "loss": 0.43678741455078124,
+ "mean_token_accuracy": 0.8619599856436253,
+ "num_tokens": 4401934.0,
+ "step": 1900
+ },
+ {
+ "epoch": 4.579360289680145,
+ "eval_entropy": 0.5085621437664782,
+ "eval_loss": 0.6766273379325867,
+ "eval_mean_token_accuracy": 0.8227389695939054,
+ "eval_num_tokens": 4401934.0,
+ "eval_runtime": 89.4661,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 1900
+ },
+ {
+ "entropy": 0.4356316838413477,
+ "epoch": 4.627640313820157,
+ "grad_norm": 1.3870011568069458,
+ "learning_rate": 0.0002372530330023795,
+ "loss": 0.41338543891906737,
+ "mean_token_accuracy": 0.8642870113253593,
+ "num_tokens": 4449745.0,
+ "step": 1920
+ },
+ {
+ "epoch": 4.627640313820157,
+ "eval_entropy": 0.4954608427674583,
+ "eval_loss": 0.6783625483512878,
+ "eval_mean_token_accuracy": 0.8214637864841504,
+ "eval_num_tokens": 4449745.0,
+ "eval_runtime": 89.2291,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 1920
+ },
+ {
+ "entropy": 0.4468349639326334,
+ "epoch": 4.675920337960169,
+ "grad_norm": 0.8028655648231506,
+ "learning_rate": 0.00023431991948050538,
+ "loss": 0.4324653625488281,
+ "mean_token_accuracy": 0.8607208795845509,
+ "num_tokens": 4495074.0,
+ "step": 1940
+ },
+ {
+ "epoch": 4.675920337960169,
+ "eval_entropy": 0.5259480836351266,
+ "eval_loss": 0.6809864044189453,
+ "eval_mean_token_accuracy": 0.8200277019752545,
+ "eval_num_tokens": 4495074.0,
+ "eval_runtime": 89.8247,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 1940
+ },
+ {
+ "entropy": 0.4610355503857136,
+ "epoch": 4.724200362100181,
+ "grad_norm": 0.8928298950195312,
+ "learning_rate": 0.0002313721022353953,
+ "loss": 0.4314274787902832,
+ "mean_token_accuracy": 0.8559959702193737,
+ "num_tokens": 4541053.0,
+ "step": 1960
+ },
+ {
+ "epoch": 4.724200362100181,
+ "eval_entropy": 0.501018373651451,
+ "eval_loss": 0.6834176182746887,
+ "eval_mean_token_accuracy": 0.8220113566082515,
+ "eval_num_tokens": 4541053.0,
+ "eval_runtime": 90.4088,
+ "eval_samples_per_second": 15.706,
+ "eval_steps_per_second": 1.969,
+ "step": 1960
+ },
+ {
+ "entropy": 0.4568953149020672,
+ "epoch": 4.772480386240193,
+ "grad_norm": 1.018485188484192,
+ "learning_rate": 0.00022841041546446032,
+ "loss": 0.4306424617767334,
+ "mean_token_accuracy": 0.8598173558712006,
+ "num_tokens": 4584210.0,
+ "step": 1980
+ },
+ {
+ "epoch": 4.772480386240193,
+ "eval_entropy": 0.47633055371514865,
+ "eval_loss": 0.6886085867881775,
+ "eval_mean_token_accuracy": 0.823554899585381,
+ "eval_num_tokens": 4584210.0,
+ "eval_runtime": 89.8185,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 1980
+ },
+ {
+ "entropy": 0.4718513362109661,
+ "epoch": 4.820760410380205,
+ "grad_norm": 0.7674988508224487,
+ "learning_rate": 0.00022543569729002318,
+ "loss": 0.44841318130493163,
+ "mean_token_accuracy": 0.8566767282783985,
+ "num_tokens": 4630254.0,
+ "step": 2000
+ },
+ {
+ "epoch": 4.820760410380205,
+ "eval_entropy": 0.49708309284086977,
+ "eval_loss": 0.6878687739372253,
+ "eval_mean_token_accuracy": 0.8217844698536262,
+ "eval_num_tokens": 4630254.0,
+ "eval_runtime": 89.7497,
+ "eval_samples_per_second": 15.822,
+ "eval_steps_per_second": 1.983,
+ "step": 2000
+ },
+ {
+ "entropy": 0.44017268233001233,
+ "epoch": 4.869040434520217,
+ "grad_norm": 0.6380876898765564,
+ "learning_rate": 0.00022244878952213976,
+ "loss": 0.42589750289916994,
+ "mean_token_accuracy": 0.8615814067423344,
+ "num_tokens": 4679265.0,
+ "step": 2020
+ },
+ {
+ "epoch": 4.869040434520217,
+ "eval_entropy": 0.5042208635740066,
+ "eval_loss": 0.6718228459358215,
+ "eval_mean_token_accuracy": 0.8232053056191863,
+ "eval_num_tokens": 4679265.0,
+ "eval_runtime": 89.7331,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 2020
+ },
+ {
+ "entropy": 0.4654298175126314,
+ "epoch": 4.917320458660229,
+ "grad_norm": 0.8285323977470398,
+ "learning_rate": 0.0002194505374203766,
+ "loss": 0.44148526191711424,
+ "mean_token_accuracy": 0.8545936144888401,
+ "num_tokens": 4723592.0,
+ "step": 2040
+ },
+ {
+ "epoch": 4.917320458660229,
+ "eval_entropy": 0.4957777807551823,
+ "eval_loss": 0.6694385409355164,
+ "eval_mean_token_accuracy": 0.8243840380331103,
+ "eval_num_tokens": 4723592.0,
+ "eval_runtime": 89.5293,
+ "eval_samples_per_second": 15.861,
+ "eval_steps_per_second": 1.988,
+ "step": 2040
+ },
+ {
+ "entropy": 0.46651641055941584,
+ "epoch": 4.965600482800241,
+ "grad_norm": 0.946260929107666,
+ "learning_rate": 0.00021644178945461268,
+ "loss": 0.43698992729187014,
+ "mean_token_accuracy": 0.8586322009563446,
+ "num_tokens": 4767253.0,
+ "step": 2060
+ },
+ {
+ "epoch": 4.965600482800241,
+ "eval_entropy": 0.5063281496254246,
+ "eval_loss": 0.6729032397270203,
+ "eval_mean_token_accuracy": 0.824299869912394,
+ "eval_num_tokens": 4767253.0,
+ "eval_runtime": 89.6207,
+ "eval_samples_per_second": 15.845,
+ "eval_steps_per_second": 1.986,
+ "step": 2060
+ },
+ {
+ "entropy": 0.4135968596130222,
+ "epoch": 5.012070006035003,
+ "grad_norm": 0.6853868961334229,
+ "learning_rate": 0.0002134233970649322,
+ "loss": 0.3800280332565308,
+ "mean_token_accuracy": 0.8756770677380747,
+ "num_tokens": 4816587.0,
+ "step": 2080
+ },
+ {
+ "epoch": 5.012070006035003,
+ "eval_entropy": 0.4107752423942759,
+ "eval_loss": 0.7204955220222473,
+ "eval_mean_token_accuracy": 0.8226918417416261,
+ "eval_num_tokens": 4816587.0,
+ "eval_runtime": 89.5817,
+ "eval_samples_per_second": 15.851,
+ "eval_steps_per_second": 1.987,
+ "step": 2080
+ },
+ {
+ "entropy": 0.316746598854661,
+ "epoch": 5.060350030175015,
+ "grad_norm": 0.6939145922660828,
+ "learning_rate": 0.00021039621442067732,
+ "loss": 0.2942554473876953,
+ "mean_token_accuracy": 0.90015934035182,
+ "num_tokens": 4861071.0,
+ "step": 2100
+ },
+ {
+ "epoch": 5.060350030175015,
+ "eval_entropy": 0.4214997919422857,
+ "eval_loss": 0.7268305420875549,
+ "eval_mean_token_accuracy": 0.8216965369294199,
+ "eval_num_tokens": 4861071.0,
+ "eval_runtime": 89.8077,
+ "eval_samples_per_second": 15.812,
+ "eval_steps_per_second": 1.982,
+ "step": 2100
+ },
+ {
+ "entropy": 0.3187096672132611,
+ "epoch": 5.108630054315027,
+ "grad_norm": 0.9134758114814758,
+ "learning_rate": 0.00020736109817872823,
+ "loss": 0.2839747190475464,
+ "mean_token_accuracy": 0.9006688073277473,
+ "num_tokens": 4904742.0,
+ "step": 2120
+ },
+ {
+ "epoch": 5.108630054315027,
+ "eval_entropy": 0.4375539443800958,
+ "eval_loss": 0.7280111908912659,
+ "eval_mean_token_accuracy": 0.8186122694712007,
+ "eval_num_tokens": 4904742.0,
+ "eval_runtime": 89.5868,
+ "eval_samples_per_second": 15.851,
+ "eval_steps_per_second": 1.987,
+ "step": 2120
+ },
+ {
+ "entropy": 0.3137735094875097,
+ "epoch": 5.15691007845504,
+ "grad_norm": 0.7324768304824829,
+ "learning_rate": 0.00020431890724107943,
+ "loss": 0.28816254138946534,
+ "mean_token_accuracy": 0.9023615352809429,
+ "num_tokens": 4950795.0,
+ "step": 2140
+ },
+ {
+ "epoch": 5.15691007845504,
+ "eval_entropy": 0.42255786312430094,
+ "eval_loss": 0.7335684895515442,
+ "eval_mean_token_accuracy": 0.8213735473959634,
+ "eval_num_tokens": 4950795.0,
+ "eval_runtime": 89.5711,
+ "eval_samples_per_second": 15.853,
+ "eval_steps_per_second": 1.987,
+ "step": 2140
+ },
+ {
+ "entropy": 0.3310150146484375,
+ "epoch": 5.2051901025950515,
+ "grad_norm": 0.7819830179214478,
+ "learning_rate": 0.00020127050251178062,
+ "loss": 0.3020722150802612,
+ "mean_token_accuracy": 0.8957489147782326,
+ "num_tokens": 4994886.0,
+ "step": 2160
+ },
+ {
+ "epoch": 5.2051901025950515,
+ "eval_entropy": 0.43374256186940696,
+ "eval_loss": 0.7270973920822144,
+ "eval_mean_token_accuracy": 0.821890847066815,
+ "eval_num_tokens": 4994886.0,
+ "eval_runtime": 89.7117,
+ "eval_samples_per_second": 15.828,
+ "eval_steps_per_second": 1.984,
+ "step": 2160
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.1585509166656102e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2180/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2180/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2180/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2180/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2180/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2180/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2180/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2180/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2180/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2180/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2180/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2180/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2180/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2180/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..7086584e4e9e572daadcb4d242e851934a239fa2
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2180/trainer_state.json
@@ -0,0 +1,2323 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 5.253470126735063,
+ "eval_steps": 20,
+ "global_step": 2180,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ },
+ {
+ "entropy": 0.4528769593685865,
+ "epoch": 4.579360289680145,
+ "grad_norm": 0.7440020442008972,
+ "learning_rate": 0.0002401706127645863,
+ "loss": 0.43678741455078124,
+ "mean_token_accuracy": 0.8619599856436253,
+ "num_tokens": 4401934.0,
+ "step": 1900
+ },
+ {
+ "epoch": 4.579360289680145,
+ "eval_entropy": 0.5085621437664782,
+ "eval_loss": 0.6766273379325867,
+ "eval_mean_token_accuracy": 0.8227389695939054,
+ "eval_num_tokens": 4401934.0,
+ "eval_runtime": 89.4661,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 1900
+ },
+ {
+ "entropy": 0.4356316838413477,
+ "epoch": 4.627640313820157,
+ "grad_norm": 1.3870011568069458,
+ "learning_rate": 0.0002372530330023795,
+ "loss": 0.41338543891906737,
+ "mean_token_accuracy": 0.8642870113253593,
+ "num_tokens": 4449745.0,
+ "step": 1920
+ },
+ {
+ "epoch": 4.627640313820157,
+ "eval_entropy": 0.4954608427674583,
+ "eval_loss": 0.6783625483512878,
+ "eval_mean_token_accuracy": 0.8214637864841504,
+ "eval_num_tokens": 4449745.0,
+ "eval_runtime": 89.2291,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 1920
+ },
+ {
+ "entropy": 0.4468349639326334,
+ "epoch": 4.675920337960169,
+ "grad_norm": 0.8028655648231506,
+ "learning_rate": 0.00023431991948050538,
+ "loss": 0.4324653625488281,
+ "mean_token_accuracy": 0.8607208795845509,
+ "num_tokens": 4495074.0,
+ "step": 1940
+ },
+ {
+ "epoch": 4.675920337960169,
+ "eval_entropy": 0.5259480836351266,
+ "eval_loss": 0.6809864044189453,
+ "eval_mean_token_accuracy": 0.8200277019752545,
+ "eval_num_tokens": 4495074.0,
+ "eval_runtime": 89.8247,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 1940
+ },
+ {
+ "entropy": 0.4610355503857136,
+ "epoch": 4.724200362100181,
+ "grad_norm": 0.8928298950195312,
+ "learning_rate": 0.0002313721022353953,
+ "loss": 0.4314274787902832,
+ "mean_token_accuracy": 0.8559959702193737,
+ "num_tokens": 4541053.0,
+ "step": 1960
+ },
+ {
+ "epoch": 4.724200362100181,
+ "eval_entropy": 0.501018373651451,
+ "eval_loss": 0.6834176182746887,
+ "eval_mean_token_accuracy": 0.8220113566082515,
+ "eval_num_tokens": 4541053.0,
+ "eval_runtime": 90.4088,
+ "eval_samples_per_second": 15.706,
+ "eval_steps_per_second": 1.969,
+ "step": 1960
+ },
+ {
+ "entropy": 0.4568953149020672,
+ "epoch": 4.772480386240193,
+ "grad_norm": 1.018485188484192,
+ "learning_rate": 0.00022841041546446032,
+ "loss": 0.4306424617767334,
+ "mean_token_accuracy": 0.8598173558712006,
+ "num_tokens": 4584210.0,
+ "step": 1980
+ },
+ {
+ "epoch": 4.772480386240193,
+ "eval_entropy": 0.47633055371514865,
+ "eval_loss": 0.6886085867881775,
+ "eval_mean_token_accuracy": 0.823554899585381,
+ "eval_num_tokens": 4584210.0,
+ "eval_runtime": 89.8185,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 1980
+ },
+ {
+ "entropy": 0.4718513362109661,
+ "epoch": 4.820760410380205,
+ "grad_norm": 0.7674988508224487,
+ "learning_rate": 0.00022543569729002318,
+ "loss": 0.44841318130493163,
+ "mean_token_accuracy": 0.8566767282783985,
+ "num_tokens": 4630254.0,
+ "step": 2000
+ },
+ {
+ "epoch": 4.820760410380205,
+ "eval_entropy": 0.49708309284086977,
+ "eval_loss": 0.6878687739372253,
+ "eval_mean_token_accuracy": 0.8217844698536262,
+ "eval_num_tokens": 4630254.0,
+ "eval_runtime": 89.7497,
+ "eval_samples_per_second": 15.822,
+ "eval_steps_per_second": 1.983,
+ "step": 2000
+ },
+ {
+ "entropy": 0.44017268233001233,
+ "epoch": 4.869040434520217,
+ "grad_norm": 0.6380876898765564,
+ "learning_rate": 0.00022244878952213976,
+ "loss": 0.42589750289916994,
+ "mean_token_accuracy": 0.8615814067423344,
+ "num_tokens": 4679265.0,
+ "step": 2020
+ },
+ {
+ "epoch": 4.869040434520217,
+ "eval_entropy": 0.5042208635740066,
+ "eval_loss": 0.6718228459358215,
+ "eval_mean_token_accuracy": 0.8232053056191863,
+ "eval_num_tokens": 4679265.0,
+ "eval_runtime": 89.7331,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 2020
+ },
+ {
+ "entropy": 0.4654298175126314,
+ "epoch": 4.917320458660229,
+ "grad_norm": 0.8285323977470398,
+ "learning_rate": 0.0002194505374203766,
+ "loss": 0.44148526191711424,
+ "mean_token_accuracy": 0.8545936144888401,
+ "num_tokens": 4723592.0,
+ "step": 2040
+ },
+ {
+ "epoch": 4.917320458660229,
+ "eval_entropy": 0.4957777807551823,
+ "eval_loss": 0.6694385409355164,
+ "eval_mean_token_accuracy": 0.8243840380331103,
+ "eval_num_tokens": 4723592.0,
+ "eval_runtime": 89.5293,
+ "eval_samples_per_second": 15.861,
+ "eval_steps_per_second": 1.988,
+ "step": 2040
+ },
+ {
+ "entropy": 0.46651641055941584,
+ "epoch": 4.965600482800241,
+ "grad_norm": 0.946260929107666,
+ "learning_rate": 0.00021644178945461268,
+ "loss": 0.43698992729187014,
+ "mean_token_accuracy": 0.8586322009563446,
+ "num_tokens": 4767253.0,
+ "step": 2060
+ },
+ {
+ "epoch": 4.965600482800241,
+ "eval_entropy": 0.5063281496254246,
+ "eval_loss": 0.6729032397270203,
+ "eval_mean_token_accuracy": 0.824299869912394,
+ "eval_num_tokens": 4767253.0,
+ "eval_runtime": 89.6207,
+ "eval_samples_per_second": 15.845,
+ "eval_steps_per_second": 1.986,
+ "step": 2060
+ },
+ {
+ "entropy": 0.4135968596130222,
+ "epoch": 5.012070006035003,
+ "grad_norm": 0.6853868961334229,
+ "learning_rate": 0.0002134233970649322,
+ "loss": 0.3800280332565308,
+ "mean_token_accuracy": 0.8756770677380747,
+ "num_tokens": 4816587.0,
+ "step": 2080
+ },
+ {
+ "epoch": 5.012070006035003,
+ "eval_entropy": 0.4107752423942759,
+ "eval_loss": 0.7204955220222473,
+ "eval_mean_token_accuracy": 0.8226918417416261,
+ "eval_num_tokens": 4816587.0,
+ "eval_runtime": 89.5817,
+ "eval_samples_per_second": 15.851,
+ "eval_steps_per_second": 1.987,
+ "step": 2080
+ },
+ {
+ "entropy": 0.316746598854661,
+ "epoch": 5.060350030175015,
+ "grad_norm": 0.6939145922660828,
+ "learning_rate": 0.00021039621442067732,
+ "loss": 0.2942554473876953,
+ "mean_token_accuracy": 0.90015934035182,
+ "num_tokens": 4861071.0,
+ "step": 2100
+ },
+ {
+ "epoch": 5.060350030175015,
+ "eval_entropy": 0.4214997919422857,
+ "eval_loss": 0.7268305420875549,
+ "eval_mean_token_accuracy": 0.8216965369294199,
+ "eval_num_tokens": 4861071.0,
+ "eval_runtime": 89.8077,
+ "eval_samples_per_second": 15.812,
+ "eval_steps_per_second": 1.982,
+ "step": 2100
+ },
+ {
+ "entropy": 0.3187096672132611,
+ "epoch": 5.108630054315027,
+ "grad_norm": 0.9134758114814758,
+ "learning_rate": 0.00020736109817872823,
+ "loss": 0.2839747190475464,
+ "mean_token_accuracy": 0.9006688073277473,
+ "num_tokens": 4904742.0,
+ "step": 2120
+ },
+ {
+ "epoch": 5.108630054315027,
+ "eval_entropy": 0.4375539443800958,
+ "eval_loss": 0.7280111908912659,
+ "eval_mean_token_accuracy": 0.8186122694712007,
+ "eval_num_tokens": 4904742.0,
+ "eval_runtime": 89.5868,
+ "eval_samples_per_second": 15.851,
+ "eval_steps_per_second": 1.987,
+ "step": 2120
+ },
+ {
+ "entropy": 0.3137735094875097,
+ "epoch": 5.15691007845504,
+ "grad_norm": 0.7324768304824829,
+ "learning_rate": 0.00020431890724107943,
+ "loss": 0.28816254138946534,
+ "mean_token_accuracy": 0.9023615352809429,
+ "num_tokens": 4950795.0,
+ "step": 2140
+ },
+ {
+ "epoch": 5.15691007845504,
+ "eval_entropy": 0.42255786312430094,
+ "eval_loss": 0.7335684895515442,
+ "eval_mean_token_accuracy": 0.8213735473959634,
+ "eval_num_tokens": 4950795.0,
+ "eval_runtime": 89.5711,
+ "eval_samples_per_second": 15.853,
+ "eval_steps_per_second": 1.987,
+ "step": 2140
+ },
+ {
+ "entropy": 0.3310150146484375,
+ "epoch": 5.2051901025950515,
+ "grad_norm": 0.7819830179214478,
+ "learning_rate": 0.00020127050251178062,
+ "loss": 0.3020722150802612,
+ "mean_token_accuracy": 0.8957489147782326,
+ "num_tokens": 4994886.0,
+ "step": 2160
+ },
+ {
+ "epoch": 5.2051901025950515,
+ "eval_entropy": 0.43374256186940696,
+ "eval_loss": 0.7270973920822144,
+ "eval_mean_token_accuracy": 0.821890847066815,
+ "eval_num_tokens": 4994886.0,
+ "eval_runtime": 89.7117,
+ "eval_samples_per_second": 15.828,
+ "eval_steps_per_second": 1.984,
+ "step": 2160
+ },
+ {
+ "entropy": 0.31917148139327767,
+ "epoch": 5.253470126735063,
+ "grad_norm": 0.8766310811042786,
+ "learning_rate": 0.00019821674665331112,
+ "loss": 0.29281470775604246,
+ "mean_token_accuracy": 0.8986819669604301,
+ "num_tokens": 5043663.0,
+ "step": 2180
+ },
+ {
+ "epoch": 5.253470126735063,
+ "eval_entropy": 0.4188300051381079,
+ "eval_loss": 0.7247599959373474,
+ "eval_mean_token_accuracy": 0.8234477113471942,
+ "eval_num_tokens": 5043663.0,
+ "eval_runtime": 89.384,
+ "eval_samples_per_second": 15.887,
+ "eval_steps_per_second": 1.991,
+ "step": 2180
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.1793047877365965e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-220/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-220/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-220/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-220/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-220/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-220/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-220/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-220/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-220/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-220/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-220/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-220/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-220/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-220/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..6c830e84e02e6e812b3d8ba0a1d8343883b11c2c
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-220/trainer_state.json
@@ -0,0 +1,265 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 0.5310802655401328,
+ "eval_steps": 20,
+ "global_step": 220,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.23138670886912e+16,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2200/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2200/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2200/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2200/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2200/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2200/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2200/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2200/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2200/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2200/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2200/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2200/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2200/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2200/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..c3a443e0a6585a6293a9accaa4885706948ba47b
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2200/trainer_state.json
@@ -0,0 +1,2344 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 5.301750150875075,
+ "eval_steps": 20,
+ "global_step": 2200,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ },
+ {
+ "entropy": 0.4528769593685865,
+ "epoch": 4.579360289680145,
+ "grad_norm": 0.7440020442008972,
+ "learning_rate": 0.0002401706127645863,
+ "loss": 0.43678741455078124,
+ "mean_token_accuracy": 0.8619599856436253,
+ "num_tokens": 4401934.0,
+ "step": 1900
+ },
+ {
+ "epoch": 4.579360289680145,
+ "eval_entropy": 0.5085621437664782,
+ "eval_loss": 0.6766273379325867,
+ "eval_mean_token_accuracy": 0.8227389695939054,
+ "eval_num_tokens": 4401934.0,
+ "eval_runtime": 89.4661,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 1900
+ },
+ {
+ "entropy": 0.4356316838413477,
+ "epoch": 4.627640313820157,
+ "grad_norm": 1.3870011568069458,
+ "learning_rate": 0.0002372530330023795,
+ "loss": 0.41338543891906737,
+ "mean_token_accuracy": 0.8642870113253593,
+ "num_tokens": 4449745.0,
+ "step": 1920
+ },
+ {
+ "epoch": 4.627640313820157,
+ "eval_entropy": 0.4954608427674583,
+ "eval_loss": 0.6783625483512878,
+ "eval_mean_token_accuracy": 0.8214637864841504,
+ "eval_num_tokens": 4449745.0,
+ "eval_runtime": 89.2291,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 1920
+ },
+ {
+ "entropy": 0.4468349639326334,
+ "epoch": 4.675920337960169,
+ "grad_norm": 0.8028655648231506,
+ "learning_rate": 0.00023431991948050538,
+ "loss": 0.4324653625488281,
+ "mean_token_accuracy": 0.8607208795845509,
+ "num_tokens": 4495074.0,
+ "step": 1940
+ },
+ {
+ "epoch": 4.675920337960169,
+ "eval_entropy": 0.5259480836351266,
+ "eval_loss": 0.6809864044189453,
+ "eval_mean_token_accuracy": 0.8200277019752545,
+ "eval_num_tokens": 4495074.0,
+ "eval_runtime": 89.8247,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 1940
+ },
+ {
+ "entropy": 0.4610355503857136,
+ "epoch": 4.724200362100181,
+ "grad_norm": 0.8928298950195312,
+ "learning_rate": 0.0002313721022353953,
+ "loss": 0.4314274787902832,
+ "mean_token_accuracy": 0.8559959702193737,
+ "num_tokens": 4541053.0,
+ "step": 1960
+ },
+ {
+ "epoch": 4.724200362100181,
+ "eval_entropy": 0.501018373651451,
+ "eval_loss": 0.6834176182746887,
+ "eval_mean_token_accuracy": 0.8220113566082515,
+ "eval_num_tokens": 4541053.0,
+ "eval_runtime": 90.4088,
+ "eval_samples_per_second": 15.706,
+ "eval_steps_per_second": 1.969,
+ "step": 1960
+ },
+ {
+ "entropy": 0.4568953149020672,
+ "epoch": 4.772480386240193,
+ "grad_norm": 1.018485188484192,
+ "learning_rate": 0.00022841041546446032,
+ "loss": 0.4306424617767334,
+ "mean_token_accuracy": 0.8598173558712006,
+ "num_tokens": 4584210.0,
+ "step": 1980
+ },
+ {
+ "epoch": 4.772480386240193,
+ "eval_entropy": 0.47633055371514865,
+ "eval_loss": 0.6886085867881775,
+ "eval_mean_token_accuracy": 0.823554899585381,
+ "eval_num_tokens": 4584210.0,
+ "eval_runtime": 89.8185,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 1980
+ },
+ {
+ "entropy": 0.4718513362109661,
+ "epoch": 4.820760410380205,
+ "grad_norm": 0.7674988508224487,
+ "learning_rate": 0.00022543569729002318,
+ "loss": 0.44841318130493163,
+ "mean_token_accuracy": 0.8566767282783985,
+ "num_tokens": 4630254.0,
+ "step": 2000
+ },
+ {
+ "epoch": 4.820760410380205,
+ "eval_entropy": 0.49708309284086977,
+ "eval_loss": 0.6878687739372253,
+ "eval_mean_token_accuracy": 0.8217844698536262,
+ "eval_num_tokens": 4630254.0,
+ "eval_runtime": 89.7497,
+ "eval_samples_per_second": 15.822,
+ "eval_steps_per_second": 1.983,
+ "step": 2000
+ },
+ {
+ "entropy": 0.44017268233001233,
+ "epoch": 4.869040434520217,
+ "grad_norm": 0.6380876898765564,
+ "learning_rate": 0.00022244878952213976,
+ "loss": 0.42589750289916994,
+ "mean_token_accuracy": 0.8615814067423344,
+ "num_tokens": 4679265.0,
+ "step": 2020
+ },
+ {
+ "epoch": 4.869040434520217,
+ "eval_entropy": 0.5042208635740066,
+ "eval_loss": 0.6718228459358215,
+ "eval_mean_token_accuracy": 0.8232053056191863,
+ "eval_num_tokens": 4679265.0,
+ "eval_runtime": 89.7331,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 2020
+ },
+ {
+ "entropy": 0.4654298175126314,
+ "epoch": 4.917320458660229,
+ "grad_norm": 0.8285323977470398,
+ "learning_rate": 0.0002194505374203766,
+ "loss": 0.44148526191711424,
+ "mean_token_accuracy": 0.8545936144888401,
+ "num_tokens": 4723592.0,
+ "step": 2040
+ },
+ {
+ "epoch": 4.917320458660229,
+ "eval_entropy": 0.4957777807551823,
+ "eval_loss": 0.6694385409355164,
+ "eval_mean_token_accuracy": 0.8243840380331103,
+ "eval_num_tokens": 4723592.0,
+ "eval_runtime": 89.5293,
+ "eval_samples_per_second": 15.861,
+ "eval_steps_per_second": 1.988,
+ "step": 2040
+ },
+ {
+ "entropy": 0.46651641055941584,
+ "epoch": 4.965600482800241,
+ "grad_norm": 0.946260929107666,
+ "learning_rate": 0.00021644178945461268,
+ "loss": 0.43698992729187014,
+ "mean_token_accuracy": 0.8586322009563446,
+ "num_tokens": 4767253.0,
+ "step": 2060
+ },
+ {
+ "epoch": 4.965600482800241,
+ "eval_entropy": 0.5063281496254246,
+ "eval_loss": 0.6729032397270203,
+ "eval_mean_token_accuracy": 0.824299869912394,
+ "eval_num_tokens": 4767253.0,
+ "eval_runtime": 89.6207,
+ "eval_samples_per_second": 15.845,
+ "eval_steps_per_second": 1.986,
+ "step": 2060
+ },
+ {
+ "entropy": 0.4135968596130222,
+ "epoch": 5.012070006035003,
+ "grad_norm": 0.6853868961334229,
+ "learning_rate": 0.0002134233970649322,
+ "loss": 0.3800280332565308,
+ "mean_token_accuracy": 0.8756770677380747,
+ "num_tokens": 4816587.0,
+ "step": 2080
+ },
+ {
+ "epoch": 5.012070006035003,
+ "eval_entropy": 0.4107752423942759,
+ "eval_loss": 0.7204955220222473,
+ "eval_mean_token_accuracy": 0.8226918417416261,
+ "eval_num_tokens": 4816587.0,
+ "eval_runtime": 89.5817,
+ "eval_samples_per_second": 15.851,
+ "eval_steps_per_second": 1.987,
+ "step": 2080
+ },
+ {
+ "entropy": 0.316746598854661,
+ "epoch": 5.060350030175015,
+ "grad_norm": 0.6939145922660828,
+ "learning_rate": 0.00021039621442067732,
+ "loss": 0.2942554473876953,
+ "mean_token_accuracy": 0.90015934035182,
+ "num_tokens": 4861071.0,
+ "step": 2100
+ },
+ {
+ "epoch": 5.060350030175015,
+ "eval_entropy": 0.4214997919422857,
+ "eval_loss": 0.7268305420875549,
+ "eval_mean_token_accuracy": 0.8216965369294199,
+ "eval_num_tokens": 4861071.0,
+ "eval_runtime": 89.8077,
+ "eval_samples_per_second": 15.812,
+ "eval_steps_per_second": 1.982,
+ "step": 2100
+ },
+ {
+ "entropy": 0.3187096672132611,
+ "epoch": 5.108630054315027,
+ "grad_norm": 0.9134758114814758,
+ "learning_rate": 0.00020736109817872823,
+ "loss": 0.2839747190475464,
+ "mean_token_accuracy": 0.9006688073277473,
+ "num_tokens": 4904742.0,
+ "step": 2120
+ },
+ {
+ "epoch": 5.108630054315027,
+ "eval_entropy": 0.4375539443800958,
+ "eval_loss": 0.7280111908912659,
+ "eval_mean_token_accuracy": 0.8186122694712007,
+ "eval_num_tokens": 4904742.0,
+ "eval_runtime": 89.5868,
+ "eval_samples_per_second": 15.851,
+ "eval_steps_per_second": 1.987,
+ "step": 2120
+ },
+ {
+ "entropy": 0.3137735094875097,
+ "epoch": 5.15691007845504,
+ "grad_norm": 0.7324768304824829,
+ "learning_rate": 0.00020431890724107943,
+ "loss": 0.28816254138946534,
+ "mean_token_accuracy": 0.9023615352809429,
+ "num_tokens": 4950795.0,
+ "step": 2140
+ },
+ {
+ "epoch": 5.15691007845504,
+ "eval_entropy": 0.42255786312430094,
+ "eval_loss": 0.7335684895515442,
+ "eval_mean_token_accuracy": 0.8213735473959634,
+ "eval_num_tokens": 4950795.0,
+ "eval_runtime": 89.5711,
+ "eval_samples_per_second": 15.853,
+ "eval_steps_per_second": 1.987,
+ "step": 2140
+ },
+ {
+ "entropy": 0.3310150146484375,
+ "epoch": 5.2051901025950515,
+ "grad_norm": 0.7819830179214478,
+ "learning_rate": 0.00020127050251178062,
+ "loss": 0.3020722150802612,
+ "mean_token_accuracy": 0.8957489147782326,
+ "num_tokens": 4994886.0,
+ "step": 2160
+ },
+ {
+ "epoch": 5.2051901025950515,
+ "eval_entropy": 0.43374256186940696,
+ "eval_loss": 0.7270973920822144,
+ "eval_mean_token_accuracy": 0.821890847066815,
+ "eval_num_tokens": 4994886.0,
+ "eval_runtime": 89.7117,
+ "eval_samples_per_second": 15.828,
+ "eval_steps_per_second": 1.984,
+ "step": 2160
+ },
+ {
+ "entropy": 0.31917148139327767,
+ "epoch": 5.253470126735063,
+ "grad_norm": 0.8766310811042786,
+ "learning_rate": 0.00019821674665331112,
+ "loss": 0.29281470775604246,
+ "mean_token_accuracy": 0.8986819669604301,
+ "num_tokens": 5043663.0,
+ "step": 2180
+ },
+ {
+ "epoch": 5.253470126735063,
+ "eval_entropy": 0.4188300051381079,
+ "eval_loss": 0.7247599959373474,
+ "eval_mean_token_accuracy": 0.8234477113471942,
+ "eval_num_tokens": 5043663.0,
+ "eval_runtime": 89.384,
+ "eval_samples_per_second": 15.887,
+ "eval_steps_per_second": 1.991,
+ "step": 2180
+ },
+ {
+ "entropy": 0.3252310147508979,
+ "epoch": 5.301750150875075,
+ "grad_norm": 0.8721100687980652,
+ "learning_rate": 0.0001951585038424563,
+ "loss": 0.2924081325531006,
+ "mean_token_accuracy": 0.8981824897229671,
+ "num_tokens": 5087959.0,
+ "step": 2200
+ },
+ {
+ "epoch": 5.301750150875075,
+ "eval_entropy": 0.4114788512835342,
+ "eval_loss": 0.735135555267334,
+ "eval_mean_token_accuracy": 0.8240512961082245,
+ "eval_num_tokens": 5087959.0,
+ "eval_runtime": 89.3499,
+ "eval_samples_per_second": 15.893,
+ "eval_steps_per_second": 1.992,
+ "step": 2200
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.1984771208425677e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2220/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2220/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2220/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2220/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2220/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2220/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2220/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2220/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2220/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2220/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2220/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2220/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2220/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2220/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..607887d7c741f3e0707803dd0bd68accaef453c8
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2220/trainer_state.json
@@ -0,0 +1,2365 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 5.350030175015087,
+ "eval_steps": 20,
+ "global_step": 2220,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ },
+ {
+ "entropy": 0.4528769593685865,
+ "epoch": 4.579360289680145,
+ "grad_norm": 0.7440020442008972,
+ "learning_rate": 0.0002401706127645863,
+ "loss": 0.43678741455078124,
+ "mean_token_accuracy": 0.8619599856436253,
+ "num_tokens": 4401934.0,
+ "step": 1900
+ },
+ {
+ "epoch": 4.579360289680145,
+ "eval_entropy": 0.5085621437664782,
+ "eval_loss": 0.6766273379325867,
+ "eval_mean_token_accuracy": 0.8227389695939054,
+ "eval_num_tokens": 4401934.0,
+ "eval_runtime": 89.4661,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 1900
+ },
+ {
+ "entropy": 0.4356316838413477,
+ "epoch": 4.627640313820157,
+ "grad_norm": 1.3870011568069458,
+ "learning_rate": 0.0002372530330023795,
+ "loss": 0.41338543891906737,
+ "mean_token_accuracy": 0.8642870113253593,
+ "num_tokens": 4449745.0,
+ "step": 1920
+ },
+ {
+ "epoch": 4.627640313820157,
+ "eval_entropy": 0.4954608427674583,
+ "eval_loss": 0.6783625483512878,
+ "eval_mean_token_accuracy": 0.8214637864841504,
+ "eval_num_tokens": 4449745.0,
+ "eval_runtime": 89.2291,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 1920
+ },
+ {
+ "entropy": 0.4468349639326334,
+ "epoch": 4.675920337960169,
+ "grad_norm": 0.8028655648231506,
+ "learning_rate": 0.00023431991948050538,
+ "loss": 0.4324653625488281,
+ "mean_token_accuracy": 0.8607208795845509,
+ "num_tokens": 4495074.0,
+ "step": 1940
+ },
+ {
+ "epoch": 4.675920337960169,
+ "eval_entropy": 0.5259480836351266,
+ "eval_loss": 0.6809864044189453,
+ "eval_mean_token_accuracy": 0.8200277019752545,
+ "eval_num_tokens": 4495074.0,
+ "eval_runtime": 89.8247,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 1940
+ },
+ {
+ "entropy": 0.4610355503857136,
+ "epoch": 4.724200362100181,
+ "grad_norm": 0.8928298950195312,
+ "learning_rate": 0.0002313721022353953,
+ "loss": 0.4314274787902832,
+ "mean_token_accuracy": 0.8559959702193737,
+ "num_tokens": 4541053.0,
+ "step": 1960
+ },
+ {
+ "epoch": 4.724200362100181,
+ "eval_entropy": 0.501018373651451,
+ "eval_loss": 0.6834176182746887,
+ "eval_mean_token_accuracy": 0.8220113566082515,
+ "eval_num_tokens": 4541053.0,
+ "eval_runtime": 90.4088,
+ "eval_samples_per_second": 15.706,
+ "eval_steps_per_second": 1.969,
+ "step": 1960
+ },
+ {
+ "entropy": 0.4568953149020672,
+ "epoch": 4.772480386240193,
+ "grad_norm": 1.018485188484192,
+ "learning_rate": 0.00022841041546446032,
+ "loss": 0.4306424617767334,
+ "mean_token_accuracy": 0.8598173558712006,
+ "num_tokens": 4584210.0,
+ "step": 1980
+ },
+ {
+ "epoch": 4.772480386240193,
+ "eval_entropy": 0.47633055371514865,
+ "eval_loss": 0.6886085867881775,
+ "eval_mean_token_accuracy": 0.823554899585381,
+ "eval_num_tokens": 4584210.0,
+ "eval_runtime": 89.8185,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 1980
+ },
+ {
+ "entropy": 0.4718513362109661,
+ "epoch": 4.820760410380205,
+ "grad_norm": 0.7674988508224487,
+ "learning_rate": 0.00022543569729002318,
+ "loss": 0.44841318130493163,
+ "mean_token_accuracy": 0.8566767282783985,
+ "num_tokens": 4630254.0,
+ "step": 2000
+ },
+ {
+ "epoch": 4.820760410380205,
+ "eval_entropy": 0.49708309284086977,
+ "eval_loss": 0.6878687739372253,
+ "eval_mean_token_accuracy": 0.8217844698536262,
+ "eval_num_tokens": 4630254.0,
+ "eval_runtime": 89.7497,
+ "eval_samples_per_second": 15.822,
+ "eval_steps_per_second": 1.983,
+ "step": 2000
+ },
+ {
+ "entropy": 0.44017268233001233,
+ "epoch": 4.869040434520217,
+ "grad_norm": 0.6380876898765564,
+ "learning_rate": 0.00022244878952213976,
+ "loss": 0.42589750289916994,
+ "mean_token_accuracy": 0.8615814067423344,
+ "num_tokens": 4679265.0,
+ "step": 2020
+ },
+ {
+ "epoch": 4.869040434520217,
+ "eval_entropy": 0.5042208635740066,
+ "eval_loss": 0.6718228459358215,
+ "eval_mean_token_accuracy": 0.8232053056191863,
+ "eval_num_tokens": 4679265.0,
+ "eval_runtime": 89.7331,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 2020
+ },
+ {
+ "entropy": 0.4654298175126314,
+ "epoch": 4.917320458660229,
+ "grad_norm": 0.8285323977470398,
+ "learning_rate": 0.0002194505374203766,
+ "loss": 0.44148526191711424,
+ "mean_token_accuracy": 0.8545936144888401,
+ "num_tokens": 4723592.0,
+ "step": 2040
+ },
+ {
+ "epoch": 4.917320458660229,
+ "eval_entropy": 0.4957777807551823,
+ "eval_loss": 0.6694385409355164,
+ "eval_mean_token_accuracy": 0.8243840380331103,
+ "eval_num_tokens": 4723592.0,
+ "eval_runtime": 89.5293,
+ "eval_samples_per_second": 15.861,
+ "eval_steps_per_second": 1.988,
+ "step": 2040
+ },
+ {
+ "entropy": 0.46651641055941584,
+ "epoch": 4.965600482800241,
+ "grad_norm": 0.946260929107666,
+ "learning_rate": 0.00021644178945461268,
+ "loss": 0.43698992729187014,
+ "mean_token_accuracy": 0.8586322009563446,
+ "num_tokens": 4767253.0,
+ "step": 2060
+ },
+ {
+ "epoch": 4.965600482800241,
+ "eval_entropy": 0.5063281496254246,
+ "eval_loss": 0.6729032397270203,
+ "eval_mean_token_accuracy": 0.824299869912394,
+ "eval_num_tokens": 4767253.0,
+ "eval_runtime": 89.6207,
+ "eval_samples_per_second": 15.845,
+ "eval_steps_per_second": 1.986,
+ "step": 2060
+ },
+ {
+ "entropy": 0.4135968596130222,
+ "epoch": 5.012070006035003,
+ "grad_norm": 0.6853868961334229,
+ "learning_rate": 0.0002134233970649322,
+ "loss": 0.3800280332565308,
+ "mean_token_accuracy": 0.8756770677380747,
+ "num_tokens": 4816587.0,
+ "step": 2080
+ },
+ {
+ "epoch": 5.012070006035003,
+ "eval_entropy": 0.4107752423942759,
+ "eval_loss": 0.7204955220222473,
+ "eval_mean_token_accuracy": 0.8226918417416261,
+ "eval_num_tokens": 4816587.0,
+ "eval_runtime": 89.5817,
+ "eval_samples_per_second": 15.851,
+ "eval_steps_per_second": 1.987,
+ "step": 2080
+ },
+ {
+ "entropy": 0.316746598854661,
+ "epoch": 5.060350030175015,
+ "grad_norm": 0.6939145922660828,
+ "learning_rate": 0.00021039621442067732,
+ "loss": 0.2942554473876953,
+ "mean_token_accuracy": 0.90015934035182,
+ "num_tokens": 4861071.0,
+ "step": 2100
+ },
+ {
+ "epoch": 5.060350030175015,
+ "eval_entropy": 0.4214997919422857,
+ "eval_loss": 0.7268305420875549,
+ "eval_mean_token_accuracy": 0.8216965369294199,
+ "eval_num_tokens": 4861071.0,
+ "eval_runtime": 89.8077,
+ "eval_samples_per_second": 15.812,
+ "eval_steps_per_second": 1.982,
+ "step": 2100
+ },
+ {
+ "entropy": 0.3187096672132611,
+ "epoch": 5.108630054315027,
+ "grad_norm": 0.9134758114814758,
+ "learning_rate": 0.00020736109817872823,
+ "loss": 0.2839747190475464,
+ "mean_token_accuracy": 0.9006688073277473,
+ "num_tokens": 4904742.0,
+ "step": 2120
+ },
+ {
+ "epoch": 5.108630054315027,
+ "eval_entropy": 0.4375539443800958,
+ "eval_loss": 0.7280111908912659,
+ "eval_mean_token_accuracy": 0.8186122694712007,
+ "eval_num_tokens": 4904742.0,
+ "eval_runtime": 89.5868,
+ "eval_samples_per_second": 15.851,
+ "eval_steps_per_second": 1.987,
+ "step": 2120
+ },
+ {
+ "entropy": 0.3137735094875097,
+ "epoch": 5.15691007845504,
+ "grad_norm": 0.7324768304824829,
+ "learning_rate": 0.00020431890724107943,
+ "loss": 0.28816254138946534,
+ "mean_token_accuracy": 0.9023615352809429,
+ "num_tokens": 4950795.0,
+ "step": 2140
+ },
+ {
+ "epoch": 5.15691007845504,
+ "eval_entropy": 0.42255786312430094,
+ "eval_loss": 0.7335684895515442,
+ "eval_mean_token_accuracy": 0.8213735473959634,
+ "eval_num_tokens": 4950795.0,
+ "eval_runtime": 89.5711,
+ "eval_samples_per_second": 15.853,
+ "eval_steps_per_second": 1.987,
+ "step": 2140
+ },
+ {
+ "entropy": 0.3310150146484375,
+ "epoch": 5.2051901025950515,
+ "grad_norm": 0.7819830179214478,
+ "learning_rate": 0.00020127050251178062,
+ "loss": 0.3020722150802612,
+ "mean_token_accuracy": 0.8957489147782326,
+ "num_tokens": 4994886.0,
+ "step": 2160
+ },
+ {
+ "epoch": 5.2051901025950515,
+ "eval_entropy": 0.43374256186940696,
+ "eval_loss": 0.7270973920822144,
+ "eval_mean_token_accuracy": 0.821890847066815,
+ "eval_num_tokens": 4994886.0,
+ "eval_runtime": 89.7117,
+ "eval_samples_per_second": 15.828,
+ "eval_steps_per_second": 1.984,
+ "step": 2160
+ },
+ {
+ "entropy": 0.31917148139327767,
+ "epoch": 5.253470126735063,
+ "grad_norm": 0.8766310811042786,
+ "learning_rate": 0.00019821674665331112,
+ "loss": 0.29281470775604246,
+ "mean_token_accuracy": 0.8986819669604301,
+ "num_tokens": 5043663.0,
+ "step": 2180
+ },
+ {
+ "epoch": 5.253470126735063,
+ "eval_entropy": 0.4188300051381079,
+ "eval_loss": 0.7247599959373474,
+ "eval_mean_token_accuracy": 0.8234477113471942,
+ "eval_num_tokens": 5043663.0,
+ "eval_runtime": 89.384,
+ "eval_samples_per_second": 15.887,
+ "eval_steps_per_second": 1.991,
+ "step": 2180
+ },
+ {
+ "entropy": 0.3252310147508979,
+ "epoch": 5.301750150875075,
+ "grad_norm": 0.8721100687980652,
+ "learning_rate": 0.0001951585038424563,
+ "loss": 0.2924081325531006,
+ "mean_token_accuracy": 0.8981824897229671,
+ "num_tokens": 5087959.0,
+ "step": 2200
+ },
+ {
+ "epoch": 5.301750150875075,
+ "eval_entropy": 0.4114788512835342,
+ "eval_loss": 0.735135555267334,
+ "eval_mean_token_accuracy": 0.8240512961082245,
+ "eval_num_tokens": 5087959.0,
+ "eval_runtime": 89.3499,
+ "eval_samples_per_second": 15.893,
+ "eval_steps_per_second": 1.992,
+ "step": 2200
+ },
+ {
+ "entropy": 0.3175549603998661,
+ "epoch": 5.350030175015087,
+ "grad_norm": 0.8240578174591064,
+ "learning_rate": 0.00019209663952575616,
+ "loss": 0.2883902072906494,
+ "mean_token_accuracy": 0.9001291915774345,
+ "num_tokens": 5136121.0,
+ "step": 2220
+ },
+ {
+ "epoch": 5.350030175015087,
+ "eval_entropy": 0.417589228474692,
+ "eval_loss": 0.7371609807014465,
+ "eval_mean_token_accuracy": 0.8225449795803327,
+ "eval_num_tokens": 5136121.0,
+ "eval_runtime": 89.1634,
+ "eval_samples_per_second": 15.926,
+ "eval_steps_per_second": 1.996,
+ "step": 2220
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.218701419042181e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2240/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2240/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2240/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2240/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2240/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2240/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2240/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2240/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2240/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '\n' }}
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
+ {{- args_value }}
+ {{- '\n\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '\n' }}
+ {%- endfor %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+ {%- elif message.role == "tool" %}
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
+ {{- '<|im_start|>user' }}
+ {%- endif %}
+ {{- '\n\n' }}
+ {{- content }}
+ {{- '\n' }}
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
+ {{- '<|im_end|>\n' }}
+ {%- elif loop.last %}
+ {{- '<|im_end|>\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- raise_exception('Unexpected message role.') }}
+ {%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+ {{- '<|im_start|>assistant\n' }}
+ {%- if enable_thinking is defined and enable_thinking is false %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n' }}
+ {%- endif %}
+{%- endif %}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2240/tokenizer_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2240/tokenizer_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..b4a37b2a6fd3ab3317cd7bac72855be1a843b2bb
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2240/tokenizer_config.json
@@ -0,0 +1,31 @@
+{
+ "add_prefix_space": false,
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "backend": "tokenizers",
+ "bos_token": null,
+ "clean_up_tokenization_spaces": false,
+ "eos_token": "<|endoftext|>",
+ "errors": "replace",
+ "image_token": "<|image_pad|>",
+ "is_local": false,
+ "model_max_length": 262144,
+ "model_specific_special_tokens": {
+ "audio_bos_token": "<|audio_start|>",
+ "audio_eos_token": "<|audio_end|>",
+ "audio_token": "<|audio_pad|>",
+ "image_token": "<|image_pad|>",
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+ },
+ "pad_token": "<|endoftext|>",
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
+ "split_special_tokens": false,
+ "tokenizer_class": "TokenizersBackend",
+ "unk_token": null,
+ "video_token": "<|video_pad|>",
+ "vision_bos_token": "<|vision_start|>",
+ "vision_eos_token": "<|vision_end|>"
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2240/trainer_state.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2240/trainer_state.json
new file mode 100644
index 0000000000000000000000000000000000000000..15138fc1d6df8fa6ca798fe536222ea5422c9740
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2240/trainer_state.json
@@ -0,0 +1,2386 @@
+{
+ "best_global_step": null,
+ "best_metric": null,
+ "best_model_checkpoint": null,
+ "epoch": 5.3983101991551,
+ "eval_steps": 20,
+ "global_step": 2240,
+ "is_hyper_param_search": false,
+ "is_local_process_zero": true,
+ "is_world_process_zero": true,
+ "log_history": [
+ {
+ "entropy": 2.024795526266098,
+ "epoch": 0.04828002414001207,
+ "grad_norm": 2.945375919342041,
+ "learning_rate": 1.6698127428345936e-05,
+ "loss": 1.7825420379638672,
+ "mean_token_accuracy": 0.6315708436071873,
+ "num_tokens": 47778.0,
+ "step": 20
+ },
+ {
+ "epoch": 0.04828002414001207,
+ "eval_entropy": 1.29921497588747,
+ "eval_loss": 1.169131875038147,
+ "eval_mean_token_accuracy": 0.7185706342204233,
+ "eval_num_tokens": 47778.0,
+ "eval_runtime": 91.7898,
+ "eval_samples_per_second": 15.47,
+ "eval_steps_per_second": 1.939,
+ "step": 20
+ },
+ {
+ "entropy": 1.0346344470977784,
+ "epoch": 0.09656004828002414,
+ "grad_norm": 1.7916979789733887,
+ "learning_rate": 3.427510366871008e-05,
+ "loss": 0.9480395317077637,
+ "mean_token_accuracy": 0.7543429024517536,
+ "num_tokens": 95981.0,
+ "step": 40
+ },
+ {
+ "epoch": 0.09656004828002414,
+ "eval_entropy": 0.9673695430326997,
+ "eval_loss": 0.8588430285453796,
+ "eval_mean_token_accuracy": 0.7737233002534073,
+ "eval_num_tokens": 95981.0,
+ "eval_runtime": 89.8219,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 40
+ },
+ {
+ "entropy": 0.8977566614747048,
+ "epoch": 0.14484007242003621,
+ "grad_norm": 1.7972187995910645,
+ "learning_rate": 5.185207990907422e-05,
+ "loss": 0.8029604911804199,
+ "mean_token_accuracy": 0.7818016350269318,
+ "num_tokens": 142784.0,
+ "step": 60
+ },
+ {
+ "epoch": 0.14484007242003621,
+ "eval_entropy": 0.8436836982041263,
+ "eval_loss": 0.7893033623695374,
+ "eval_mean_token_accuracy": 0.7858914984076211,
+ "eval_num_tokens": 142784.0,
+ "eval_runtime": 89.9349,
+ "eval_samples_per_second": 15.789,
+ "eval_steps_per_second": 1.979,
+ "step": 60
+ },
+ {
+ "entropy": 0.8650321319699288,
+ "epoch": 0.19312009656004828,
+ "grad_norm": 1.2128729820251465,
+ "learning_rate": 6.942905614943836e-05,
+ "loss": 0.77417893409729,
+ "mean_token_accuracy": 0.7875479467213153,
+ "num_tokens": 187338.0,
+ "step": 80
+ },
+ {
+ "epoch": 0.19312009656004828,
+ "eval_entropy": 0.8209857180547179,
+ "eval_loss": 0.7572280168533325,
+ "eval_mean_token_accuracy": 0.7925970383574453,
+ "eval_num_tokens": 187338.0,
+ "eval_runtime": 89.9459,
+ "eval_samples_per_second": 15.787,
+ "eval_steps_per_second": 1.979,
+ "step": 80
+ },
+ {
+ "entropy": 0.8096660938113928,
+ "epoch": 0.24140012070006034,
+ "grad_norm": 1.1254287958145142,
+ "learning_rate": 8.700603238980251e-05,
+ "loss": 0.7170331001281738,
+ "mean_token_accuracy": 0.7994555257260799,
+ "num_tokens": 236038.0,
+ "step": 100
+ },
+ {
+ "epoch": 0.24140012070006034,
+ "eval_entropy": 0.792780208788561,
+ "eval_loss": 0.7442984580993652,
+ "eval_mean_token_accuracy": 0.7954704573984896,
+ "eval_num_tokens": 236038.0,
+ "eval_runtime": 90.294,
+ "eval_samples_per_second": 15.726,
+ "eval_steps_per_second": 1.971,
+ "step": 100
+ },
+ {
+ "entropy": 0.7979829199612141,
+ "epoch": 0.28968014484007243,
+ "grad_norm": 1.2140507698059082,
+ "learning_rate": 0.00010458300863016666,
+ "loss": 0.7141861915588379,
+ "mean_token_accuracy": 0.8000014744699001,
+ "num_tokens": 282918.0,
+ "step": 120
+ },
+ {
+ "epoch": 0.28968014484007243,
+ "eval_entropy": 0.7938692486018277,
+ "eval_loss": 0.7279797196388245,
+ "eval_mean_token_accuracy": 0.7989798734027348,
+ "eval_num_tokens": 282918.0,
+ "eval_runtime": 89.7189,
+ "eval_samples_per_second": 15.827,
+ "eval_steps_per_second": 1.984,
+ "step": 120
+ },
+ {
+ "entropy": 0.7916347607970238,
+ "epoch": 0.33796016898008446,
+ "grad_norm": 1.0506458282470703,
+ "learning_rate": 0.0001221599848705308,
+ "loss": 0.7117055416107178,
+ "mean_token_accuracy": 0.7970046654343605,
+ "num_tokens": 331305.0,
+ "step": 140
+ },
+ {
+ "epoch": 0.33796016898008446,
+ "eval_entropy": 0.8052781063519167,
+ "eval_loss": 0.722576916217804,
+ "eval_mean_token_accuracy": 0.7982530523551984,
+ "eval_num_tokens": 331305.0,
+ "eval_runtime": 90.0376,
+ "eval_samples_per_second": 15.771,
+ "eval_steps_per_second": 1.977,
+ "step": 140
+ },
+ {
+ "entropy": 0.7880666613578796,
+ "epoch": 0.38624019312009655,
+ "grad_norm": 0.8409207463264465,
+ "learning_rate": 0.00013973696111089492,
+ "loss": 0.7052523136138916,
+ "mean_token_accuracy": 0.8000123649835587,
+ "num_tokens": 375876.0,
+ "step": 160
+ },
+ {
+ "epoch": 0.38624019312009655,
+ "eval_entropy": 0.7787443134891853,
+ "eval_loss": 0.7137772440910339,
+ "eval_mean_token_accuracy": 0.8014209287220173,
+ "eval_num_tokens": 375876.0,
+ "eval_runtime": 90.2562,
+ "eval_samples_per_second": 15.733,
+ "eval_steps_per_second": 1.972,
+ "step": 160
+ },
+ {
+ "entropy": 0.7696007996797561,
+ "epoch": 0.43452021726010864,
+ "grad_norm": 1.0182477235794067,
+ "learning_rate": 0.00015731393735125907,
+ "loss": 0.7008797645568847,
+ "mean_token_accuracy": 0.8007259473204613,
+ "num_tokens": 423913.0,
+ "step": 180
+ },
+ {
+ "epoch": 0.43452021726010864,
+ "eval_entropy": 0.752259682068664,
+ "eval_loss": 0.7171286344528198,
+ "eval_mean_token_accuracy": 0.80052122574174,
+ "eval_num_tokens": 423913.0,
+ "eval_runtime": 90.0797,
+ "eval_samples_per_second": 15.764,
+ "eval_steps_per_second": 1.976,
+ "step": 180
+ },
+ {
+ "entropy": 0.7828126326203346,
+ "epoch": 0.4828002414001207,
+ "grad_norm": 0.9553632736206055,
+ "learning_rate": 0.0001748909135916232,
+ "loss": 0.7046853542327881,
+ "mean_token_accuracy": 0.7988562889397144,
+ "num_tokens": 472395.0,
+ "step": 200
+ },
+ {
+ "epoch": 0.4828002414001207,
+ "eval_entropy": 0.7389638169427936,
+ "eval_loss": 0.7187935709953308,
+ "eval_mean_token_accuracy": 0.7995899999409579,
+ "eval_num_tokens": 472395.0,
+ "eval_runtime": 89.9223,
+ "eval_samples_per_second": 15.791,
+ "eval_steps_per_second": 1.979,
+ "step": 200
+ },
+ {
+ "entropy": 0.7909286297857762,
+ "epoch": 0.5310802655401328,
+ "grad_norm": 1.134459376335144,
+ "learning_rate": 0.00019246788983198735,
+ "loss": 0.6989312171936035,
+ "mean_token_accuracy": 0.7991565138101577,
+ "num_tokens": 517049.0,
+ "step": 220
+ },
+ {
+ "epoch": 0.5310802655401328,
+ "eval_entropy": 0.7496965199374082,
+ "eval_loss": 0.7106321454048157,
+ "eval_mean_token_accuracy": 0.8046183442131857,
+ "eval_num_tokens": 517049.0,
+ "eval_runtime": 89.6899,
+ "eval_samples_per_second": 15.832,
+ "eval_steps_per_second": 1.985,
+ "step": 220
+ },
+ {
+ "entropy": 0.7651963323354721,
+ "epoch": 0.5793602896801449,
+ "grad_norm": 0.9557585716247559,
+ "learning_rate": 0.0002100448660723515,
+ "loss": 0.6984559059143066,
+ "mean_token_accuracy": 0.8019823372364044,
+ "num_tokens": 562836.0,
+ "step": 240
+ },
+ {
+ "epoch": 0.5793602896801449,
+ "eval_entropy": 0.7615563079212488,
+ "eval_loss": 0.7098539471626282,
+ "eval_mean_token_accuracy": 0.8047233508544022,
+ "eval_num_tokens": 562836.0,
+ "eval_runtime": 90.3752,
+ "eval_samples_per_second": 15.712,
+ "eval_steps_per_second": 1.97,
+ "step": 240
+ },
+ {
+ "entropy": 0.787507726252079,
+ "epoch": 0.627640313820157,
+ "grad_norm": 0.9279311299324036,
+ "learning_rate": 0.00022762184231271568,
+ "loss": 0.708138370513916,
+ "mean_token_accuracy": 0.8013308703899383,
+ "num_tokens": 606612.0,
+ "step": 260
+ },
+ {
+ "epoch": 0.627640313820157,
+ "eval_entropy": 0.7856640896100676,
+ "eval_loss": 0.7191570401191711,
+ "eval_mean_token_accuracy": 0.8024677249153008,
+ "eval_num_tokens": 606612.0,
+ "eval_runtime": 93.7336,
+ "eval_samples_per_second": 15.149,
+ "eval_steps_per_second": 1.899,
+ "step": 260
+ },
+ {
+ "entropy": 0.768489234894514,
+ "epoch": 0.6759203379601689,
+ "grad_norm": 0.7992776036262512,
+ "learning_rate": 0.0002451988185530798,
+ "loss": 0.7035849571228028,
+ "mean_token_accuracy": 0.8031105332076549,
+ "num_tokens": 655920.0,
+ "step": 280
+ },
+ {
+ "epoch": 0.6759203379601689,
+ "eval_entropy": 0.7682498740346244,
+ "eval_loss": 0.7149022221565247,
+ "eval_mean_token_accuracy": 0.7973497995499814,
+ "eval_num_tokens": 655920.0,
+ "eval_runtime": 89.7309,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 280
+ },
+ {
+ "entropy": 0.7720955617725849,
+ "epoch": 0.724200362100181,
+ "grad_norm": 1.1931514739990234,
+ "learning_rate": 0.0002627757947934439,
+ "loss": 0.6965402603149414,
+ "mean_token_accuracy": 0.8040348663926125,
+ "num_tokens": 702638.0,
+ "step": 300
+ },
+ {
+ "epoch": 0.724200362100181,
+ "eval_entropy": 0.7214708304807042,
+ "eval_loss": 0.7350823879241943,
+ "eval_mean_token_accuracy": 0.7942911925610532,
+ "eval_num_tokens": 702638.0,
+ "eval_runtime": 89.6451,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 300
+ },
+ {
+ "entropy": 0.7930095426738262,
+ "epoch": 0.7724803862401931,
+ "grad_norm": 0.8994467258453369,
+ "learning_rate": 0.0002803527710338081,
+ "loss": 0.7331492900848389,
+ "mean_token_accuracy": 0.796185302734375,
+ "num_tokens": 749377.0,
+ "step": 320
+ },
+ {
+ "epoch": 0.7724803862401931,
+ "eval_entropy": 0.8142535488926963,
+ "eval_loss": 0.7237834334373474,
+ "eval_mean_token_accuracy": 0.799495273761535,
+ "eval_num_tokens": 749377.0,
+ "eval_runtime": 90.6759,
+ "eval_samples_per_second": 15.66,
+ "eval_steps_per_second": 1.963,
+ "step": 320
+ },
+ {
+ "entropy": 0.7825102441012859,
+ "epoch": 0.8207604103802052,
+ "grad_norm": 1.8224239349365234,
+ "learning_rate": 0.00029792974727417223,
+ "loss": 0.7331175804138184,
+ "mean_token_accuracy": 0.7977667152881622,
+ "num_tokens": 798874.0,
+ "step": 340
+ },
+ {
+ "epoch": 0.8207604103802052,
+ "eval_entropy": 0.7327135443017724,
+ "eval_loss": 0.7388784289360046,
+ "eval_mean_token_accuracy": 0.8007041252730938,
+ "eval_num_tokens": 798874.0,
+ "eval_runtime": 89.7692,
+ "eval_samples_per_second": 15.818,
+ "eval_steps_per_second": 1.983,
+ "step": 340
+ },
+ {
+ "entropy": 0.819331557303667,
+ "epoch": 0.8690404345202173,
+ "grad_norm": 1.0407779216766357,
+ "learning_rate": 0.0003155067235145364,
+ "loss": 0.7456695556640625,
+ "mean_token_accuracy": 0.7936980701982975,
+ "num_tokens": 840825.0,
+ "step": 360
+ },
+ {
+ "epoch": 0.8690404345202173,
+ "eval_entropy": 0.8065585420372781,
+ "eval_loss": 0.735519528388977,
+ "eval_mean_token_accuracy": 0.7974206296245704,
+ "eval_num_tokens": 840825.0,
+ "eval_runtime": 89.4336,
+ "eval_samples_per_second": 15.878,
+ "eval_steps_per_second": 1.99,
+ "step": 360
+ },
+ {
+ "entropy": 0.8114653021097183,
+ "epoch": 0.9173204586602294,
+ "grad_norm": 1.1209490299224854,
+ "learning_rate": 0.0003330836997549005,
+ "loss": 0.7340899467468261,
+ "mean_token_accuracy": 0.7914272703230381,
+ "num_tokens": 882885.0,
+ "step": 380
+ },
+ {
+ "epoch": 0.9173204586602294,
+ "eval_entropy": 0.8025527285056168,
+ "eval_loss": 0.7295576930046082,
+ "eval_mean_token_accuracy": 0.7990576628218876,
+ "eval_num_tokens": 882885.0,
+ "eval_runtime": 89.9579,
+ "eval_samples_per_second": 15.785,
+ "eval_steps_per_second": 1.979,
+ "step": 380
+ },
+ {
+ "entropy": 0.8042497783899307,
+ "epoch": 0.9656004828002414,
+ "grad_norm": 1.2437798976898193,
+ "learning_rate": 0.0003506606759952647,
+ "loss": 0.7381744384765625,
+ "mean_token_accuracy": 0.7967362694442273,
+ "num_tokens": 927456.0,
+ "step": 400
+ },
+ {
+ "epoch": 0.9656004828002414,
+ "eval_entropy": 0.7717011336530193,
+ "eval_loss": 0.7436173558235168,
+ "eval_mean_token_accuracy": 0.7975849482450592,
+ "eval_num_tokens": 927456.0,
+ "eval_runtime": 89.8411,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 400
+ },
+ {
+ "entropy": 0.7987750010056929,
+ "epoch": 1.012070006035003,
+ "grad_norm": 1.2390918731689453,
+ "learning_rate": 0.000364721224843344,
+ "loss": 0.7317790031433106,
+ "mean_token_accuracy": 0.7986087551364651,
+ "num_tokens": 972662.0,
+ "step": 420
+ },
+ {
+ "epoch": 1.012070006035003,
+ "eval_entropy": 0.7333170604170038,
+ "eval_loss": 0.7456948757171631,
+ "eval_mean_token_accuracy": 0.7964897577682238,
+ "eval_num_tokens": 972662.0,
+ "eval_runtime": 89.7288,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 420
+ },
+ {
+ "entropy": 0.7372976846992969,
+ "epoch": 1.060350030175015,
+ "grad_norm": 1.1552740335464478,
+ "learning_rate": 0.00036468510102269575,
+ "loss": 0.6734249114990234,
+ "mean_token_accuracy": 0.805834824591875,
+ "num_tokens": 1022525.0,
+ "step": 440
+ },
+ {
+ "epoch": 1.060350030175015,
+ "eval_entropy": 0.7130355526891987,
+ "eval_loss": 0.7555333375930786,
+ "eval_mean_token_accuracy": 0.7961652781186479,
+ "eval_num_tokens": 1022525.0,
+ "eval_runtime": 89.8713,
+ "eval_samples_per_second": 15.8,
+ "eval_steps_per_second": 1.981,
+ "step": 440
+ },
+ {
+ "entropy": 0.7465668715536594,
+ "epoch": 1.1086300543150271,
+ "grad_norm": 1.1991469860076904,
+ "learning_rate": 0.0003645973816745066,
+ "loss": 0.6871879577636719,
+ "mean_token_accuracy": 0.8042450234293937,
+ "num_tokens": 1069194.0,
+ "step": 460
+ },
+ {
+ "epoch": 1.1086300543150271,
+ "eval_entropy": 0.7772979542110743,
+ "eval_loss": 0.7561826705932617,
+ "eval_mean_token_accuracy": 0.7953876427720102,
+ "eval_num_tokens": 1069194.0,
+ "eval_runtime": 89.8395,
+ "eval_samples_per_second": 15.806,
+ "eval_steps_per_second": 1.981,
+ "step": 460
+ },
+ {
+ "entropy": 0.7477855801582336,
+ "epoch": 1.1569100784550392,
+ "grad_norm": 1.325486660003662,
+ "learning_rate": 0.00036445809162231435,
+ "loss": 0.6836830139160156,
+ "mean_token_accuracy": 0.8021619468927383,
+ "num_tokens": 1117570.0,
+ "step": 480
+ },
+ {
+ "epoch": 1.1569100784550392,
+ "eval_entropy": 0.7098110676481483,
+ "eval_loss": 0.758007287979126,
+ "eval_mean_token_accuracy": 0.7909953189030122,
+ "eval_num_tokens": 1117570.0,
+ "eval_runtime": 89.6101,
+ "eval_samples_per_second": 15.846,
+ "eval_steps_per_second": 1.986,
+ "step": 480
+ },
+ {
+ "entropy": 0.7632691070437432,
+ "epoch": 1.2051901025950513,
+ "grad_norm": 1.0122262239456177,
+ "learning_rate": 0.0003642672702835562,
+ "loss": 0.7071213245391845,
+ "mean_token_accuracy": 0.8000680930912495,
+ "num_tokens": 1161717.0,
+ "step": 500
+ },
+ {
+ "epoch": 1.2051901025950513,
+ "eval_entropy": 0.7645382720432924,
+ "eval_loss": 0.7492606043815613,
+ "eval_mean_token_accuracy": 0.7976858676149604,
+ "eval_num_tokens": 1161717.0,
+ "eval_runtime": 89.6092,
+ "eval_samples_per_second": 15.847,
+ "eval_steps_per_second": 1.986,
+ "step": 500
+ },
+ {
+ "entropy": 0.7623726457357407,
+ "epoch": 1.2534701267350634,
+ "grad_norm": 1.367774486541748,
+ "learning_rate": 0.00036402497165841384,
+ "loss": 0.6965175151824952,
+ "mean_token_accuracy": 0.8016501650214195,
+ "num_tokens": 1209390.0,
+ "step": 520
+ },
+ {
+ "epoch": 1.2534701267350634,
+ "eval_entropy": 0.7868083715438843,
+ "eval_loss": 0.7481449842453003,
+ "eval_mean_token_accuracy": 0.799330582779445,
+ "eval_num_tokens": 1209390.0,
+ "eval_runtime": 89.8241,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 520
+ },
+ {
+ "entropy": 0.7535207405686378,
+ "epoch": 1.3017501508750755,
+ "grad_norm": 1.213614821434021,
+ "learning_rate": 0.00036373126431453207,
+ "loss": 0.7078310489654541,
+ "mean_token_accuracy": 0.7998907208442688,
+ "num_tokens": 1256611.0,
+ "step": 540
+ },
+ {
+ "epoch": 1.3017501508750755,
+ "eval_entropy": 0.7829091950748743,
+ "eval_loss": 0.7416072487831116,
+ "eval_mean_token_accuracy": 0.7969146250339036,
+ "eval_num_tokens": 1256611.0,
+ "eval_runtime": 89.4607,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 540
+ },
+ {
+ "entropy": 0.7527363449335098,
+ "epoch": 1.3500301750150876,
+ "grad_norm": 1.0929982662200928,
+ "learning_rate": 0.00036338623136761495,
+ "loss": 0.7058369636535644,
+ "mean_token_accuracy": 0.7997087433934211,
+ "num_tokens": 1302325.0,
+ "step": 560
+ },
+ {
+ "epoch": 1.3500301750150876,
+ "eval_entropy": 0.728419938114252,
+ "eval_loss": 0.744873046875,
+ "eval_mean_token_accuracy": 0.7966634296299366,
+ "eval_num_tokens": 1302325.0,
+ "eval_runtime": 89.6597,
+ "eval_samples_per_second": 15.838,
+ "eval_steps_per_second": 1.985,
+ "step": 560
+ },
+ {
+ "entropy": 0.7406099259853363,
+ "epoch": 1.3983101991550995,
+ "grad_norm": 1.409364104270935,
+ "learning_rate": 0.00036298997045790515,
+ "loss": 0.6902852058410645,
+ "mean_token_accuracy": 0.805386807769537,
+ "num_tokens": 1343844.0,
+ "step": 580
+ },
+ {
+ "epoch": 1.3983101991550995,
+ "eval_entropy": 0.6991632758231645,
+ "eval_loss": 0.7480319738388062,
+ "eval_mean_token_accuracy": 0.7970068006033308,
+ "eval_num_tokens": 1343844.0,
+ "eval_runtime": 89.8329,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 580
+ },
+ {
+ "entropy": 0.7459449704736472,
+ "epoch": 1.4465902232951118,
+ "grad_norm": 1.7053111791610718,
+ "learning_rate": 0.00036254259372255277,
+ "loss": 0.7044640064239502,
+ "mean_token_accuracy": 0.8005018755793571,
+ "num_tokens": 1389471.0,
+ "step": 600
+ },
+ {
+ "epoch": 1.4465902232951118,
+ "eval_entropy": 0.7454566660891758,
+ "eval_loss": 0.7551484704017639,
+ "eval_mean_token_accuracy": 0.7897093731365846,
+ "eval_num_tokens": 1389471.0,
+ "eval_runtime": 89.8313,
+ "eval_samples_per_second": 15.807,
+ "eval_steps_per_second": 1.981,
+ "step": 600
+ },
+ {
+ "entropy": 0.7621391348540782,
+ "epoch": 1.4948702474351236,
+ "grad_norm": 1.174443244934082,
+ "learning_rate": 0.000362044227763882,
+ "loss": 0.7111573696136475,
+ "mean_token_accuracy": 0.798752411454916,
+ "num_tokens": 1435975.0,
+ "step": 620
+ },
+ {
+ "epoch": 1.4948702474351236,
+ "eval_entropy": 0.7664387949397055,
+ "eval_loss": 0.7365804314613342,
+ "eval_mean_token_accuracy": 0.7963101200843126,
+ "eval_num_tokens": 1435975.0,
+ "eval_runtime": 89.9623,
+ "eval_samples_per_second": 15.784,
+ "eval_steps_per_second": 1.979,
+ "step": 620
+ },
+ {
+ "entropy": 0.7424944408237935,
+ "epoch": 1.5431502715751357,
+ "grad_norm": 1.2898184061050415,
+ "learning_rate": 0.000361495013613564,
+ "loss": 0.6913110733032226,
+ "mean_token_accuracy": 0.8009015507996082,
+ "num_tokens": 1482035.0,
+ "step": 640
+ },
+ {
+ "epoch": 1.5431502715751357,
+ "eval_entropy": 0.7255479529332579,
+ "eval_loss": 0.7345991730690002,
+ "eval_mean_token_accuracy": 0.8013592945056015,
+ "eval_num_tokens": 1482035.0,
+ "eval_runtime": 89.817,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 640
+ },
+ {
+ "entropy": 0.7303367637097835,
+ "epoch": 1.5914302957151478,
+ "grad_norm": 1.1310476064682007,
+ "learning_rate": 0.00036089510669270666,
+ "loss": 0.6875617027282714,
+ "mean_token_accuracy": 0.8035397171974182,
+ "num_tokens": 1528119.0,
+ "step": 660
+ },
+ {
+ "epoch": 1.5914302957151478,
+ "eval_entropy": 0.7101308248016271,
+ "eval_loss": 0.7394917011260986,
+ "eval_mean_token_accuracy": 0.7990792942850777,
+ "eval_num_tokens": 1528119.0,
+ "eval_runtime": 89.7431,
+ "eval_samples_per_second": 15.823,
+ "eval_steps_per_second": 1.983,
+ "step": 660
+ },
+ {
+ "entropy": 0.7387678384780884,
+ "epoch": 1.63971031985516,
+ "grad_norm": 0.9645389914512634,
+ "learning_rate": 0.0003602446767678725,
+ "loss": 0.6911204338073731,
+ "mean_token_accuracy": 0.8056333385407924,
+ "num_tokens": 1573661.0,
+ "step": 680
+ },
+ {
+ "epoch": 1.63971031985516,
+ "eval_entropy": 0.7281073737010527,
+ "eval_loss": 0.7376708984375,
+ "eval_mean_token_accuracy": 0.7997356415464637,
+ "eval_num_tokens": 1573661.0,
+ "eval_runtime": 90.0338,
+ "eval_samples_per_second": 15.772,
+ "eval_steps_per_second": 1.977,
+ "step": 680
+ },
+ {
+ "entropy": 0.7449554048478604,
+ "epoch": 1.687990343995172,
+ "grad_norm": 1.9069629907608032,
+ "learning_rate": 0.0003595439079030364,
+ "loss": 0.7042286396026611,
+ "mean_token_accuracy": 0.8018639378249646,
+ "num_tokens": 1621714.0,
+ "step": 700
+ },
+ {
+ "epoch": 1.687990343995172,
+ "eval_entropy": 0.7022633817088738,
+ "eval_loss": 0.7320939898490906,
+ "eval_mean_token_accuracy": 0.8006051541044471,
+ "eval_num_tokens": 1621714.0,
+ "eval_runtime": 89.8948,
+ "eval_samples_per_second": 15.796,
+ "eval_steps_per_second": 1.98,
+ "step": 700
+ },
+ {
+ "entropy": 0.723249051719904,
+ "epoch": 1.736270368135184,
+ "grad_norm": 0.8333877325057983,
+ "learning_rate": 0.00035879299840749777,
+ "loss": 0.7043714046478271,
+ "mean_token_accuracy": 0.8028345607221127,
+ "num_tokens": 1672118.0,
+ "step": 720
+ },
+ {
+ "epoch": 1.736270368135184,
+ "eval_entropy": 0.672436946898364,
+ "eval_loss": 0.713233470916748,
+ "eval_mean_token_accuracy": 0.8047760577684038,
+ "eval_num_tokens": 1672118.0,
+ "eval_runtime": 89.6394,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 720
+ },
+ {
+ "entropy": 0.730014206469059,
+ "epoch": 1.7845503922751962,
+ "grad_norm": 1.4580037593841553,
+ "learning_rate": 0.00035799216077976137,
+ "loss": 0.6985176563262939,
+ "mean_token_accuracy": 0.8015547141432762,
+ "num_tokens": 1720465.0,
+ "step": 740
+ },
+ {
+ "epoch": 1.7845503922751962,
+ "eval_entropy": 0.6908873795123582,
+ "eval_loss": 0.7198202013969421,
+ "eval_mean_token_accuracy": 0.8024131682481659,
+ "eval_num_tokens": 1720465.0,
+ "eval_runtime": 89.3771,
+ "eval_samples_per_second": 15.888,
+ "eval_steps_per_second": 1.992,
+ "step": 740
+ },
+ {
+ "entropy": 0.7219605796039105,
+ "epoch": 1.832830416415208,
+ "grad_norm": 1.51682710647583,
+ "learning_rate": 0.000357141621647403,
+ "loss": 0.6844424724578857,
+ "mean_token_accuracy": 0.8032813094556331,
+ "num_tokens": 1764440.0,
+ "step": 760
+ },
+ {
+ "epoch": 1.832830416415208,
+ "eval_entropy": 0.7113401109582922,
+ "eval_loss": 0.7195309400558472,
+ "eval_mean_token_accuracy": 0.8019187035185568,
+ "eval_num_tokens": 1764440.0,
+ "eval_runtime": 89.4658,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 760
+ },
+ {
+ "entropy": 0.729826743900776,
+ "epoch": 1.8811104405552204,
+ "grad_norm": 1.0838000774383545,
+ "learning_rate": 0.0003562416217029361,
+ "loss": 0.685858678817749,
+ "mean_token_accuracy": 0.8050567395985126,
+ "num_tokens": 1810828.0,
+ "step": 780
+ },
+ {
+ "epoch": 1.8811104405552204,
+ "eval_entropy": 0.649370267484965,
+ "eval_loss": 0.7177029252052307,
+ "eval_mean_token_accuracy": 0.8050174043419656,
+ "eval_num_tokens": 1810828.0,
+ "eval_runtime": 89.2287,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 780
+ },
+ {
+ "entropy": 0.6799364153295755,
+ "epoch": 1.9293904646952322,
+ "grad_norm": 1.1364924907684326,
+ "learning_rate": 0.00035529241563569903,
+ "loss": 0.6714428901672364,
+ "mean_token_accuracy": 0.8087341658771038,
+ "num_tokens": 1857600.0,
+ "step": 800
+ },
+ {
+ "epoch": 1.9293904646952322,
+ "eval_entropy": 0.6603799047095052,
+ "eval_loss": 0.7201809287071228,
+ "eval_mean_token_accuracy": 0.802863868769635,
+ "eval_num_tokens": 1857600.0,
+ "eval_runtime": 89.2718,
+ "eval_samples_per_second": 15.906,
+ "eval_steps_per_second": 1.994,
+ "step": 800
+ },
+ {
+ "entropy": 0.7078610375523567,
+ "epoch": 1.9776704888352445,
+ "grad_norm": 1.2191942930221558,
+ "learning_rate": 0.0003542942720597808,
+ "loss": 0.6996590614318847,
+ "mean_token_accuracy": 0.8049512051045895,
+ "num_tokens": 1901998.0,
+ "step": 820
+ },
+ {
+ "epoch": 1.9776704888352445,
+ "eval_entropy": 0.7240545264120852,
+ "eval_loss": 0.7301309108734131,
+ "eval_mean_token_accuracy": 0.8015973166133581,
+ "eval_num_tokens": 1901998.0,
+ "eval_runtime": 89.3438,
+ "eval_samples_per_second": 15.894,
+ "eval_steps_per_second": 1.992,
+ "step": 820
+ },
+ {
+ "entropy": 0.6557391556826505,
+ "epoch": 2.024140012070006,
+ "grad_norm": 1.36203134059906,
+ "learning_rate": 0.0003532474734380064,
+ "loss": 0.6327113628387451,
+ "mean_token_accuracy": 0.8142829847026181,
+ "num_tokens": 1947553.0,
+ "step": 840
+ },
+ {
+ "epoch": 2.024140012070006,
+ "eval_entropy": 0.6573808648613062,
+ "eval_loss": 0.7334365844726562,
+ "eval_mean_token_accuracy": 0.8043633893634496,
+ "eval_num_tokens": 1947553.0,
+ "eval_runtime": 89.3359,
+ "eval_samples_per_second": 15.895,
+ "eval_steps_per_second": 1.992,
+ "step": 840
+ },
+ {
+ "entropy": 0.6409128502011299,
+ "epoch": 2.0724200362100182,
+ "grad_norm": 1.5804955959320068,
+ "learning_rate": 0.0003521523160020035,
+ "loss": 0.6023559093475341,
+ "mean_token_accuracy": 0.8222251623868942,
+ "num_tokens": 1995234.0,
+ "step": 860
+ },
+ {
+ "epoch": 2.0724200362100182,
+ "eval_entropy": 0.6983234788594621,
+ "eval_loss": 0.7232135534286499,
+ "eval_mean_token_accuracy": 0.8010654268639811,
+ "eval_num_tokens": 1995234.0,
+ "eval_runtime": 89.2615,
+ "eval_samples_per_second": 15.908,
+ "eval_steps_per_second": 1.994,
+ "step": 860
+ },
+ {
+ "entropy": 0.6151149723678827,
+ "epoch": 2.12070006035003,
+ "grad_norm": 1.1975436210632324,
+ "learning_rate": 0.00035100910966837193,
+ "loss": 0.5803564548492431,
+ "mean_token_accuracy": 0.8256325736641884,
+ "num_tokens": 2040639.0,
+ "step": 880
+ },
+ {
+ "epoch": 2.12070006035003,
+ "eval_entropy": 0.6636555573243773,
+ "eval_loss": 0.7205662727355957,
+ "eval_mean_token_accuracy": 0.8036768369460374,
+ "eval_num_tokens": 2040639.0,
+ "eval_runtime": 89.2957,
+ "eval_samples_per_second": 15.902,
+ "eval_steps_per_second": 1.993,
+ "step": 880
+ },
+ {
+ "entropy": 0.6201727617532015,
+ "epoch": 2.1689800844900424,
+ "grad_norm": 0.9907131195068359,
+ "learning_rate": 0.0003498181779509813,
+ "loss": 0.6101872444152832,
+ "mean_token_accuracy": 0.8196512959897518,
+ "num_tokens": 2090604.0,
+ "step": 900
+ },
+ {
+ "epoch": 2.1689800844900424,
+ "eval_entropy": 0.6473788086617931,
+ "eval_loss": 0.7062397003173828,
+ "eval_mean_token_accuracy": 0.8078362184963869,
+ "eval_num_tokens": 2090604.0,
+ "eval_runtime": 88.6198,
+ "eval_samples_per_second": 16.024,
+ "eval_steps_per_second": 2.009,
+ "step": 900
+ },
+ {
+ "entropy": 0.6097332935780286,
+ "epoch": 2.2172601086300543,
+ "grad_norm": 0.9055633544921875,
+ "learning_rate": 0.00034857985786942026,
+ "loss": 0.5956094264984131,
+ "mean_token_accuracy": 0.8240575887262821,
+ "num_tokens": 2142246.0,
+ "step": 920
+ },
+ {
+ "epoch": 2.2172601086300543,
+ "eval_entropy": 0.612629868341296,
+ "eval_loss": 0.7106770277023315,
+ "eval_mean_token_accuracy": 0.805324277181304,
+ "eval_num_tokens": 2142246.0,
+ "eval_runtime": 89.4383,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 920
+ },
+ {
+ "entropy": 0.6182701248675585,
+ "epoch": 2.2655401327700666,
+ "grad_norm": 0.9513736367225647,
+ "learning_rate": 0.00034729449985362404,
+ "loss": 0.6019269466400147,
+ "mean_token_accuracy": 0.8187106013298034,
+ "num_tokens": 2183144.0,
+ "step": 940
+ },
+ {
+ "epoch": 2.2655401327700666,
+ "eval_entropy": 0.6437647456533453,
+ "eval_loss": 0.7173502445220947,
+ "eval_mean_token_accuracy": 0.8056439573175451,
+ "eval_num_tokens": 2183144.0,
+ "eval_runtime": 88.9568,
+ "eval_samples_per_second": 15.963,
+ "eval_steps_per_second": 2.001,
+ "step": 940
+ },
+ {
+ "entropy": 0.6508570514619351,
+ "epoch": 2.3138201569100785,
+ "grad_norm": 0.9606207013130188,
+ "learning_rate": 0.0003459624676447067,
+ "loss": 0.6296409606933594,
+ "mean_token_accuracy": 0.8167732007801533,
+ "num_tokens": 2228098.0,
+ "step": 960
+ },
+ {
+ "epoch": 2.3138201569100785,
+ "eval_entropy": 0.6288247135248077,
+ "eval_loss": 0.7160675525665283,
+ "eval_mean_token_accuracy": 0.8066010689467527,
+ "eval_num_tokens": 2228098.0,
+ "eval_runtime": 89.0044,
+ "eval_samples_per_second": 15.954,
+ "eval_steps_per_second": 2.0,
+ "step": 960
+ },
+ {
+ "entropy": 0.6380081418901682,
+ "epoch": 2.3621001810500903,
+ "grad_norm": 0.9677181839942932,
+ "learning_rate": 0.000344584138192027,
+ "loss": 0.6210546493530273,
+ "mean_token_accuracy": 0.8186135642230511,
+ "num_tokens": 2273062.0,
+ "step": 980
+ },
+ {
+ "epoch": 2.3621001810500903,
+ "eval_entropy": 0.6407805031604981,
+ "eval_loss": 0.7055043578147888,
+ "eval_mean_token_accuracy": 0.8070103157772107,
+ "eval_num_tokens": 2273062.0,
+ "eval_runtime": 89.5606,
+ "eval_samples_per_second": 15.855,
+ "eval_steps_per_second": 1.987,
+ "step": 980
+ },
+ {
+ "entropy": 0.6310251638293266,
+ "epoch": 2.4103802051901027,
+ "grad_norm": 2.1046407222747803,
+ "learning_rate": 0.000343159901546516,
+ "loss": 0.6165830135345459,
+ "mean_token_accuracy": 0.8207855716347694,
+ "num_tokens": 2319305.0,
+ "step": 1000
+ },
+ {
+ "epoch": 2.4103802051901027,
+ "eval_entropy": 0.6461507470420237,
+ "eval_loss": 0.7292136549949646,
+ "eval_mean_token_accuracy": 0.8030418261383356,
+ "eval_num_tokens": 2319305.0,
+ "eval_runtime": 89.315,
+ "eval_samples_per_second": 15.899,
+ "eval_steps_per_second": 1.993,
+ "step": 1000
+ },
+ {
+ "entropy": 0.6256617344915867,
+ "epoch": 2.4586602293301145,
+ "grad_norm": 1.1768349409103394,
+ "learning_rate": 0.0003416901607502972,
+ "loss": 0.6142709255218506,
+ "mean_token_accuracy": 0.8220669947564602,
+ "num_tokens": 2368592.0,
+ "step": 1020
+ },
+ {
+ "epoch": 2.4586602293301145,
+ "eval_entropy": 0.6150523993406403,
+ "eval_loss": 0.7019563317298889,
+ "eval_mean_token_accuracy": 0.8094323099998946,
+ "eval_num_tokens": 2368592.0,
+ "eval_runtime": 89.419,
+ "eval_samples_per_second": 15.88,
+ "eval_steps_per_second": 1.991,
+ "step": 1020
+ },
+ {
+ "entropy": 0.6126735735684633,
+ "epoch": 2.506940253470127,
+ "grad_norm": 1.8065496683120728,
+ "learning_rate": 0.00034017533172263055,
+ "loss": 0.6056458473205566,
+ "mean_token_accuracy": 0.8203762136399746,
+ "num_tokens": 2412859.0,
+ "step": 1040
+ },
+ {
+ "epoch": 2.506940253470127,
+ "eval_entropy": 0.6544345456562685,
+ "eval_loss": 0.7086517214775085,
+ "eval_mean_token_accuracy": 0.806709805901131,
+ "eval_num_tokens": 2412859.0,
+ "eval_runtime": 88.981,
+ "eval_samples_per_second": 15.958,
+ "eval_steps_per_second": 2.0,
+ "step": 1040
+ },
+ {
+ "entropy": 0.6311754353344441,
+ "epoch": 2.5552202776101387,
+ "grad_norm": 1.0576369762420654,
+ "learning_rate": 0.00033861584314221236,
+ "loss": 0.6176392555236816,
+ "mean_token_accuracy": 0.8177410490810871,
+ "num_tokens": 2458201.0,
+ "step": 1060
+ },
+ {
+ "epoch": 2.5552202776101387,
+ "eval_entropy": 0.6273799200406235,
+ "eval_loss": 0.708163321018219,
+ "eval_mean_token_accuracy": 0.8032017099053672,
+ "eval_num_tokens": 2458201.0,
+ "eval_runtime": 89.7668,
+ "eval_samples_per_second": 15.819,
+ "eval_steps_per_second": 1.983,
+ "step": 1060
+ },
+ {
+ "entropy": 0.6200249589979648,
+ "epoch": 2.603500301750151,
+ "grad_norm": 0.997065007686615,
+ "learning_rate": 0.0003370121363258637,
+ "loss": 0.6203888893127442,
+ "mean_token_accuracy": 0.8196275025606156,
+ "num_tokens": 2505382.0,
+ "step": 1080
+ },
+ {
+ "epoch": 2.603500301750151,
+ "eval_entropy": 0.6457576811983344,
+ "eval_loss": 0.6972290277481079,
+ "eval_mean_token_accuracy": 0.8095980849158898,
+ "eval_num_tokens": 2505382.0,
+ "eval_runtime": 89.2027,
+ "eval_samples_per_second": 15.919,
+ "eval_steps_per_second": 1.995,
+ "step": 1080
+ },
+ {
+ "entropy": 0.626810473203659,
+ "epoch": 2.651780325890163,
+ "grad_norm": 0.865475058555603,
+ "learning_rate": 0.0003353646651036438,
+ "loss": 0.6102585792541504,
+ "mean_token_accuracy": 0.819708751142025,
+ "num_tokens": 2552927.0,
+ "step": 1100
+ },
+ {
+ "epoch": 2.651780325890163,
+ "eval_entropy": 0.6412022431914726,
+ "eval_loss": 0.7036139965057373,
+ "eval_mean_token_accuracy": 0.8070317135098275,
+ "eval_num_tokens": 2552927.0,
+ "eval_runtime": 89.4733,
+ "eval_samples_per_second": 15.871,
+ "eval_steps_per_second": 1.989,
+ "step": 1100
+ },
+ {
+ "entropy": 0.6581619590520859,
+ "epoch": 2.700060350030175,
+ "grad_norm": 1.00307035446167,
+ "learning_rate": 0.00033367389569042064,
+ "loss": 0.6397720336914062,
+ "mean_token_accuracy": 0.81413309648633,
+ "num_tokens": 2597426.0,
+ "step": 1120
+ },
+ {
+ "epoch": 2.700060350030175,
+ "eval_entropy": 0.6288736402318719,
+ "eval_loss": 0.6966825723648071,
+ "eval_mean_token_accuracy": 0.8094039669867312,
+ "eval_num_tokens": 2597426.0,
+ "eval_runtime": 88.7341,
+ "eval_samples_per_second": 16.003,
+ "eval_steps_per_second": 2.006,
+ "step": 1120
+ },
+ {
+ "entropy": 0.6399412982165813,
+ "epoch": 2.748340374170187,
+ "grad_norm": 0.8589329719543457,
+ "learning_rate": 0.0003319403065539384,
+ "loss": 0.6347338199615479,
+ "mean_token_accuracy": 0.8139534957706929,
+ "num_tokens": 2640514.0,
+ "step": 1140
+ },
+ {
+ "epoch": 2.748340374170187,
+ "eval_entropy": 0.683491913455256,
+ "eval_loss": 0.6984149217605591,
+ "eval_mean_token_accuracy": 0.8083989469522841,
+ "eval_num_tokens": 2640514.0,
+ "eval_runtime": 89.3929,
+ "eval_samples_per_second": 15.885,
+ "eval_steps_per_second": 1.991,
+ "step": 1140
+ },
+ {
+ "entropy": 0.6264835461974144,
+ "epoch": 2.796620398310199,
+ "grad_norm": 0.7807853817939758,
+ "learning_rate": 0.0003301643882794163,
+ "loss": 0.6162629127502441,
+ "mean_token_accuracy": 0.8216292977333068,
+ "num_tokens": 2687730.0,
+ "step": 1160
+ },
+ {
+ "epoch": 2.796620398310199,
+ "eval_entropy": 0.6578792159476977,
+ "eval_loss": 0.6921441555023193,
+ "eval_mean_token_accuracy": 0.8094299177775223,
+ "eval_num_tokens": 2687730.0,
+ "eval_runtime": 89.2356,
+ "eval_samples_per_second": 15.913,
+ "eval_steps_per_second": 1.995,
+ "step": 1160
+ },
+ {
+ "entropy": 0.6269129011780024,
+ "epoch": 2.8449004224502112,
+ "grad_norm": 1.1675593852996826,
+ "learning_rate": 0.0003283466434307189,
+ "loss": 0.6235532760620117,
+ "mean_token_accuracy": 0.8188897788524627,
+ "num_tokens": 2732930.0,
+ "step": 1180
+ },
+ {
+ "epoch": 2.8449004224502112,
+ "eval_entropy": 0.6533248447970058,
+ "eval_loss": 0.696849524974823,
+ "eval_mean_token_accuracy": 0.8102703713968898,
+ "eval_num_tokens": 2732930.0,
+ "eval_runtime": 88.2602,
+ "eval_samples_per_second": 16.089,
+ "eval_steps_per_second": 2.017,
+ "step": 1180
+ },
+ {
+ "entropy": 0.6259873781353236,
+ "epoch": 2.8931804465902236,
+ "grad_norm": 0.9596486687660217,
+ "learning_rate": 0.00032648758640813655,
+ "loss": 0.6118728160858155,
+ "mean_token_accuracy": 0.8171859130263328,
+ "num_tokens": 2779450.0,
+ "step": 1200
+ },
+ {
+ "epoch": 2.8931804465902236,
+ "eval_entropy": 0.6212078754821521,
+ "eval_loss": 0.6942870020866394,
+ "eval_mean_token_accuracy": 0.810484343365337,
+ "eval_num_tokens": 2779450.0,
+ "eval_runtime": 89.4375,
+ "eval_samples_per_second": 15.877,
+ "eval_steps_per_second": 1.99,
+ "step": 1200
+ },
+ {
+ "entropy": 0.6142537571489811,
+ "epoch": 2.9414604707302354,
+ "grad_norm": 0.9625058770179749,
+ "learning_rate": 0.00032458774330281615,
+ "loss": 0.6060261249542236,
+ "mean_token_accuracy": 0.8223113484680653,
+ "num_tokens": 2827819.0,
+ "step": 1220
+ },
+ {
+ "epoch": 2.9414604707302354,
+ "eval_entropy": 0.6553665667437436,
+ "eval_loss": 0.6835269331932068,
+ "eval_mean_token_accuracy": 0.8134723968720168,
+ "eval_num_tokens": 2827819.0,
+ "eval_runtime": 89.5092,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1220
+ },
+ {
+ "entropy": 0.6301972426474094,
+ "epoch": 2.9897404948702473,
+ "grad_norm": 0.9657333493232727,
+ "learning_rate": 0.0003226476517478835,
+ "loss": 0.6277759075164795,
+ "mean_token_accuracy": 0.8164266042411328,
+ "num_tokens": 2872493.0,
+ "step": 1240
+ },
+ {
+ "epoch": 2.9897404948702473,
+ "eval_entropy": 0.6594253118788258,
+ "eval_loss": 0.6820935606956482,
+ "eval_mean_token_accuracy": 0.8132368198941263,
+ "eval_num_tokens": 2872493.0,
+ "eval_runtime": 89.6211,
+ "eval_samples_per_second": 15.844,
+ "eval_steps_per_second": 1.986,
+ "step": 1240
+ },
+ {
+ "entropy": 0.5346100655469027,
+ "epoch": 3.036210018105009,
+ "grad_norm": 0.8541185259819031,
+ "learning_rate": 0.0003206678607662996,
+ "loss": 0.5261692047119141,
+ "mean_token_accuracy": 0.8382289425119177,
+ "num_tokens": 2918249.0,
+ "step": 1260
+ },
+ {
+ "epoch": 3.036210018105009,
+ "eval_entropy": 0.5626645662476507,
+ "eval_loss": 0.7028542160987854,
+ "eval_mean_token_accuracy": 0.8133395572056931,
+ "eval_num_tokens": 2918249.0,
+ "eval_runtime": 89.1267,
+ "eval_samples_per_second": 15.932,
+ "eval_steps_per_second": 1.997,
+ "step": 1260
+ },
+ {
+ "entropy": 0.5001101832836866,
+ "epoch": 3.084490042245021,
+ "grad_norm": 0.8208815455436707,
+ "learning_rate": 0.0003186489306154935,
+ "loss": 0.4850132942199707,
+ "mean_token_accuracy": 0.8468983843922615,
+ "num_tokens": 2968353.0,
+ "step": 1280
+ },
+ {
+ "epoch": 3.084490042245021,
+ "eval_entropy": 0.5889462046743779,
+ "eval_loss": 0.691015899181366,
+ "eval_mean_token_accuracy": 0.8135226466012805,
+ "eval_num_tokens": 2968353.0,
+ "eval_runtime": 88.9677,
+ "eval_samples_per_second": 15.961,
+ "eval_steps_per_second": 2.001,
+ "step": 1280
+ },
+ {
+ "entropy": 0.5242766339331866,
+ "epoch": 3.1327700663850333,
+ "grad_norm": 0.7893072962760925,
+ "learning_rate": 0.0003165914326288163,
+ "loss": 0.5061577320098877,
+ "mean_token_accuracy": 0.8425608821213245,
+ "num_tokens": 3018291.0,
+ "step": 1300
+ },
+ {
+ "epoch": 3.1327700663850333,
+ "eval_entropy": 0.5543691962957382,
+ "eval_loss": 0.6913734078407288,
+ "eval_mean_token_accuracy": 0.8154889680026622,
+ "eval_num_tokens": 3018291.0,
+ "eval_runtime": 89.6963,
+ "eval_samples_per_second": 15.831,
+ "eval_steps_per_second": 1.984,
+ "step": 1300
+ },
+ {
+ "entropy": 0.5327721010893584,
+ "epoch": 3.181050090525045,
+ "grad_norm": 0.8992080092430115,
+ "learning_rate": 0.0003144959490538604,
+ "loss": 0.5068036556243897,
+ "mean_token_accuracy": 0.8420542575418949,
+ "num_tokens": 3070712.0,
+ "step": 1320
+ },
+ {
+ "epoch": 3.181050090525045,
+ "eval_entropy": 0.5734540803378887,
+ "eval_loss": 0.6948716044425964,
+ "eval_mean_token_accuracy": 0.8137217944257715,
+ "eval_num_tokens": 3070712.0,
+ "eval_runtime": 89.4597,
+ "eval_samples_per_second": 15.873,
+ "eval_steps_per_second": 1.99,
+ "step": 1320
+ },
+ {
+ "entropy": 0.5247377116233111,
+ "epoch": 3.2293301146650575,
+ "grad_norm": 0.9291845560073853,
+ "learning_rate": 0.0003123630728876902,
+ "loss": 0.5023272037506104,
+ "mean_token_accuracy": 0.8435182586312294,
+ "num_tokens": 3116564.0,
+ "step": 1340
+ },
+ {
+ "epoch": 3.2293301146650575,
+ "eval_entropy": 0.5492362935891312,
+ "eval_loss": 0.6980633735656738,
+ "eval_mean_token_accuracy": 0.8155082919624415,
+ "eval_num_tokens": 3116564.0,
+ "eval_runtime": 89.2124,
+ "eval_samples_per_second": 15.917,
+ "eval_steps_per_second": 1.995,
+ "step": 1340
+ },
+ {
+ "entropy": 0.5380499072372913,
+ "epoch": 3.2776101388050694,
+ "grad_norm": 1.0477524995803833,
+ "learning_rate": 0.00031019340770903136,
+ "loss": 0.5264591693878173,
+ "mean_token_accuracy": 0.8391959741711617,
+ "num_tokens": 3164054.0,
+ "step": 1360
+ },
+ {
+ "epoch": 3.2776101388050694,
+ "eval_entropy": 0.5896503307511297,
+ "eval_loss": 0.6997144222259521,
+ "eval_mean_token_accuracy": 0.8124373062942805,
+ "eval_num_tokens": 3164054.0,
+ "eval_runtime": 89.1147,
+ "eval_samples_per_second": 15.935,
+ "eval_steps_per_second": 1.997,
+ "step": 1360
+ },
+ {
+ "entropy": 0.5460653610527515,
+ "epoch": 3.3258901629450817,
+ "grad_norm": 1.2473069429397583,
+ "learning_rate": 0.00030798756750746476,
+ "loss": 0.5319677829742432,
+ "mean_token_accuracy": 0.8364320226013661,
+ "num_tokens": 3207419.0,
+ "step": 1380
+ },
+ {
+ "epoch": 3.3258901629450817,
+ "eval_entropy": 0.5659052182114526,
+ "eval_loss": 0.6924049258232117,
+ "eval_mean_token_accuracy": 0.8153053585732921,
+ "eval_num_tokens": 3207419.0,
+ "eval_runtime": 89.6468,
+ "eval_samples_per_second": 15.84,
+ "eval_steps_per_second": 1.986,
+ "step": 1380
+ },
+ {
+ "entropy": 0.5367407951503992,
+ "epoch": 3.3741701870850935,
+ "grad_norm": 1.2679712772369385,
+ "learning_rate": 0.00030574617650967485,
+ "loss": 0.5190927028656006,
+ "mean_token_accuracy": 0.8399469174444676,
+ "num_tokens": 3252679.0,
+ "step": 1400
+ },
+ {
+ "epoch": 3.3741701870850935,
+ "eval_entropy": 0.5581759144081159,
+ "eval_loss": 0.6933317184448242,
+ "eval_mean_token_accuracy": 0.8139757862251796,
+ "eval_num_tokens": 3252679.0,
+ "eval_runtime": 89.0749,
+ "eval_samples_per_second": 15.942,
+ "eval_steps_per_second": 1.998,
+ "step": 1400
+ },
+ {
+ "entropy": 0.530736118927598,
+ "epoch": 3.4224502112251054,
+ "grad_norm": 0.9717612862586975,
+ "learning_rate": 0.0003034698690028009,
+ "loss": 0.512363052368164,
+ "mean_token_accuracy": 0.8419127956032753,
+ "num_tokens": 3297309.0,
+ "step": 1420
+ },
+ {
+ "epoch": 3.4224502112251054,
+ "eval_entropy": 0.5233335836549823,
+ "eval_loss": 0.6958550214767456,
+ "eval_mean_token_accuracy": 0.8158778501360604,
+ "eval_num_tokens": 3297309.0,
+ "eval_runtime": 88.8664,
+ "eval_samples_per_second": 15.979,
+ "eval_steps_per_second": 2.003,
+ "step": 1420
+ },
+ {
+ "entropy": 0.5327631575986743,
+ "epoch": 3.4707302353651177,
+ "grad_norm": 0.7649208903312683,
+ "learning_rate": 0.00030115928915494116,
+ "loss": 0.5202179431915284,
+ "mean_token_accuracy": 0.8384683802723885,
+ "num_tokens": 3341825.0,
+ "step": 1440
+ },
+ {
+ "epoch": 3.4707302353651177,
+ "eval_entropy": 0.5960972657364406,
+ "eval_loss": 0.6984783411026001,
+ "eval_mean_token_accuracy": 0.8095930597085631,
+ "eval_num_tokens": 3341825.0,
+ "eval_runtime": 89.5168,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1440
+ },
+ {
+ "entropy": 0.5428155034780502,
+ "epoch": 3.51901025950513,
+ "grad_norm": 1.2790874242782593,
+ "learning_rate": 0.00029881509083286116,
+ "loss": 0.524717378616333,
+ "mean_token_accuracy": 0.8346644505858422,
+ "num_tokens": 3388463.0,
+ "step": 1460
+ },
+ {
+ "epoch": 3.51901025950513,
+ "eval_entropy": 0.550318562917495,
+ "eval_loss": 0.6988471746444702,
+ "eval_mean_token_accuracy": 0.8132969897784544,
+ "eval_num_tokens": 3388463.0,
+ "eval_runtime": 89.2598,
+ "eval_samples_per_second": 15.909,
+ "eval_steps_per_second": 1.994,
+ "step": 1460
+ },
+ {
+ "entropy": 0.5489993002265692,
+ "epoch": 3.567290283645142,
+ "grad_norm": 1.055821418762207,
+ "learning_rate": 0.00029643793741695687,
+ "loss": 0.5421150207519532,
+ "mean_token_accuracy": 0.8345710933208466,
+ "num_tokens": 3433731.0,
+ "step": 1480
+ },
+ {
+ "epoch": 3.567290283645142,
+ "eval_entropy": 0.5816821718818685,
+ "eval_loss": 0.6853668689727783,
+ "eval_mean_token_accuracy": 0.8142332360985574,
+ "eval_num_tokens": 3433731.0,
+ "eval_runtime": 89.0424,
+ "eval_samples_per_second": 15.947,
+ "eval_steps_per_second": 1.999,
+ "step": 1480
+ },
+ {
+ "entropy": 0.5451365381479263,
+ "epoch": 3.6155703077851538,
+ "grad_norm": 0.9027532935142517,
+ "learning_rate": 0.00029402850161352583,
+ "loss": 0.5463609218597412,
+ "mean_token_accuracy": 0.8369908526539802,
+ "num_tokens": 3478600.0,
+ "step": 1500
+ },
+ {
+ "epoch": 3.6155703077851538,
+ "eval_entropy": 0.5344233209832331,
+ "eval_loss": 0.6794531345367432,
+ "eval_mean_token_accuracy": 0.8177759824843889,
+ "eval_num_tokens": 3478600.0,
+ "eval_runtime": 89.123,
+ "eval_samples_per_second": 15.933,
+ "eval_steps_per_second": 1.997,
+ "step": 1500
+ },
+ {
+ "entropy": 0.5521467242389917,
+ "epoch": 3.663850331925166,
+ "grad_norm": 1.4606311321258545,
+ "learning_rate": 0.0002915874652643997,
+ "loss": 0.5421902179718018,
+ "mean_token_accuracy": 0.8339455373585224,
+ "num_tokens": 3521787.0,
+ "step": 1520
+ },
+ {
+ "epoch": 3.663850331925166,
+ "eval_entropy": 0.5807709094513668,
+ "eval_loss": 0.6949691772460938,
+ "eval_mean_token_accuracy": 0.8121141439743256,
+ "eval_num_tokens": 3521787.0,
+ "eval_runtime": 88.5437,
+ "eval_samples_per_second": 16.037,
+ "eval_steps_per_second": 2.01,
+ "step": 1520
+ },
+ {
+ "entropy": 0.5455610822886229,
+ "epoch": 3.712130356065178,
+ "grad_norm": 0.7269843220710754,
+ "learning_rate": 0.0002891155191539905,
+ "loss": 0.5407682895660401,
+ "mean_token_accuracy": 0.8393648222088814,
+ "num_tokens": 3569036.0,
+ "step": 1540
+ },
+ {
+ "epoch": 3.712130356065178,
+ "eval_entropy": 0.6032205542151847,
+ "eval_loss": 0.6790974140167236,
+ "eval_mean_token_accuracy": 0.8165463112043531,
+ "eval_num_tokens": 3569036.0,
+ "eval_runtime": 89.517,
+ "eval_samples_per_second": 15.863,
+ "eval_steps_per_second": 1.988,
+ "step": 1540
+ },
+ {
+ "entropy": 0.5474163968116045,
+ "epoch": 3.7604103802051903,
+ "grad_norm": 0.8148866295814514,
+ "learning_rate": 0.00028661336281380717,
+ "loss": 0.5394441604614257,
+ "mean_token_accuracy": 0.8336028568446636,
+ "num_tokens": 3614313.0,
+ "step": 1560
+ },
+ {
+ "epoch": 3.7604103802051903,
+ "eval_entropy": 0.5797032238392348,
+ "eval_loss": 0.670870304107666,
+ "eval_mean_token_accuracy": 0.8182691593518417,
+ "eval_num_tokens": 3614313.0,
+ "eval_runtime": 89.4468,
+ "eval_samples_per_second": 15.875,
+ "eval_steps_per_second": 1.99,
+ "step": 1560
+ },
+ {
+ "entropy": 0.5511750865727663,
+ "epoch": 3.808690404345202,
+ "grad_norm": 1.3441656827926636,
+ "learning_rate": 0.0002840817043244964,
+ "loss": 0.5378933906555176,
+ "mean_token_accuracy": 0.8349151819944381,
+ "num_tokens": 3659786.0,
+ "step": 1580
+ },
+ {
+ "epoch": 3.808690404345202,
+ "eval_entropy": 0.5297858750217417,
+ "eval_loss": 0.680833637714386,
+ "eval_mean_token_accuracy": 0.8180924275617921,
+ "eval_num_tokens": 3659786.0,
+ "eval_runtime": 89.5417,
+ "eval_samples_per_second": 15.859,
+ "eval_steps_per_second": 1.988,
+ "step": 1580
+ },
+ {
+ "entropy": 0.5379613988101483,
+ "epoch": 3.856970428485214,
+ "grad_norm": 0.9813932776451111,
+ "learning_rate": 0.00028152126011546396,
+ "loss": 0.5340342044830322,
+ "mean_token_accuracy": 0.8346010789275169,
+ "num_tokens": 3706362.0,
+ "step": 1600
+ },
+ {
+ "epoch": 3.856970428485214,
+ "eval_entropy": 0.5811898510777549,
+ "eval_loss": 0.6764267683029175,
+ "eval_mean_token_accuracy": 0.8180528236239144,
+ "eval_num_tokens": 3706362.0,
+ "eval_runtime": 88.7831,
+ "eval_samples_per_second": 15.994,
+ "eval_steps_per_second": 2.005,
+ "step": 1600
+ },
+ {
+ "entropy": 0.5607876226305961,
+ "epoch": 3.9052504526252263,
+ "grad_norm": 0.9617928266525269,
+ "learning_rate": 0.00027893275476213383,
+ "loss": 0.5434338569641113,
+ "mean_token_accuracy": 0.8332898400723934,
+ "num_tokens": 3753013.0,
+ "step": 1620
+ },
+ {
+ "epoch": 3.9052504526252263,
+ "eval_entropy": 0.5474644318390428,
+ "eval_loss": 0.6740226745605469,
+ "eval_mean_token_accuracy": 0.8185851306058047,
+ "eval_num_tokens": 3753013.0,
+ "eval_runtime": 89.3246,
+ "eval_samples_per_second": 15.897,
+ "eval_steps_per_second": 1.993,
+ "step": 1620
+ },
+ {
+ "entropy": 0.5354843523353339,
+ "epoch": 3.9535304767652386,
+ "grad_norm": 1.5320990085601807,
+ "learning_rate": 0.0002763169207809021,
+ "loss": 0.5251852989196777,
+ "mean_token_accuracy": 0.8416872084140777,
+ "num_tokens": 3800024.0,
+ "step": 1640
+ },
+ {
+ "epoch": 3.9535304767652386,
+ "eval_entropy": 0.5436856410141742,
+ "eval_loss": 0.6675190925598145,
+ "eval_mean_token_accuracy": 0.8202396507343549,
+ "eval_num_tokens": 3800024.0,
+ "eval_runtime": 89.6328,
+ "eval_samples_per_second": 15.842,
+ "eval_steps_per_second": 1.986,
+ "step": 1640
+ },
+ {
+ "entropy": 0.5418571938167919,
+ "epoch": 4.0,
+ "grad_norm": 4.690354347229004,
+ "learning_rate": 0.000273674498421843,
+ "loss": 0.548275089263916,
+ "mean_token_accuracy": 0.836557479647847,
+ "num_tokens": 3844304.0,
+ "step": 1660
+ },
+ {
+ "epoch": 4.0,
+ "eval_entropy": 0.5497228983747825,
+ "eval_loss": 0.6621812582015991,
+ "eval_mean_token_accuracy": 0.8205679802412398,
+ "eval_num_tokens": 3844304.0,
+ "eval_runtime": 89.6272,
+ "eval_samples_per_second": 15.843,
+ "eval_steps_per_second": 1.986,
+ "step": 1660
+ },
+ {
+ "entropy": 0.4195931971073151,
+ "epoch": 4.048280024140012,
+ "grad_norm": 0.6678048968315125,
+ "learning_rate": 0.0002710062354592273,
+ "loss": 0.3970795154571533,
+ "mean_token_accuracy": 0.8719995342195034,
+ "num_tokens": 3890738.0,
+ "step": 1680
+ },
+ {
+ "epoch": 4.048280024140012,
+ "eval_entropy": 0.46495268753405367,
+ "eval_loss": 0.7061581611633301,
+ "eval_mean_token_accuracy": 0.8190107995204712,
+ "eval_num_tokens": 3890738.0,
+ "eval_runtime": 89.6387,
+ "eval_samples_per_second": 15.841,
+ "eval_steps_per_second": 1.986,
+ "step": 1680
+ },
+ {
+ "entropy": 0.4263830740004778,
+ "epoch": 4.096560048280024,
+ "grad_norm": 0.7026669383049011,
+ "learning_rate": 0.000268312886979911,
+ "loss": 0.398905086517334,
+ "mean_token_accuracy": 0.8691012300550938,
+ "num_tokens": 3938205.0,
+ "step": 1700
+ },
+ {
+ "epoch": 4.096560048280024,
+ "eval_entropy": 0.45755320670229666,
+ "eval_loss": 0.7102291584014893,
+ "eval_mean_token_accuracy": 0.819212955370378,
+ "eval_num_tokens": 3938205.0,
+ "eval_runtime": 89.2777,
+ "eval_samples_per_second": 15.905,
+ "eval_steps_per_second": 1.994,
+ "step": 1700
+ },
+ {
+ "entropy": 0.41640330031514167,
+ "epoch": 4.1448400724200365,
+ "grad_norm": 0.752216100692749,
+ "learning_rate": 0.00026559521516965437,
+ "loss": 0.39748263359069824,
+ "mean_token_accuracy": 0.8684553548693656,
+ "num_tokens": 3986350.0,
+ "step": 1720
+ },
+ {
+ "epoch": 4.1448400724200365,
+ "eval_entropy": 0.469178682297803,
+ "eval_loss": 0.7004125714302063,
+ "eval_mean_token_accuracy": 0.8194386949030201,
+ "eval_num_tokens": 3986350.0,
+ "eval_runtime": 89.8428,
+ "eval_samples_per_second": 15.805,
+ "eval_steps_per_second": 1.981,
+ "step": 1720
+ },
+ {
+ "entropy": 0.42079071439802646,
+ "epoch": 4.193120096560048,
+ "grad_norm": 0.8349477648735046,
+ "learning_rate": 0.0002628539890974329,
+ "loss": 0.40129623413085935,
+ "mean_token_accuracy": 0.8686790131032467,
+ "num_tokens": 4033541.0,
+ "step": 1740
+ },
+ {
+ "epoch": 4.193120096560048,
+ "eval_entropy": 0.48358210871058904,
+ "eval_loss": 0.70196533203125,
+ "eval_mean_token_accuracy": 0.818096743875675,
+ "eval_num_tokens": 4033541.0,
+ "eval_runtime": 90.444,
+ "eval_samples_per_second": 15.7,
+ "eval_steps_per_second": 1.968,
+ "step": 1740
+ },
+ {
+ "entropy": 0.4285690750926733,
+ "epoch": 4.24140012070006,
+ "grad_norm": 0.781974732875824,
+ "learning_rate": 0.00026008998449779933,
+ "loss": 0.40457854270935056,
+ "mean_token_accuracy": 0.8679066866636276,
+ "num_tokens": 4079253.0,
+ "step": 1760
+ },
+ {
+ "epoch": 4.24140012070006,
+ "eval_entropy": 0.4653420230645812,
+ "eval_loss": 0.7041526436805725,
+ "eval_mean_token_accuracy": 0.8190121017815022,
+ "eval_num_tokens": 4079253.0,
+ "eval_runtime": 89.5248,
+ "eval_samples_per_second": 15.862,
+ "eval_steps_per_second": 1.988,
+ "step": 1760
+ },
+ {
+ "entropy": 0.4335357129573822,
+ "epoch": 4.289680144840072,
+ "grad_norm": 0.8383703827857971,
+ "learning_rate": 0.0002573039835513604,
+ "loss": 0.41876921653747556,
+ "mean_token_accuracy": 0.8663770146667957,
+ "num_tokens": 4125495.0,
+ "step": 1780
+ },
+ {
+ "epoch": 4.289680144840072,
+ "eval_entropy": 0.5015502416350869,
+ "eval_loss": 0.6904259324073792,
+ "eval_mean_token_accuracy": 0.8200626497188311,
+ "eval_num_tokens": 4125495.0,
+ "eval_runtime": 89.7014,
+ "eval_samples_per_second": 15.83,
+ "eval_steps_per_second": 1.984,
+ "step": 1780
+ },
+ {
+ "entropy": 0.4258002445101738,
+ "epoch": 4.337960168980085,
+ "grad_norm": 0.7313240766525269,
+ "learning_rate": 0.00025449677466342867,
+ "loss": 0.40549774169921876,
+ "mean_token_accuracy": 0.8665009342133999,
+ "num_tokens": 4173422.0,
+ "step": 1800
+ },
+ {
+ "epoch": 4.337960168980085,
+ "eval_entropy": 0.49225394293833313,
+ "eval_loss": 0.6914838552474976,
+ "eval_mean_token_accuracy": 0.8200667657878962,
+ "eval_num_tokens": 4173422.0,
+ "eval_runtime": 91.6837,
+ "eval_samples_per_second": 15.488,
+ "eval_steps_per_second": 1.941,
+ "step": 1800
+ },
+ {
+ "entropy": 0.4372248936444521,
+ "epoch": 4.386240193120097,
+ "grad_norm": 0.8803642392158508,
+ "learning_rate": 0.0002516691522409134,
+ "loss": 0.41895318031311035,
+ "mean_token_accuracy": 0.861507610976696,
+ "num_tokens": 4216243.0,
+ "step": 1820
+ },
+ {
+ "epoch": 4.386240193120097,
+ "eval_entropy": 0.514270243852326,
+ "eval_loss": 0.6958855390548706,
+ "eval_mean_token_accuracy": 0.8197078490525149,
+ "eval_num_tokens": 4216243.0,
+ "eval_runtime": 89.6012,
+ "eval_samples_per_second": 15.848,
+ "eval_steps_per_second": 1.987,
+ "step": 1820
+ },
+ {
+ "entropy": 0.4372284222394228,
+ "epoch": 4.434520217260109,
+ "grad_norm": 0.7187909483909607,
+ "learning_rate": 0.0002488219164675126,
+ "loss": 0.4192944049835205,
+ "mean_token_accuracy": 0.8656419426202774,
+ "num_tokens": 4262463.0,
+ "step": 1840
+ },
+ {
+ "epoch": 4.434520217260109,
+ "eval_entropy": 0.4986347949571824,
+ "eval_loss": 0.6924281120300293,
+ "eval_mean_token_accuracy": 0.8201444835475321,
+ "eval_num_tokens": 4262463.0,
+ "eval_runtime": 89.6844,
+ "eval_samples_per_second": 15.833,
+ "eval_steps_per_second": 1.985,
+ "step": 1840
+ },
+ {
+ "entropy": 0.4428512301295996,
+ "epoch": 4.4828002414001205,
+ "grad_norm": 0.8485888242721558,
+ "learning_rate": 0.00024595587307727054,
+ "loss": 0.41988000869750974,
+ "mean_token_accuracy": 0.8625609986484051,
+ "num_tokens": 4311451.0,
+ "step": 1860
+ },
+ {
+ "epoch": 4.4828002414001205,
+ "eval_entropy": 0.49857732888018147,
+ "eval_loss": 0.6870604753494263,
+ "eval_mean_token_accuracy": 0.8219694069932016,
+ "eval_num_tokens": 4311451.0,
+ "eval_runtime": 89.7342,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 1860
+ },
+ {
+ "entropy": 0.44114561565220356,
+ "epoch": 4.531080265540133,
+ "grad_norm": 0.8288350701332092,
+ "learning_rate": 0.00024307183312656478,
+ "loss": 0.41482858657836913,
+ "mean_token_accuracy": 0.8646543353796006,
+ "num_tokens": 4357724.0,
+ "step": 1880
+ },
+ {
+ "epoch": 4.531080265540133,
+ "eval_entropy": 0.49263144577487133,
+ "eval_loss": 0.6941611766815186,
+ "eval_mean_token_accuracy": 0.8205747092038058,
+ "eval_num_tokens": 4357724.0,
+ "eval_runtime": 89.5086,
+ "eval_samples_per_second": 15.864,
+ "eval_steps_per_second": 1.989,
+ "step": 1880
+ },
+ {
+ "entropy": 0.4528769593685865,
+ "epoch": 4.579360289680145,
+ "grad_norm": 0.7440020442008972,
+ "learning_rate": 0.0002401706127645863,
+ "loss": 0.43678741455078124,
+ "mean_token_accuracy": 0.8619599856436253,
+ "num_tokens": 4401934.0,
+ "step": 1900
+ },
+ {
+ "epoch": 4.579360289680145,
+ "eval_entropy": 0.5085621437664782,
+ "eval_loss": 0.6766273379325867,
+ "eval_mean_token_accuracy": 0.8227389695939054,
+ "eval_num_tokens": 4401934.0,
+ "eval_runtime": 89.4661,
+ "eval_samples_per_second": 15.872,
+ "eval_steps_per_second": 1.99,
+ "step": 1900
+ },
+ {
+ "entropy": 0.4356316838413477,
+ "epoch": 4.627640313820157,
+ "grad_norm": 1.3870011568069458,
+ "learning_rate": 0.0002372530330023795,
+ "loss": 0.41338543891906737,
+ "mean_token_accuracy": 0.8642870113253593,
+ "num_tokens": 4449745.0,
+ "step": 1920
+ },
+ {
+ "epoch": 4.627640313820157,
+ "eval_entropy": 0.4954608427674583,
+ "eval_loss": 0.6783625483512878,
+ "eval_mean_token_accuracy": 0.8214637864841504,
+ "eval_num_tokens": 4449745.0,
+ "eval_runtime": 89.2291,
+ "eval_samples_per_second": 15.914,
+ "eval_steps_per_second": 1.995,
+ "step": 1920
+ },
+ {
+ "entropy": 0.4468349639326334,
+ "epoch": 4.675920337960169,
+ "grad_norm": 0.8028655648231506,
+ "learning_rate": 0.00023431991948050538,
+ "loss": 0.4324653625488281,
+ "mean_token_accuracy": 0.8607208795845509,
+ "num_tokens": 4495074.0,
+ "step": 1940
+ },
+ {
+ "epoch": 4.675920337960169,
+ "eval_entropy": 0.5259480836351266,
+ "eval_loss": 0.6809864044189453,
+ "eval_mean_token_accuracy": 0.8200277019752545,
+ "eval_num_tokens": 4495074.0,
+ "eval_runtime": 89.8247,
+ "eval_samples_per_second": 15.809,
+ "eval_steps_per_second": 1.982,
+ "step": 1940
+ },
+ {
+ "entropy": 0.4610355503857136,
+ "epoch": 4.724200362100181,
+ "grad_norm": 0.8928298950195312,
+ "learning_rate": 0.0002313721022353953,
+ "loss": 0.4314274787902832,
+ "mean_token_accuracy": 0.8559959702193737,
+ "num_tokens": 4541053.0,
+ "step": 1960
+ },
+ {
+ "epoch": 4.724200362100181,
+ "eval_entropy": 0.501018373651451,
+ "eval_loss": 0.6834176182746887,
+ "eval_mean_token_accuracy": 0.8220113566082515,
+ "eval_num_tokens": 4541053.0,
+ "eval_runtime": 90.4088,
+ "eval_samples_per_second": 15.706,
+ "eval_steps_per_second": 1.969,
+ "step": 1960
+ },
+ {
+ "entropy": 0.4568953149020672,
+ "epoch": 4.772480386240193,
+ "grad_norm": 1.018485188484192,
+ "learning_rate": 0.00022841041546446032,
+ "loss": 0.4306424617767334,
+ "mean_token_accuracy": 0.8598173558712006,
+ "num_tokens": 4584210.0,
+ "step": 1980
+ },
+ {
+ "epoch": 4.772480386240193,
+ "eval_entropy": 0.47633055371514865,
+ "eval_loss": 0.6886085867881775,
+ "eval_mean_token_accuracy": 0.823554899585381,
+ "eval_num_tokens": 4584210.0,
+ "eval_runtime": 89.8185,
+ "eval_samples_per_second": 15.81,
+ "eval_steps_per_second": 1.982,
+ "step": 1980
+ },
+ {
+ "entropy": 0.4718513362109661,
+ "epoch": 4.820760410380205,
+ "grad_norm": 0.7674988508224487,
+ "learning_rate": 0.00022543569729002318,
+ "loss": 0.44841318130493163,
+ "mean_token_accuracy": 0.8566767282783985,
+ "num_tokens": 4630254.0,
+ "step": 2000
+ },
+ {
+ "epoch": 4.820760410380205,
+ "eval_entropy": 0.49708309284086977,
+ "eval_loss": 0.6878687739372253,
+ "eval_mean_token_accuracy": 0.8217844698536262,
+ "eval_num_tokens": 4630254.0,
+ "eval_runtime": 89.7497,
+ "eval_samples_per_second": 15.822,
+ "eval_steps_per_second": 1.983,
+ "step": 2000
+ },
+ {
+ "entropy": 0.44017268233001233,
+ "epoch": 4.869040434520217,
+ "grad_norm": 0.6380876898765564,
+ "learning_rate": 0.00022244878952213976,
+ "loss": 0.42589750289916994,
+ "mean_token_accuracy": 0.8615814067423344,
+ "num_tokens": 4679265.0,
+ "step": 2020
+ },
+ {
+ "epoch": 4.869040434520217,
+ "eval_entropy": 0.5042208635740066,
+ "eval_loss": 0.6718228459358215,
+ "eval_mean_token_accuracy": 0.8232053056191863,
+ "eval_num_tokens": 4679265.0,
+ "eval_runtime": 89.7331,
+ "eval_samples_per_second": 15.825,
+ "eval_steps_per_second": 1.984,
+ "step": 2020
+ },
+ {
+ "entropy": 0.4654298175126314,
+ "epoch": 4.917320458660229,
+ "grad_norm": 0.8285323977470398,
+ "learning_rate": 0.0002194505374203766,
+ "loss": 0.44148526191711424,
+ "mean_token_accuracy": 0.8545936144888401,
+ "num_tokens": 4723592.0,
+ "step": 2040
+ },
+ {
+ "epoch": 4.917320458660229,
+ "eval_entropy": 0.4957777807551823,
+ "eval_loss": 0.6694385409355164,
+ "eval_mean_token_accuracy": 0.8243840380331103,
+ "eval_num_tokens": 4723592.0,
+ "eval_runtime": 89.5293,
+ "eval_samples_per_second": 15.861,
+ "eval_steps_per_second": 1.988,
+ "step": 2040
+ },
+ {
+ "entropy": 0.46651641055941584,
+ "epoch": 4.965600482800241,
+ "grad_norm": 0.946260929107666,
+ "learning_rate": 0.00021644178945461268,
+ "loss": 0.43698992729187014,
+ "mean_token_accuracy": 0.8586322009563446,
+ "num_tokens": 4767253.0,
+ "step": 2060
+ },
+ {
+ "epoch": 4.965600482800241,
+ "eval_entropy": 0.5063281496254246,
+ "eval_loss": 0.6729032397270203,
+ "eval_mean_token_accuracy": 0.824299869912394,
+ "eval_num_tokens": 4767253.0,
+ "eval_runtime": 89.6207,
+ "eval_samples_per_second": 15.845,
+ "eval_steps_per_second": 1.986,
+ "step": 2060
+ },
+ {
+ "entropy": 0.4135968596130222,
+ "epoch": 5.012070006035003,
+ "grad_norm": 0.6853868961334229,
+ "learning_rate": 0.0002134233970649322,
+ "loss": 0.3800280332565308,
+ "mean_token_accuracy": 0.8756770677380747,
+ "num_tokens": 4816587.0,
+ "step": 2080
+ },
+ {
+ "epoch": 5.012070006035003,
+ "eval_entropy": 0.4107752423942759,
+ "eval_loss": 0.7204955220222473,
+ "eval_mean_token_accuracy": 0.8226918417416261,
+ "eval_num_tokens": 4816587.0,
+ "eval_runtime": 89.5817,
+ "eval_samples_per_second": 15.851,
+ "eval_steps_per_second": 1.987,
+ "step": 2080
+ },
+ {
+ "entropy": 0.316746598854661,
+ "epoch": 5.060350030175015,
+ "grad_norm": 0.6939145922660828,
+ "learning_rate": 0.00021039621442067732,
+ "loss": 0.2942554473876953,
+ "mean_token_accuracy": 0.90015934035182,
+ "num_tokens": 4861071.0,
+ "step": 2100
+ },
+ {
+ "epoch": 5.060350030175015,
+ "eval_entropy": 0.4214997919422857,
+ "eval_loss": 0.7268305420875549,
+ "eval_mean_token_accuracy": 0.8216965369294199,
+ "eval_num_tokens": 4861071.0,
+ "eval_runtime": 89.8077,
+ "eval_samples_per_second": 15.812,
+ "eval_steps_per_second": 1.982,
+ "step": 2100
+ },
+ {
+ "entropy": 0.3187096672132611,
+ "epoch": 5.108630054315027,
+ "grad_norm": 0.9134758114814758,
+ "learning_rate": 0.00020736109817872823,
+ "loss": 0.2839747190475464,
+ "mean_token_accuracy": 0.9006688073277473,
+ "num_tokens": 4904742.0,
+ "step": 2120
+ },
+ {
+ "epoch": 5.108630054315027,
+ "eval_entropy": 0.4375539443800958,
+ "eval_loss": 0.7280111908912659,
+ "eval_mean_token_accuracy": 0.8186122694712007,
+ "eval_num_tokens": 4904742.0,
+ "eval_runtime": 89.5868,
+ "eval_samples_per_second": 15.851,
+ "eval_steps_per_second": 1.987,
+ "step": 2120
+ },
+ {
+ "entropy": 0.3137735094875097,
+ "epoch": 5.15691007845504,
+ "grad_norm": 0.7324768304824829,
+ "learning_rate": 0.00020431890724107943,
+ "loss": 0.28816254138946534,
+ "mean_token_accuracy": 0.9023615352809429,
+ "num_tokens": 4950795.0,
+ "step": 2140
+ },
+ {
+ "epoch": 5.15691007845504,
+ "eval_entropy": 0.42255786312430094,
+ "eval_loss": 0.7335684895515442,
+ "eval_mean_token_accuracy": 0.8213735473959634,
+ "eval_num_tokens": 4950795.0,
+ "eval_runtime": 89.5711,
+ "eval_samples_per_second": 15.853,
+ "eval_steps_per_second": 1.987,
+ "step": 2140
+ },
+ {
+ "entropy": 0.3310150146484375,
+ "epoch": 5.2051901025950515,
+ "grad_norm": 0.7819830179214478,
+ "learning_rate": 0.00020127050251178062,
+ "loss": 0.3020722150802612,
+ "mean_token_accuracy": 0.8957489147782326,
+ "num_tokens": 4994886.0,
+ "step": 2160
+ },
+ {
+ "epoch": 5.2051901025950515,
+ "eval_entropy": 0.43374256186940696,
+ "eval_loss": 0.7270973920822144,
+ "eval_mean_token_accuracy": 0.821890847066815,
+ "eval_num_tokens": 4994886.0,
+ "eval_runtime": 89.7117,
+ "eval_samples_per_second": 15.828,
+ "eval_steps_per_second": 1.984,
+ "step": 2160
+ },
+ {
+ "entropy": 0.31917148139327767,
+ "epoch": 5.253470126735063,
+ "grad_norm": 0.8766310811042786,
+ "learning_rate": 0.00019821674665331112,
+ "loss": 0.29281470775604246,
+ "mean_token_accuracy": 0.8986819669604301,
+ "num_tokens": 5043663.0,
+ "step": 2180
+ },
+ {
+ "epoch": 5.253470126735063,
+ "eval_entropy": 0.4188300051381079,
+ "eval_loss": 0.7247599959373474,
+ "eval_mean_token_accuracy": 0.8234477113471942,
+ "eval_num_tokens": 5043663.0,
+ "eval_runtime": 89.384,
+ "eval_samples_per_second": 15.887,
+ "eval_steps_per_second": 1.991,
+ "step": 2180
+ },
+ {
+ "entropy": 0.3252310147508979,
+ "epoch": 5.301750150875075,
+ "grad_norm": 0.8721100687980652,
+ "learning_rate": 0.0001951585038424563,
+ "loss": 0.2924081325531006,
+ "mean_token_accuracy": 0.8981824897229671,
+ "num_tokens": 5087959.0,
+ "step": 2200
+ },
+ {
+ "epoch": 5.301750150875075,
+ "eval_entropy": 0.4114788512835342,
+ "eval_loss": 0.735135555267334,
+ "eval_mean_token_accuracy": 0.8240512961082245,
+ "eval_num_tokens": 5087959.0,
+ "eval_runtime": 89.3499,
+ "eval_samples_per_second": 15.893,
+ "eval_steps_per_second": 1.992,
+ "step": 2200
+ },
+ {
+ "entropy": 0.3175549603998661,
+ "epoch": 5.350030175015087,
+ "grad_norm": 0.8240578174591064,
+ "learning_rate": 0.00019209663952575616,
+ "loss": 0.2883902072906494,
+ "mean_token_accuracy": 0.9001291915774345,
+ "num_tokens": 5136121.0,
+ "step": 2220
+ },
+ {
+ "epoch": 5.350030175015087,
+ "eval_entropy": 0.417589228474692,
+ "eval_loss": 0.7371609807014465,
+ "eval_mean_token_accuracy": 0.8225449795803327,
+ "eval_num_tokens": 5136121.0,
+ "eval_runtime": 89.1634,
+ "eval_samples_per_second": 15.926,
+ "eval_steps_per_second": 1.996,
+ "step": 2220
+ },
+ {
+ "entropy": 0.32724177204072474,
+ "epoch": 5.3983101991551,
+ "grad_norm": 0.9292752742767334,
+ "learning_rate": 0.00018903202017459399,
+ "loss": 0.3026577472686768,
+ "mean_token_accuracy": 0.8965070247650146,
+ "num_tokens": 5182144.0,
+ "step": 2240
+ },
+ {
+ "epoch": 5.3983101991551,
+ "eval_entropy": 0.4234565015924111,
+ "eval_loss": 0.7295495867729187,
+ "eval_mean_token_accuracy": 0.8232441668430072,
+ "eval_num_tokens": 5182144.0,
+ "eval_runtime": 89.016,
+ "eval_samples_per_second": 15.952,
+ "eval_steps_per_second": 2.0,
+ "step": 2240
+ }
+ ],
+ "logging_steps": 20,
+ "max_steps": 4150,
+ "num_input_tokens_seen": 0,
+ "num_train_epochs": 10,
+ "save_steps": 20,
+ "stateful_callbacks": {
+ "TrainerControl": {
+ "args": {
+ "should_epoch_stop": false,
+ "should_evaluate": false,
+ "should_log": false,
+ "should_save": true,
+ "should_training_stop": false
+ },
+ "attributes": {}
+ }
+ },
+ "total_flos": 2.238685165496832e+17,
+ "train_batch_size": 4,
+ "trial_name": null,
+ "trial_params": null
+}
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2260/README.md b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2260/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..41e6c854e77830e9ea767c8c35f8c82a65c1ba35
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2260/README.md
@@ -0,0 +1,209 @@
+---
+base_model: Qwen/Qwen3.5-4B-Base
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen3.5-4B-Base
+- lora
+- sft
+- transformers
+- trl
+---
+
+# Model Card for Model ID
+
+
+
+
+
+## Model Details
+
+### Model Description
+
+
+
+
+
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+
+### Model Sources [optional]
+
+
+
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+
+## Uses
+
+
+
+### Direct Use
+
+
+
+[More Information Needed]
+
+### Downstream Use [optional]
+
+
+
+[More Information Needed]
+
+### Out-of-Scope Use
+
+
+
+[More Information Needed]
+
+## Bias, Risks, and Limitations
+
+
+
+[More Information Needed]
+
+### Recommendations
+
+
+
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+
+## How to Get Started with the Model
+
+Use the code below to get started with the model.
+
+[More Information Needed]
+
+## Training Details
+
+### Training Data
+
+
+
+[More Information Needed]
+
+### Training Procedure
+
+
+
+#### Preprocessing [optional]
+
+[More Information Needed]
+
+
+#### Training Hyperparameters
+
+- **Training regime:** [More Information Needed]
+
+#### Speeds, Sizes, Times [optional]
+
+
+
+[More Information Needed]
+
+## Evaluation
+
+
+
+### Testing Data, Factors & Metrics
+
+#### Testing Data
+
+
+
+[More Information Needed]
+
+#### Factors
+
+
+
+[More Information Needed]
+
+#### Metrics
+
+
+
+[More Information Needed]
+
+### Results
+
+[More Information Needed]
+
+#### Summary
+
+
+
+## Model Examination [optional]
+
+
+
+[More Information Needed]
+
+## Environmental Impact
+
+
+
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+
+## Technical Specifications [optional]
+
+### Model Architecture and Objective
+
+[More Information Needed]
+
+### Compute Infrastructure
+
+[More Information Needed]
+
+#### Hardware
+
+[More Information Needed]
+
+#### Software
+
+[More Information Needed]
+
+## Citation [optional]
+
+
+
+**BibTeX:**
+
+[More Information Needed]
+
+**APA:**
+
+[More Information Needed]
+
+## Glossary [optional]
+
+
+
+[More Information Needed]
+
+## More Information [optional]
+
+[More Information Needed]
+
+## Model Card Authors [optional]
+
+[More Information Needed]
+
+## Model Card Contact
+
+[More Information Needed]
+### Framework versions
+
+- PEFT 0.18.1
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2260/adapter_config.json b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2260/adapter_config.json
new file mode 100644
index 0000000000000000000000000000000000000000..7ce4b072fcbe5b636867e13937c56fedee3b68b4
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2260/adapter_config.json
@@ -0,0 +1,46 @@
+{
+ "alora_invocation_tokens": null,
+ "alpha_pattern": {},
+ "arrow_config": null,
+ "auto_mapping": null,
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B-Base",
+ "bias": "none",
+ "corda_config": null,
+ "ensure_weight_tying": false,
+ "eva_config": null,
+ "exclude_modules": null,
+ "fan_in_fan_out": false,
+ "inference_mode": true,
+ "init_lora_weights": true,
+ "layer_replication": null,
+ "layers_pattern": null,
+ "layers_to_transform": null,
+ "loftq_config": {},
+ "lora_alpha": 256,
+ "lora_bias": false,
+ "lora_dropout": 0.02728394274790723,
+ "megatron_config": null,
+ "megatron_core": "megatron.core",
+ "modules_to_save": null,
+ "peft_type": "LORA",
+ "peft_version": "0.18.1",
+ "qalora_group_size": 16,
+ "r": 128,
+ "rank_pattern": {},
+ "revision": null,
+ "target_modules": [
+ "gate_proj",
+ "v_proj",
+ "up_proj",
+ "down_proj",
+ "o_proj",
+ "q_proj",
+ "k_proj"
+ ],
+ "target_parameters": null,
+ "task_type": "CAUSAL_LM",
+ "trainable_token_indices": null,
+ "use_dora": false,
+ "use_qalora": false,
+ "use_rslora": false
+}
\ No newline at end of file
diff --git a/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2260/chat_template.jinja b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2260/chat_template.jinja
new file mode 100644
index 0000000000000000000000000000000000000000..a585dec894e63da457d9440ec6aa7caa16d20860
--- /dev/null
+++ b/overgeneralisation_original_Estonian/Qwen3.5-4B-Base_overgeneralisation_splits_original_features_train_overgeneralisation_splits_original_features_test1/checkpoint-2260/chat_template.jinja
@@ -0,0 +1,154 @@
+{%- set image_count = namespace(value=0) %}
+{%- set video_count = namespace(value=0) %}
+{%- macro render_content(content, do_vision_count, is_system_content=false) %}
+ {%- if content is string %}
+ {{- content }}
+ {%- elif content is iterable and content is not mapping %}
+ {%- for item in content %}
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain images.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set image_count.value = image_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
+ {%- elif 'video' in item or item.type == 'video' %}
+ {%- if is_system_content %}
+ {{- raise_exception('System message cannot contain videos.') }}
+ {%- endif %}
+ {%- if do_vision_count %}
+ {%- set video_count.value = video_count.value + 1 %}
+ {%- endif %}
+ {%- if add_vision_id %}
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
+ {%- endif %}
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
+ {%- elif 'text' in item %}
+ {{- item.text }}
+ {%- else %}
+ {{- raise_exception('Unexpected item type in content.') }}
+ {%- endif %}
+ {%- endfor %}
+ {%- elif content is none or content is undefined %}
+ {{- '' }}
+ {%- else %}
+ {{- raise_exception('Unexpected content type.') }}
+ {%- endif %}
+{%- endmacro %}
+{%- if not messages %}
+ {{- raise_exception('No messages provided.') }}
+{%- endif %}
+{%- if tools and tools is iterable and tools is not mapping %}
+ {{- '<|im_start|>system\n' }}
+ {{- "# Tools\n\nYou have access to the following functions:\n\n" }}
+ {%- for tool in tools %}
+ {{- "\n" }}
+ {{- tool | tojson }}
+ {%- endfor %}
+ {{- "\n" }}
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {%- if content %}
+ {{- '\n\n' + content }}
+ {%- endif %}
+ {%- endif %}
+ {{- '<|im_end|>\n' }}
+{%- else %}
+ {%- if messages[0].role == 'system' %}
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
+ {%- endif %}
+{%- endif %}
+{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
+{%- for message in messages[::-1] %}
+ {%- set index = (messages|length - 1) - loop.index0 %}
+ {%- if ns.multi_step_tool and message.role == "user" %}
+ {%- set content = render_content(message.content, false)|trim %}
+ {%- if not(content.startswith('') and content.endswith('')) %}
+ {%- set ns.multi_step_tool = false %}
+ {%- set ns.last_query_index = index %}
+ {%- endif %}
+ {%- endif %}
+{%- endfor %}
+{%- if ns.multi_step_tool %}
+ {{- raise_exception('No user query found in messages.') }}
+{%- endif %}
+{%- for message in messages %}
+ {%- set content = render_content(message.content, true)|trim %}
+ {%- if message.role == "system" %}
+ {%- if not loop.first %}
+ {{- raise_exception('System message must be at the beginning.') }}
+ {%- endif %}
+ {%- elif message.role == "user" %}
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
+ {%- elif message.role == "assistant" %}
+ {%- set reasoning_content = '' %}
+ {%- if message.reasoning_content is string %}
+ {%- set reasoning_content = message.reasoning_content %}
+ {%- else %}
+ {%- if '' in content %}
+ {%- set reasoning_content = content.split('')[0].rstrip('\n').split('')[-1].lstrip('\n') %}
+ {%- set content = content.split('')[-1].lstrip('\n') %}
+ {%- endif %}
+ {%- endif %}
+ {%- set reasoning_content = reasoning_content|trim %}
+ {%- if loop.index0 > ns.last_query_index %}
+ {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }}
+ {%- else %}
+ {{- '<|im_start|>' + message.role + '\n' + content }}
+ {%- endif %}
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
+ {%- for tool_call in message.tool_calls %}
+ {%- if tool_call.function is defined %}
+ {%- set tool_call = tool_call.function %}
+ {%- endif %}
+ {%- if loop.first %}
+ {%- if content|trim %}
+ {{- '\n\n\n\n' }}
+ {%- else %}
+ {{- '\n\n' }}
+ {%- endif %}
+ {%- else %}
+ {{- '\n\n\n' }}
+ {%- endif %}
+ {%- if tool_call.arguments is defined %}
+ {%- for args_name, args_value in tool_call.arguments|items %}
+ {{- '