Instructions to use furproxy/27b-4-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use furproxy/27b-4-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("/workspace/models/Qwen3.6-27B") model = PeftModel.from_pretrained(base_model, "furproxy/27b-4-lora") - Transformers
How to use furproxy/27b-4-lora with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="furproxy/27b-4-lora") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("furproxy/27b-4-lora", dtype="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use furproxy/27b-4-lora with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "furproxy/27b-4-lora" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "furproxy/27b-4-lora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/furproxy/27b-4-lora
- SGLang
How to use furproxy/27b-4-lora with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "furproxy/27b-4-lora" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "furproxy/27b-4-lora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "furproxy/27b-4-lora" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "furproxy/27b-4-lora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use furproxy/27b-4-lora with Docker Model Runner:
docker model run hf.co/furproxy/27b-4-lora
Upload folder using huggingface_hub
Browse files- .gitattributes +1 -0
- .ipynb_checkpoints/README-checkpoint.md +66 -0
- .ipynb_checkpoints/trainer_log-checkpoint.jsonl +252 -0
- README.md +66 -0
- adapter_config.json +55 -0
- adapter_model.safetensors +3 -0
- all_results.json +9 -0
- chat_template.jinja +154 -0
- processor_config.json +60 -0
- tokenizer.json +3 -0
- tokenizer_config.json +37 -0
- train_results.json +9 -0
- trainer_log.jsonl +0 -0
- trainer_state.json +0 -0
- training_args.bin +3 -0
- training_loss.png +0 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
.ipynb_checkpoints/README-checkpoint.md
ADDED
|
@@ -0,0 +1,66 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
library_name: peft
|
| 3 |
+
license: other
|
| 4 |
+
base_model: Qwen3.6-27B
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:/workspace/models/Qwen3.6-27B
|
| 7 |
+
- llama-factory
|
| 8 |
+
- lora
|
| 9 |
+
- transformers
|
| 10 |
+
pipeline_tag: text-generation
|
| 11 |
+
model-index:
|
| 12 |
+
- name: qwen35_caption_galore
|
| 13 |
+
results: []
|
| 14 |
+
---
|
| 15 |
+
|
| 16 |
+
<!-- This model card has been generated automatically according to the information the Trainer had access to. You
|
| 17 |
+
should probably proofread and complete it, then remove this comment. -->
|
| 18 |
+
|
| 19 |
+
# qwen35_caption_galore
|
| 20 |
+
|
| 21 |
+
This model is a fine-tuned version of [/workspace/models/Qwen3.6-27B](https://huggingface.co//workspace/models/Qwen3.6-27B) on the my_caption dataset.
|
| 22 |
+
|
| 23 |
+
## Model description
|
| 24 |
+
|
| 25 |
+
More information needed
|
| 26 |
+
|
| 27 |
+
## Intended uses & limitations
|
| 28 |
+
|
| 29 |
+
More information needed
|
| 30 |
+
|
| 31 |
+
## Training and evaluation data
|
| 32 |
+
|
| 33 |
+
More information needed
|
| 34 |
+
|
| 35 |
+
## Training procedure
|
| 36 |
+
|
| 37 |
+
### Training hyperparameters
|
| 38 |
+
|
| 39 |
+
The following hyperparameters were used during training:
|
| 40 |
+
- family_to_adamw_lr = {
|
| 41 |
+
"language": _fallback(getattr(training_args, "language_adamw_lr", 2e-5), language_lr),
|
| 42 |
+
"vision": _fallback(getattr(training_args, "vision_adamw_lr", 2e-5), vision_lr),
|
| 43 |
+
"merger": _fallback(getattr(training_args, "merger_adamw_lr", 4e-5), merger_lr),
|
| 44 |
+
}
|
| 45 |
+
- train_batch_size: 1
|
| 46 |
+
- eval_batch_size: 8
|
| 47 |
+
- seed: 42
|
| 48 |
+
- distributed_type: multi-GPU
|
| 49 |
+
- gradient_accumulation_steps: 24
|
| 50 |
+
- total_train_batch_size: 24
|
| 51 |
+
- optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
|
| 52 |
+
- lr_scheduler_type: cosine_with_min_lr
|
| 53 |
+
- lr_scheduler_warmup_steps: 0.03
|
| 54 |
+
- num_epochs: 3
|
| 55 |
+
|
| 56 |
+
### Training results
|
| 57 |
+
|
| 58 |
+
|
| 59 |
+
|
| 60 |
+
### Framework versions
|
| 61 |
+
|
| 62 |
+
- PEFT 0.18.1
|
| 63 |
+
- Transformers 5.5.3
|
| 64 |
+
- Pytorch 2.11.0+cu128
|
| 65 |
+
- Datasets 4.0.0
|
| 66 |
+
- Tokenizers 0.22.2
|
.ipynb_checkpoints/trainer_log-checkpoint.jsonl
ADDED
|
@@ -0,0 +1,252 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"current_steps": 2, "total_steps": 1638, "loss": 2.6876986026763916, "lr": 4.0000000000000003e-07, "epoch": 0.003663003663003663, "percentage": 0.12, "elapsed_time": "0:01:16", "remaining_time": "17:27:26"}
|
| 2 |
+
{"current_steps": 4, "total_steps": 1638, "loss": 1.6663331985473633, "lr": 1.2000000000000002e-06, "epoch": 0.007326007326007326, "percentage": 0.24, "elapsed_time": "0:02:19", "remaining_time": "15:51:28"}
|
| 3 |
+
{"current_steps": 6, "total_steps": 1638, "loss": 1.881505012512207, "lr": 2.0000000000000003e-06, "epoch": 0.01098901098901099, "percentage": 0.37, "elapsed_time": "0:03:26", "remaining_time": "15:38:10"}
|
| 4 |
+
{"current_steps": 8, "total_steps": 1638, "loss": 2.063166618347168, "lr": 2.8000000000000003e-06, "epoch": 0.014652014652014652, "percentage": 0.49, "elapsed_time": "0:04:40", "remaining_time": "15:52:54"}
|
| 5 |
+
{"current_steps": 10, "total_steps": 1638, "loss": 2.217334747314453, "lr": 3.6000000000000003e-06, "epoch": 0.018315018315018316, "percentage": 0.61, "elapsed_time": "0:05:53", "remaining_time": "15:59:13"}
|
| 6 |
+
{"current_steps": 12, "total_steps": 1638, "loss": 2.016745090484619, "lr": 4.4e-06, "epoch": 0.02197802197802198, "percentage": 0.73, "elapsed_time": "0:07:05", "remaining_time": "16:01:10"}
|
| 7 |
+
{"current_steps": 14, "total_steps": 1638, "loss": 1.754088044166565, "lr": 5.2e-06, "epoch": 0.02564102564102564, "percentage": 0.85, "elapsed_time": "0:08:18", "remaining_time": "16:03:15"}
|
| 8 |
+
{"current_steps": 16, "total_steps": 1638, "loss": 1.8193743228912354, "lr": 6e-06, "epoch": 0.029304029304029304, "percentage": 0.98, "elapsed_time": "0:09:29", "remaining_time": "16:02:09"}
|
| 9 |
+
{"current_steps": 18, "total_steps": 1638, "loss": 1.7517926692962646, "lr": 6.800000000000001e-06, "epoch": 0.03296703296703297, "percentage": 1.1, "elapsed_time": "0:10:42", "remaining_time": "16:03:57"}
|
| 10 |
+
{"current_steps": 20, "total_steps": 1638, "loss": 1.740407943725586, "lr": 7.600000000000001e-06, "epoch": 0.03663003663003663, "percentage": 1.22, "elapsed_time": "0:12:00", "remaining_time": "16:11:17"}
|
| 11 |
+
{"current_steps": 22, "total_steps": 1638, "loss": 1.112322449684143, "lr": 8.400000000000001e-06, "epoch": 0.040293040293040296, "percentage": 1.34, "elapsed_time": "0:12:56", "remaining_time": "15:50:04"}
|
| 12 |
+
{"current_steps": 24, "total_steps": 1638, "loss": 1.2775133848190308, "lr": 9.200000000000002e-06, "epoch": 0.04395604395604396, "percentage": 1.47, "elapsed_time": "0:14:05", "remaining_time": "15:47:14"}
|
| 13 |
+
{"current_steps": 26, "total_steps": 1638, "loss": 1.4573379755020142, "lr": 1e-05, "epoch": 0.047619047619047616, "percentage": 1.59, "elapsed_time": "0:15:20", "remaining_time": "15:51:24"}
|
| 14 |
+
{"current_steps": 28, "total_steps": 1638, "loss": 1.476364016532898, "lr": 1.0800000000000002e-05, "epoch": 0.05128205128205128, "percentage": 1.71, "elapsed_time": "0:16:16", "remaining_time": "15:35:44"}
|
| 15 |
+
{"current_steps": 30, "total_steps": 1638, "loss": 1.1695544719696045, "lr": 1.16e-05, "epoch": 0.054945054945054944, "percentage": 1.83, "elapsed_time": "0:17:25", "remaining_time": "15:33:51"}
|
| 16 |
+
{"current_steps": 32, "total_steps": 1638, "loss": 1.108648419380188, "lr": 1.2400000000000002e-05, "epoch": 0.05860805860805861, "percentage": 1.95, "elapsed_time": "0:18:39", "remaining_time": "15:36:37"}
|
| 17 |
+
{"current_steps": 34, "total_steps": 1638, "loss": 1.2354477643966675, "lr": 1.3200000000000002e-05, "epoch": 0.06227106227106227, "percentage": 2.08, "elapsed_time": "0:19:41", "remaining_time": "15:29:02"}
|
| 18 |
+
{"current_steps": 36, "total_steps": 1638, "loss": 1.6114126443862915, "lr": 1.4e-05, "epoch": 0.06593406593406594, "percentage": 2.2, "elapsed_time": "0:20:58", "remaining_time": "15:33:24"}
|
| 19 |
+
{"current_steps": 38, "total_steps": 1638, "loss": 1.375102162361145, "lr": 1.48e-05, "epoch": 0.0695970695970696, "percentage": 2.32, "elapsed_time": "0:22:12", "remaining_time": "15:34:52"}
|
| 20 |
+
{"current_steps": 40, "total_steps": 1638, "loss": 1.4309669733047485, "lr": 1.5600000000000003e-05, "epoch": 0.07326007326007326, "percentage": 2.44, "elapsed_time": "0:23:23", "remaining_time": "15:34:16"}
|
| 21 |
+
{"current_steps": 42, "total_steps": 1638, "loss": 1.1326210498809814, "lr": 1.64e-05, "epoch": 0.07692307692307693, "percentage": 2.56, "elapsed_time": "0:24:35", "remaining_time": "15:34:21"}
|
| 22 |
+
{"current_steps": 44, "total_steps": 1638, "loss": 1.532023549079895, "lr": 1.72e-05, "epoch": 0.08058608058608059, "percentage": 2.69, "elapsed_time": "0:25:53", "remaining_time": "15:37:56"}
|
| 23 |
+
{"current_steps": 46, "total_steps": 1638, "loss": 1.6225450038909912, "lr": 1.8e-05, "epoch": 0.08424908424908426, "percentage": 2.81, "elapsed_time": "0:27:06", "remaining_time": "15:38:19"}
|
| 24 |
+
{"current_steps": 48, "total_steps": 1638, "loss": 1.0662028789520264, "lr": 1.88e-05, "epoch": 0.08791208791208792, "percentage": 2.93, "elapsed_time": "0:28:20", "remaining_time": "15:38:48"}
|
| 25 |
+
{"current_steps": 50, "total_steps": 1638, "loss": 1.509545087814331, "lr": 1.9600000000000002e-05, "epoch": 0.09157509157509157, "percentage": 3.05, "elapsed_time": "0:29:32", "remaining_time": "15:38:29"}
|
| 26 |
+
{"current_steps": 52, "total_steps": 1638, "loss": 0.7481744885444641, "lr": 1.999998238790087e-05, "epoch": 0.09523809523809523, "percentage": 3.17, "elapsed_time": "0:30:28", "remaining_time": "15:29:36"}
|
| 27 |
+
{"current_steps": 54, "total_steps": 1638, "loss": 0.9735360145568848, "lr": 1.999984149152137e-05, "epoch": 0.0989010989010989, "percentage": 3.3, "elapsed_time": "0:31:41", "remaining_time": "15:29:27"}
|
| 28 |
+
{"current_steps": 56, "total_steps": 1638, "loss": 1.3465101718902588, "lr": 1.999955970096814e-05, "epoch": 0.10256410256410256, "percentage": 3.42, "elapsed_time": "0:32:54", "remaining_time": "15:29:46"}
|
| 29 |
+
{"current_steps": 58, "total_steps": 1638, "loss": 1.1941382884979248, "lr": 1.9999137020652663e-05, "epoch": 0.10622710622710622, "percentage": 3.54, "elapsed_time": "0:34:04", "remaining_time": "15:28:05"}
|
| 30 |
+
{"current_steps": 60, "total_steps": 1638, "loss": 1.4114629030227661, "lr": 1.999857345719207e-05, "epoch": 0.10989010989010989, "percentage": 3.66, "elapsed_time": "0:35:16", "remaining_time": "15:27:43"}
|
| 31 |
+
{"current_steps": 62, "total_steps": 1638, "loss": 1.428280234336853, "lr": 1.9997869019409047e-05, "epoch": 0.11355311355311355, "percentage": 3.79, "elapsed_time": "0:36:29", "remaining_time": "15:27:39"}
|
| 32 |
+
{"current_steps": 64, "total_steps": 1638, "loss": 1.3913908004760742, "lr": 1.9997023718331707e-05, "epoch": 0.11721611721611722, "percentage": 3.91, "elapsed_time": "0:37:47", "remaining_time": "15:29:22"}
|
| 33 |
+
{"current_steps": 66, "total_steps": 1638, "loss": 1.3538874387741089, "lr": 1.9996037567193388e-05, "epoch": 0.12087912087912088, "percentage": 4.03, "elapsed_time": "0:39:05", "remaining_time": "15:31:13"}
|
| 34 |
+
{"current_steps": 68, "total_steps": 1638, "loss": 1.3233115673065186, "lr": 1.9994910581432466e-05, "epoch": 0.12454212454212454, "percentage": 4.15, "elapsed_time": "0:40:24", "remaining_time": "15:33:04"}
|
| 35 |
+
{"current_steps": 70, "total_steps": 1638, "loss": 1.1092506647109985, "lr": 1.9993642778692116e-05, "epoch": 0.1282051282051282, "percentage": 4.27, "elapsed_time": "0:41:21", "remaining_time": "15:26:24"}
|
| 36 |
+
{"current_steps": 72, "total_steps": 1638, "loss": 1.431759238243103, "lr": 1.999223417882002e-05, "epoch": 0.13186813186813187, "percentage": 4.4, "elapsed_time": "0:42:38", "remaining_time": "15:27:19"}
|
| 37 |
+
{"current_steps": 74, "total_steps": 1638, "loss": 1.593772292137146, "lr": 1.9990684803868068e-05, "epoch": 0.13553113553113552, "percentage": 4.52, "elapsed_time": "0:43:49", "remaining_time": "15:26:24"}
|
| 38 |
+
{"current_steps": 76, "total_steps": 1638, "loss": 1.0841394662857056, "lr": 1.9988994678092007e-05, "epoch": 0.1391941391941392, "percentage": 4.64, "elapsed_time": "0:44:47", "remaining_time": "15:20:28"}
|
| 39 |
+
{"current_steps": 78, "total_steps": 1638, "loss": 1.434753656387329, "lr": 1.9987163827951077e-05, "epoch": 0.14285714285714285, "percentage": 4.76, "elapsed_time": "0:45:59", "remaining_time": "15:19:55"}
|
| 40 |
+
{"current_steps": 80, "total_steps": 1638, "loss": 1.5812965631484985, "lr": 1.998519228210756e-05, "epoch": 0.14652014652014653, "percentage": 4.88, "elapsed_time": "0:47:15", "remaining_time": "15:20:15"}
|
| 41 |
+
{"current_steps": 82, "total_steps": 1638, "loss": 1.1953579187393188, "lr": 1.998308007142638e-05, "epoch": 0.15018315018315018, "percentage": 5.01, "elapsed_time": "0:48:19", "remaining_time": "15:16:51"}
|
| 42 |
+
{"current_steps": 84, "total_steps": 1638, "loss": 1.3637341260910034, "lr": 1.9980827228974575e-05, "epoch": 0.15384615384615385, "percentage": 5.13, "elapsed_time": "0:49:35", "remaining_time": "15:17:26"}
|
| 43 |
+
{"current_steps": 86, "total_steps": 1638, "loss": 1.4977788925170898, "lr": 1.997843379002081e-05, "epoch": 0.1575091575091575, "percentage": 5.25, "elapsed_time": "0:50:48", "remaining_time": "15:16:47"}
|
| 44 |
+
{"current_steps": 88, "total_steps": 1638, "loss": 0.6990910768508911, "lr": 1.9975899792034824e-05, "epoch": 0.16117216117216118, "percentage": 5.37, "elapsed_time": "0:51:39", "remaining_time": "15:09:50"}
|
| 45 |
+
{"current_steps": 90, "total_steps": 1638, "loss": 0.8633757829666138, "lr": 1.9973225274686804e-05, "epoch": 0.16483516483516483, "percentage": 5.49, "elapsed_time": "0:52:36", "remaining_time": "15:04:45"}
|
| 46 |
+
{"current_steps": 92, "total_steps": 1638, "loss": 1.3163621425628662, "lr": 1.9970410279846816e-05, "epoch": 0.1684981684981685, "percentage": 5.62, "elapsed_time": "0:53:55", "remaining_time": "15:06:02"}
|
| 47 |
+
{"current_steps": 94, "total_steps": 1638, "loss": 1.3430674076080322, "lr": 1.9967454851584132e-05, "epoch": 0.17216117216117216, "percentage": 5.74, "elapsed_time": "0:55:11", "remaining_time": "15:06:29"}
|
| 48 |
+
{"current_steps": 96, "total_steps": 1638, "loss": 1.232496976852417, "lr": 1.996435903616651e-05, "epoch": 0.17582417582417584, "percentage": 5.86, "elapsed_time": "0:56:11", "remaining_time": "15:02:35"}
|
| 49 |
+
{"current_steps": 98, "total_steps": 1638, "loss": 1.3284939527511597, "lr": 1.9961122882059523e-05, "epoch": 0.1794871794871795, "percentage": 5.98, "elapsed_time": "0:57:29", "remaining_time": "15:03:21"}
|
| 50 |
+
{"current_steps": 100, "total_steps": 1638, "loss": 1.1766855716705322, "lr": 1.9957746439925748e-05, "epoch": 0.18315018315018314, "percentage": 6.11, "elapsed_time": "0:58:43", "remaining_time": "15:03:04"}
|
| 51 |
+
{"current_steps": 102, "total_steps": 1638, "loss": 1.2522631883621216, "lr": 1.9954229762624016e-05, "epoch": 0.18681318681318682, "percentage": 6.23, "elapsed_time": "0:59:51", "remaining_time": "15:01:28"}
|
| 52 |
+
{"current_steps": 104, "total_steps": 1638, "loss": 0.9271856546401978, "lr": 1.995057290520855e-05, "epoch": 0.19047619047619047, "percentage": 6.35, "elapsed_time": "1:01:04", "remaining_time": "15:00:58"}
|
| 53 |
+
{"current_steps": 106, "total_steps": 1638, "loss": 1.0877634286880493, "lr": 1.9946775924928132e-05, "epoch": 0.19413919413919414, "percentage": 6.47, "elapsed_time": "1:02:03", "remaining_time": "14:56:48"}
|
| 54 |
+
{"current_steps": 108, "total_steps": 1638, "loss": 1.315323829650879, "lr": 1.9942838881225183e-05, "epoch": 0.1978021978021978, "percentage": 6.59, "elapsed_time": "1:03:15", "remaining_time": "14:56:07"}
|
| 55 |
+
{"current_steps": 110, "total_steps": 1638, "loss": 1.1636862754821777, "lr": 1.9938761835734842e-05, "epoch": 0.20146520146520147, "percentage": 6.72, "elapsed_time": "1:04:20", "remaining_time": "14:53:43"}
|
| 56 |
+
{"current_steps": 112, "total_steps": 1638, "loss": 1.231845736503601, "lr": 1.9934544852284013e-05, "epoch": 0.20512820512820512, "percentage": 6.84, "elapsed_time": "1:05:39", "remaining_time": "14:54:39"}
|
| 57 |
+
{"current_steps": 114, "total_steps": 1638, "loss": 0.6122523546218872, "lr": 1.9930187996890347e-05, "epoch": 0.2087912087912088, "percentage": 6.96, "elapsed_time": "1:06:47", "remaining_time": "14:52:50"}
|
| 58 |
+
{"current_steps": 116, "total_steps": 1638, "loss": 1.3050415515899658, "lr": 1.992569133776121e-05, "epoch": 0.21245421245421245, "percentage": 7.08, "elapsed_time": "1:08:03", "remaining_time": "14:52:54"}
|
| 59 |
+
{"current_steps": 118, "total_steps": 1638, "loss": 1.3003625869750977, "lr": 1.992105494529264e-05, "epoch": 0.21611721611721613, "percentage": 7.2, "elapsed_time": "1:09:16", "remaining_time": "14:52:20"}
|
| 60 |
+
{"current_steps": 120, "total_steps": 1638, "loss": 1.3665677309036255, "lr": 1.99162788920682e-05, "epoch": 0.21978021978021978, "percentage": 7.33, "elapsed_time": "1:10:11", "remaining_time": "14:47:54"}
|
| 61 |
+
{"current_steps": 122, "total_steps": 1638, "loss": 1.2980903387069702, "lr": 1.9911363252857887e-05, "epoch": 0.22344322344322345, "percentage": 7.45, "elapsed_time": "1:11:27", "remaining_time": "14:48:00"}
|
| 62 |
+
{"current_steps": 124, "total_steps": 1638, "loss": 1.0229507684707642, "lr": 1.990630810461694e-05, "epoch": 0.2271062271062271, "percentage": 7.57, "elapsed_time": "1:12:41", "remaining_time": "14:47:32"}
|
| 63 |
+
{"current_steps": 126, "total_steps": 1638, "loss": 0.9008902311325073, "lr": 1.990111352648463e-05, "epoch": 0.23076923076923078, "percentage": 7.69, "elapsed_time": "1:13:54", "remaining_time": "14:46:52"}
|
| 64 |
+
{"current_steps": 128, "total_steps": 1638, "loss": 1.194158673286438, "lr": 1.9895779599783033e-05, "epoch": 0.23443223443223443, "percentage": 7.81, "elapsed_time": "1:15:09", "remaining_time": "14:46:33"}
|
| 65 |
+
{"current_steps": 130, "total_steps": 1638, "loss": 1.2995827198028564, "lr": 1.989030640801576e-05, "epoch": 0.23809523809523808, "percentage": 7.94, "elapsed_time": "1:16:25", "remaining_time": "14:46:37"}
|
| 66 |
+
{"current_steps": 132, "total_steps": 1638, "loss": 1.3926903009414673, "lr": 1.9884694036866624e-05, "epoch": 0.24175824175824176, "percentage": 8.06, "elapsed_time": "1:17:26", "remaining_time": "14:43:29"}
|
| 67 |
+
{"current_steps": 134, "total_steps": 1638, "loss": 1.2948912382125854, "lr": 1.9878942574198334e-05, "epoch": 0.2454212454212454, "percentage": 8.18, "elapsed_time": "1:18:41", "remaining_time": "14:43:16"}
|
| 68 |
+
{"current_steps": 136, "total_steps": 1638, "loss": 1.2796152830123901, "lr": 1.9873052110051094e-05, "epoch": 0.2490842490842491, "percentage": 8.3, "elapsed_time": "1:20:00", "remaining_time": "14:43:33"}
|
| 69 |
+
{"current_steps": 138, "total_steps": 1638, "loss": 1.0871366262435913, "lr": 1.9867022736641205e-05, "epoch": 0.25274725274725274, "percentage": 8.42, "elapsed_time": "1:21:14", "remaining_time": "14:43:07"}
|
| 70 |
+
{"current_steps": 140, "total_steps": 1638, "loss": 1.2783880233764648, "lr": 1.9860854548359615e-05, "epoch": 0.2564102564102564, "percentage": 8.55, "elapsed_time": "1:22:30", "remaining_time": "14:42:46"}
|
| 71 |
+
{"current_steps": 142, "total_steps": 1638, "loss": 1.2902638912200928, "lr": 1.9854547641770446e-05, "epoch": 0.2600732600732601, "percentage": 8.67, "elapsed_time": "1:23:36", "remaining_time": "14:40:51"}
|
| 72 |
+
{"current_steps": 144, "total_steps": 1638, "loss": 1.2590699195861816, "lr": 1.9848102115609483e-05, "epoch": 0.26373626373626374, "percentage": 8.79, "elapsed_time": "1:24:52", "remaining_time": "14:40:39"}
|
| 73 |
+
{"current_steps": 146, "total_steps": 1638, "loss": 1.407915711402893, "lr": 1.9841518070782615e-05, "epoch": 0.2673992673992674, "percentage": 8.91, "elapsed_time": "1:26:06", "remaining_time": "14:40:02"}
|
| 74 |
+
{"current_steps": 148, "total_steps": 1638, "loss": 1.316453456878662, "lr": 1.983479561036429e-05, "epoch": 0.27106227106227104, "percentage": 9.04, "elapsed_time": "1:27:19", "remaining_time": "14:39:07"}
|
| 75 |
+
{"current_steps": 150, "total_steps": 1638, "loss": 0.9313694834709167, "lr": 1.982793483959585e-05, "epoch": 0.27472527472527475, "percentage": 9.16, "elapsed_time": "1:28:18", "remaining_time": "14:36:04"}
|
| 76 |
+
{"current_steps": 152, "total_steps": 1638, "loss": 0.6459367871284485, "lr": 1.9820935865883924e-05, "epoch": 0.2783882783882784, "percentage": 9.28, "elapsed_time": "1:29:01", "remaining_time": "14:30:24"}
|
| 77 |
+
{"current_steps": 154, "total_steps": 1638, "loss": 1.0998259782791138, "lr": 1.981379879879874e-05, "epoch": 0.28205128205128205, "percentage": 9.4, "elapsed_time": "1:29:46", "remaining_time": "14:25:06"}
|
| 78 |
+
{"current_steps": 156, "total_steps": 1638, "loss": 1.3427971601486206, "lr": 1.9806523750072385e-05, "epoch": 0.2857142857142857, "percentage": 9.52, "elapsed_time": "1:31:03", "remaining_time": "14:24:59"}
|
| 79 |
+
{"current_steps": 158, "total_steps": 1638, "loss": 1.2781323194503784, "lr": 1.9799110833597093e-05, "epoch": 0.2893772893772894, "percentage": 9.65, "elapsed_time": "1:32:21", "remaining_time": "14:25:04"}
|
| 80 |
+
{"current_steps": 160, "total_steps": 1638, "loss": 0.9342338442802429, "lr": 1.9791560165423433e-05, "epoch": 0.29304029304029305, "percentage": 9.77, "elapsed_time": "1:33:33", "remaining_time": "14:24:19"}
|
| 81 |
+
{"current_steps": 162, "total_steps": 1638, "loss": 1.5323666334152222, "lr": 1.9783871863758503e-05, "epoch": 0.2967032967032967, "percentage": 9.89, "elapsed_time": "1:34:48", "remaining_time": "14:23:49"}
|
| 82 |
+
{"current_steps": 164, "total_steps": 1638, "loss": 1.0420787334442139, "lr": 1.9776046048964082e-05, "epoch": 0.30036630036630035, "percentage": 10.01, "elapsed_time": "1:36:01", "remaining_time": "14:23:04"}
|
| 83 |
+
{"current_steps": 166, "total_steps": 1638, "loss": 1.3945231437683105, "lr": 1.9768082843554737e-05, "epoch": 0.304029304029304, "percentage": 10.13, "elapsed_time": "1:37:18", "remaining_time": "14:22:55"}
|
| 84 |
+
{"current_steps": 168, "total_steps": 1638, "loss": 1.1299034357070923, "lr": 1.9759982372195918e-05, "epoch": 0.3076923076923077, "percentage": 10.26, "elapsed_time": "1:38:31", "remaining_time": "14:22:07"}
|
| 85 |
+
{"current_steps": 170, "total_steps": 1638, "loss": 1.260428786277771, "lr": 1.9751744761701984e-05, "epoch": 0.31135531135531136, "percentage": 10.38, "elapsed_time": "1:39:40", "remaining_time": "14:20:46"}
|
| 86 |
+
{"current_steps": 172, "total_steps": 1638, "loss": 1.006029725074768, "lr": 1.9743370141034248e-05, "epoch": 0.315018315018315, "percentage": 10.5, "elapsed_time": "1:40:54", "remaining_time": "14:20:05"}
|
| 87 |
+
{"current_steps": 174, "total_steps": 1638, "loss": 0.8638060688972473, "lr": 1.973485864129894e-05, "epoch": 0.31868131868131866, "percentage": 10.62, "elapsed_time": "1:41:55", "remaining_time": "14:17:34"}
|
| 88 |
+
{"current_steps": 176, "total_steps": 1638, "loss": 1.397803783416748, "lr": 1.9726210395745148e-05, "epoch": 0.32234432234432236, "percentage": 10.74, "elapsed_time": "1:43:16", "remaining_time": "14:17:53"}
|
| 89 |
+
{"current_steps": 178, "total_steps": 1638, "loss": 0.921164870262146, "lr": 1.971742553976275e-05, "epoch": 0.326007326007326, "percentage": 10.87, "elapsed_time": "1:44:24", "remaining_time": "14:16:25"}
|
| 90 |
+
{"current_steps": 180, "total_steps": 1638, "loss": 1.4893336296081543, "lr": 1.9708504210880284e-05, "epoch": 0.32967032967032966, "percentage": 10.99, "elapsed_time": "1:45:37", "remaining_time": "14:15:33"}
|
| 91 |
+
{"current_steps": 182, "total_steps": 1638, "loss": 0.977988064289093, "lr": 1.969944654876279e-05, "epoch": 0.3333333333333333, "percentage": 11.11, "elapsed_time": "1:46:49", "remaining_time": "14:14:32"}
|
| 92 |
+
{"current_steps": 184, "total_steps": 1638, "loss": 1.2434574365615845, "lr": 1.9690252695209636e-05, "epoch": 0.336996336996337, "percentage": 11.23, "elapsed_time": "1:48:07", "remaining_time": "14:14:22"}
|
| 93 |
+
{"current_steps": 186, "total_steps": 1638, "loss": 1.3150489330291748, "lr": 1.9680922794152294e-05, "epoch": 0.34065934065934067, "percentage": 11.36, "elapsed_time": "1:49:06", "remaining_time": "14:11:44"}
|
| 94 |
+
{"current_steps": 188, "total_steps": 1638, "loss": 1.170944333076477, "lr": 1.9671456991652072e-05, "epoch": 0.3443223443223443, "percentage": 11.48, "elapsed_time": "1:50:26", "remaining_time": "14:11:51"}
|
| 95 |
+
{"current_steps": 190, "total_steps": 1638, "loss": 1.2700278759002686, "lr": 1.9661855435897858e-05, "epoch": 0.34798534798534797, "percentage": 11.6, "elapsed_time": "1:51:45", "remaining_time": "14:11:40"}
|
| 96 |
+
{"current_steps": 192, "total_steps": 1638, "loss": 1.1048038005828857, "lr": 1.9652118277203767e-05, "epoch": 0.3516483516483517, "percentage": 11.72, "elapsed_time": "1:52:54", "remaining_time": "14:10:19"}
|
| 97 |
+
{"current_steps": 194, "total_steps": 1638, "loss": 1.2465182542800903, "lr": 1.9642245668006814e-05, "epoch": 0.3553113553113553, "percentage": 11.84, "elapsed_time": "1:54:07", "remaining_time": "14:09:27"}
|
| 98 |
+
{"current_steps": 196, "total_steps": 1638, "loss": 1.2548701763153076, "lr": 1.963223776286451e-05, "epoch": 0.358974358974359, "percentage": 11.97, "elapsed_time": "1:55:21", "remaining_time": "14:08:40"}
|
| 99 |
+
{"current_steps": 198, "total_steps": 1638, "loss": 0.8607481718063354, "lr": 1.9622094718452448e-05, "epoch": 0.3626373626373626, "percentage": 12.09, "elapsed_time": "1:56:29", "remaining_time": "14:07:13"}
|
| 100 |
+
{"current_steps": 200, "total_steps": 1638, "loss": 1.0138609409332275, "lr": 1.9611816693561858e-05, "epoch": 0.3663003663003663, "percentage": 12.21, "elapsed_time": "1:57:43", "remaining_time": "14:06:23"}
|
| 101 |
+
{"current_steps": 202, "total_steps": 1638, "loss": 1.4290412664413452, "lr": 1.96014038490971e-05, "epoch": 0.36996336996337, "percentage": 12.33, "elapsed_time": "1:58:57", "remaining_time": "14:05:38"}
|
| 102 |
+
{"current_steps": 204, "total_steps": 1638, "loss": 1.2090134620666504, "lr": 1.9590856348073182e-05, "epoch": 0.37362637362637363, "percentage": 12.45, "elapsed_time": "2:00:08", "remaining_time": "14:04:30"}
|
| 103 |
+
{"current_steps": 206, "total_steps": 1638, "loss": 0.7611775398254395, "lr": 1.9580174355613168e-05, "epoch": 0.3772893772893773, "percentage": 12.58, "elapsed_time": "2:01:04", "remaining_time": "14:01:37"}
|
| 104 |
+
{"current_steps": 208, "total_steps": 1638, "loss": 1.1324646472930908, "lr": 1.9569358038945617e-05, "epoch": 0.38095238095238093, "percentage": 12.7, "elapsed_time": "2:01:55", "remaining_time": "13:58:11"}
|
| 105 |
+
{"current_steps": 210, "total_steps": 1638, "loss": 1.4070419073104858, "lr": 1.9558407567401945e-05, "epoch": 0.38461538461538464, "percentage": 12.82, "elapsed_time": "2:03:05", "remaining_time": "13:57:00"}
|
| 106 |
+
{"current_steps": 212, "total_steps": 1638, "loss": 1.071973204612732, "lr": 1.9547323112413806e-05, "epoch": 0.3882783882783883, "percentage": 12.94, "elapsed_time": "2:04:21", "remaining_time": "13:56:29"}
|
| 107 |
+
{"current_steps": 214, "total_steps": 1638, "loss": 1.1344265937805176, "lr": 1.9536104847510384e-05, "epoch": 0.39194139194139194, "percentage": 13.06, "elapsed_time": "2:05:29", "remaining_time": "13:55:01"}
|
| 108 |
+
{"current_steps": 216, "total_steps": 1638, "loss": 1.2220566272735596, "lr": 1.9524752948315677e-05, "epoch": 0.3956043956043956, "percentage": 13.19, "elapsed_time": "2:06:43", "remaining_time": "13:54:14"}
|
| 109 |
+
{"current_steps": 218, "total_steps": 1638, "loss": 1.2576326131820679, "lr": 1.9513267592545752e-05, "epoch": 0.3992673992673993, "percentage": 13.31, "elapsed_time": "2:08:00", "remaining_time": "13:53:46"}
|
| 110 |
+
{"current_steps": 220, "total_steps": 1638, "loss": 0.6170648336410522, "lr": 1.9501648960005964e-05, "epoch": 0.40293040293040294, "percentage": 13.43, "elapsed_time": "2:09:09", "remaining_time": "13:52:26"}
|
| 111 |
+
{"current_steps": 222, "total_steps": 1638, "loss": 1.339916706085205, "lr": 1.948989723258815e-05, "epoch": 0.4065934065934066, "percentage": 13.55, "elapsed_time": "2:10:23", "remaining_time": "13:51:42"}
|
| 112 |
+
{"current_steps": 224, "total_steps": 1638, "loss": 1.077911615371704, "lr": 1.9478012594267757e-05, "epoch": 0.41025641025641024, "percentage": 13.68, "elapsed_time": "2:11:23", "remaining_time": "13:49:23"}
|
| 113 |
+
{"current_steps": 226, "total_steps": 1638, "loss": 1.2358698844909668, "lr": 1.946599523110099e-05, "epoch": 0.4139194139194139, "percentage": 13.8, "elapsed_time": "2:12:33", "remaining_time": "13:48:13"}
|
| 114 |
+
{"current_steps": 228, "total_steps": 1638, "loss": 1.3055531978607178, "lr": 1.945384533122187e-05, "epoch": 0.4175824175824176, "percentage": 13.92, "elapsed_time": "2:13:48", "remaining_time": "13:47:26"}
|
| 115 |
+
{"current_steps": 230, "total_steps": 1638, "loss": 1.2255327701568604, "lr": 1.9441563084839324e-05, "epoch": 0.42124542124542125, "percentage": 14.04, "elapsed_time": "2:14:52", "remaining_time": "13:45:39"}
|
| 116 |
+
{"current_steps": 232, "total_steps": 1638, "loss": 0.9998282790184021, "lr": 1.942914868423417e-05, "epoch": 0.4249084249084249, "percentage": 14.16, "elapsed_time": "2:15:53", "remaining_time": "13:43:34"}
|
| 117 |
+
{"current_steps": 234, "total_steps": 1638, "loss": 1.4894115924835205, "lr": 1.941660232375614e-05, "epoch": 0.42857142857142855, "percentage": 14.29, "elapsed_time": "2:17:07", "remaining_time": "13:42:42"}
|
| 118 |
+
{"current_steps": 236, "total_steps": 1638, "loss": 1.0169166326522827, "lr": 1.9403924199820813e-05, "epoch": 0.43223443223443225, "percentage": 14.41, "elapsed_time": "2:18:15", "remaining_time": "13:41:20"}
|
| 119 |
+
{"current_steps": 238, "total_steps": 1638, "loss": 1.0664693117141724, "lr": 1.9391114510906546e-05, "epoch": 0.4358974358974359, "percentage": 14.53, "elapsed_time": "2:19:25", "remaining_time": "13:40:11"}
|
| 120 |
+
{"current_steps": 240, "total_steps": 1638, "loss": 0.9052911996841431, "lr": 1.937817345755138e-05, "epoch": 0.43956043956043955, "percentage": 14.65, "elapsed_time": "2:20:35", "remaining_time": "13:38:56"}
|
| 121 |
+
{"current_steps": 242, "total_steps": 1638, "loss": 0.866486668586731, "lr": 1.9365101242349883e-05, "epoch": 0.4432234432234432, "percentage": 14.77, "elapsed_time": "2:21:45", "remaining_time": "13:37:45"}
|
| 122 |
+
{"current_steps": 244, "total_steps": 1638, "loss": 0.5777208209037781, "lr": 1.9351898069949985e-05, "epoch": 0.4468864468864469, "percentage": 14.9, "elapsed_time": "2:22:49", "remaining_time": "13:35:59"}
|
| 123 |
+
{"current_steps": 246, "total_steps": 1638, "loss": 1.2593415975570679, "lr": 1.9338564147049785e-05, "epoch": 0.45054945054945056, "percentage": 15.02, "elapsed_time": "2:24:01", "remaining_time": "13:34:56"}
|
| 124 |
+
{"current_steps": 248, "total_steps": 1638, "loss": 0.8774123787879944, "lr": 1.9325099682394296e-05, "epoch": 0.4542124542124542, "percentage": 15.14, "elapsed_time": "2:25:08", "remaining_time": "13:33:30"}
|
| 125 |
+
{"current_steps": 250, "total_steps": 1638, "loss": 1.2584809064865112, "lr": 1.9311504886772183e-05, "epoch": 0.45787545787545786, "percentage": 15.26, "elapsed_time": "2:26:05", "remaining_time": "13:31:07"}
|
| 126 |
+
{"current_steps": 252, "total_steps": 1638, "loss": 1.1775099039077759, "lr": 1.929777997301248e-05, "epoch": 0.46153846153846156, "percentage": 15.38, "elapsed_time": "2:27:19", "remaining_time": "13:30:15"}
|
| 127 |
+
{"current_steps": 254, "total_steps": 1638, "loss": 0.9682942628860474, "lr": 1.9283925155981228e-05, "epoch": 0.4652014652014652, "percentage": 15.51, "elapsed_time": "2:28:32", "remaining_time": "13:29:23"}
|
| 128 |
+
{"current_steps": 256, "total_steps": 1638, "loss": 1.26495361328125, "lr": 1.9269940652578143e-05, "epoch": 0.46886446886446886, "percentage": 15.63, "elapsed_time": "2:29:48", "remaining_time": "13:28:45"}
|
| 129 |
+
{"current_steps": 258, "total_steps": 1638, "loss": 1.286316990852356, "lr": 1.9255826681733194e-05, "epoch": 0.4725274725274725, "percentage": 15.75, "elapsed_time": "2:31:05", "remaining_time": "13:28:10"}
|
| 130 |
+
{"current_steps": 260, "total_steps": 1638, "loss": 0.7573358416557312, "lr": 1.924158346440319e-05, "epoch": 0.47619047619047616, "percentage": 15.87, "elapsed_time": "2:32:01", "remaining_time": "13:25:44"}
|
| 131 |
+
{"current_steps": 262, "total_steps": 1638, "loss": 1.148931622505188, "lr": 1.9227211223568317e-05, "epoch": 0.47985347985347987, "percentage": 16.0, "elapsed_time": "2:33:06", "remaining_time": "13:24:06"}
|
| 132 |
+
{"current_steps": 264, "total_steps": 1638, "loss": 1.2336255311965942, "lr": 1.9212710184228654e-05, "epoch": 0.4835164835164835, "percentage": 16.12, "elapsed_time": "2:34:22", "remaining_time": "13:23:26"}
|
| 133 |
+
{"current_steps": 266, "total_steps": 1638, "loss": 1.503554105758667, "lr": 1.9198080573400634e-05, "epoch": 0.48717948717948717, "percentage": 16.24, "elapsed_time": "2:35:36", "remaining_time": "13:22:35"}
|
| 134 |
+
{"current_steps": 268, "total_steps": 1638, "loss": 0.7954114675521851, "lr": 1.9183322620113505e-05, "epoch": 0.4908424908424908, "percentage": 16.36, "elapsed_time": "2:36:44", "remaining_time": "13:21:16"}
|
| 135 |
+
{"current_steps": 270, "total_steps": 1638, "loss": 1.2086211442947388, "lr": 1.916843655540574e-05, "epoch": 0.4945054945054945, "percentage": 16.48, "elapsed_time": "2:38:03", "remaining_time": "13:20:48"}
|
| 136 |
+
{"current_steps": 272, "total_steps": 1638, "loss": 0.8882235884666443, "lr": 1.915342261232142e-05, "epoch": 0.4981684981684982, "percentage": 16.61, "elapsed_time": "2:39:13", "remaining_time": "13:19:37"}
|
| 137 |
+
{"current_steps": 274, "total_steps": 1638, "loss": 1.2472962141036987, "lr": 1.913828102590659e-05, "epoch": 0.5018315018315018, "percentage": 16.73, "elapsed_time": "2:40:20", "remaining_time": "13:18:12"}
|
| 138 |
+
{"current_steps": 276, "total_steps": 1638, "loss": 0.8005316853523254, "lr": 1.9123012033205564e-05, "epoch": 0.5054945054945055, "percentage": 16.85, "elapsed_time": "2:41:21", "remaining_time": "13:16:15"}
|
| 139 |
+
{"current_steps": 278, "total_steps": 1638, "loss": 0.8836736083030701, "lr": 1.9107615873257234e-05, "epoch": 0.5091575091575091, "percentage": 16.97, "elapsed_time": "2:42:18", "remaining_time": "13:14:01"}
|
| 140 |
+
{"current_steps": 280, "total_steps": 1638, "loss": 1.2598307132720947, "lr": 1.909209278709131e-05, "epoch": 0.5128205128205128, "percentage": 17.09, "elapsed_time": "2:43:35", "remaining_time": "13:13:25"}
|
| 141 |
+
{"current_steps": 282, "total_steps": 1638, "loss": 1.2541738748550415, "lr": 1.9076443017724568e-05, "epoch": 0.5164835164835165, "percentage": 17.22, "elapsed_time": "2:44:49", "remaining_time": "13:12:32"}
|
| 142 |
+
{"current_steps": 284, "total_steps": 1638, "loss": 1.2553168535232544, "lr": 1.9060666810157025e-05, "epoch": 0.5201465201465202, "percentage": 17.34, "elapsed_time": "2:46:05", "remaining_time": "13:11:50"}
|
| 143 |
+
{"current_steps": 286, "total_steps": 1638, "loss": 1.0212312936782837, "lr": 1.9044764411368106e-05, "epoch": 0.5238095238095238, "percentage": 17.46, "elapsed_time": "2:47:18", "remaining_time": "13:10:56"}
|
| 144 |
+
{"current_steps": 288, "total_steps": 1638, "loss": 1.262040615081787, "lr": 1.9028736070312796e-05, "epoch": 0.5274725274725275, "percentage": 17.58, "elapsed_time": "2:48:35", "remaining_time": "13:10:15"}
|
| 145 |
+
{"current_steps": 290, "total_steps": 1638, "loss": 1.2195689678192139, "lr": 1.9012582037917713e-05, "epoch": 0.5311355311355311, "percentage": 17.7, "elapsed_time": "2:49:51", "remaining_time": "13:09:32"}
|
| 146 |
+
{"current_steps": 292, "total_steps": 1638, "loss": 0.7313263416290283, "lr": 1.8996302567077217e-05, "epoch": 0.5347985347985348, "percentage": 17.83, "elapsed_time": "2:50:48", "remaining_time": "13:07:23"}
|
| 147 |
+
{"current_steps": 294, "total_steps": 1638, "loss": 0.9493017792701721, "lr": 1.897989791264941e-05, "epoch": 0.5384615384615384, "percentage": 17.95, "elapsed_time": "2:51:46", "remaining_time": "13:05:16"}
|
| 148 |
+
{"current_steps": 296, "total_steps": 1638, "loss": 1.028235673904419, "lr": 1.8963368331452172e-05, "epoch": 0.5421245421245421, "percentage": 18.07, "elapsed_time": "2:52:48", "remaining_time": "13:03:28"}
|
| 149 |
+
{"current_steps": 298, "total_steps": 1638, "loss": 1.3035478591918945, "lr": 1.8946714082259145e-05, "epoch": 0.5457875457875457, "percentage": 18.19, "elapsed_time": "2:54:01", "remaining_time": "13:02:30"}
|
| 150 |
+
{"current_steps": 300, "total_steps": 1638, "loss": 1.2072572708129883, "lr": 1.8929935425795655e-05, "epoch": 0.5494505494505495, "percentage": 18.32, "elapsed_time": "2:55:06", "remaining_time": "13:00:57"}
|
| 151 |
+
{"current_steps": 302, "total_steps": 1638, "loss": 1.192374587059021, "lr": 1.8913032624734657e-05, "epoch": 0.5531135531135531, "percentage": 18.44, "elapsed_time": "2:56:22", "remaining_time": "13:00:15"}
|
| 152 |
+
{"current_steps": 304, "total_steps": 1638, "loss": 0.9877452850341797, "lr": 1.8896005943692614e-05, "epoch": 0.5567765567765568, "percentage": 18.56, "elapsed_time": "2:57:27", "remaining_time": "12:58:44"}
|
| 153 |
+
{"current_steps": 306, "total_steps": 1638, "loss": 0.966866672039032, "lr": 1.8878855649225346e-05, "epoch": 0.5604395604395604, "percentage": 18.68, "elapsed_time": "2:58:27", "remaining_time": "12:56:49"}
|
| 154 |
+
{"current_steps": 308, "total_steps": 1638, "loss": 1.418901801109314, "lr": 1.8861582009823868e-05, "epoch": 0.5641025641025641, "percentage": 18.8, "elapsed_time": "2:59:35", "remaining_time": "12:55:28"}
|
| 155 |
+
{"current_steps": 310, "total_steps": 1638, "loss": 0.9967933297157288, "lr": 1.884418529591018e-05, "epoch": 0.5677655677655677, "percentage": 18.93, "elapsed_time": "3:00:49", "remaining_time": "12:54:37"}
|
| 156 |
+
{"current_steps": 312, "total_steps": 1638, "loss": 1.213472604751587, "lr": 1.882666577983304e-05, "epoch": 0.5714285714285714, "percentage": 19.05, "elapsed_time": "3:01:39", "remaining_time": "12:52:04"}
|
| 157 |
+
{"current_steps": 314, "total_steps": 1638, "loss": 1.145321249961853, "lr": 1.8809023735863693e-05, "epoch": 0.575091575091575, "percentage": 19.17, "elapsed_time": "3:02:55", "remaining_time": "12:51:18"}
|
| 158 |
+
{"current_steps": 316, "total_steps": 1638, "loss": 1.2913143634796143, "lr": 1.879125944019158e-05, "epoch": 0.5787545787545788, "percentage": 19.29, "elapsed_time": "3:03:57", "remaining_time": "12:49:34"}
|
| 159 |
+
{"current_steps": 318, "total_steps": 1638, "loss": 1.129797339439392, "lr": 1.8773373170920022e-05, "epoch": 0.5824175824175825, "percentage": 19.41, "elapsed_time": "3:05:08", "remaining_time": "12:48:31"}
|
| 160 |
+
{"current_steps": 320, "total_steps": 1638, "loss": 1.345680832862854, "lr": 1.875536520806185e-05, "epoch": 0.5860805860805861, "percentage": 19.54, "elapsed_time": "3:06:20", "remaining_time": "12:47:27"}
|
| 161 |
+
{"current_steps": 322, "total_steps": 1638, "loss": 1.532901406288147, "lr": 1.8737235833535033e-05, "epoch": 0.5897435897435898, "percentage": 19.66, "elapsed_time": "3:07:33", "remaining_time": "12:46:32"}
|
| 162 |
+
{"current_steps": 324, "total_steps": 1638, "loss": 1.2808473110198975, "lr": 1.871898533115827e-05, "epoch": 0.5934065934065934, "percentage": 19.78, "elapsed_time": "3:08:51", "remaining_time": "12:45:53"}
|
| 163 |
+
{"current_steps": 326, "total_steps": 1638, "loss": 1.3697587251663208, "lr": 1.870061398664653e-05, "epoch": 0.5970695970695971, "percentage": 19.9, "elapsed_time": "3:09:58", "remaining_time": "12:44:34"}
|
| 164 |
+
{"current_steps": 328, "total_steps": 1638, "loss": 1.2338076829910278, "lr": 1.868212208760658e-05, "epoch": 0.6007326007326007, "percentage": 20.02, "elapsed_time": "3:11:15", "remaining_time": "12:43:52"}
|
| 165 |
+
{"current_steps": 330, "total_steps": 1638, "loss": 1.113355040550232, "lr": 1.8663509923532514e-05, "epoch": 0.6043956043956044, "percentage": 20.15, "elapsed_time": "3:12:27", "remaining_time": "12:42:51"}
|
| 166 |
+
{"current_steps": 332, "total_steps": 1638, "loss": 1.1921931505203247, "lr": 1.8644777785801175e-05, "epoch": 0.608058608058608, "percentage": 20.27, "elapsed_time": "3:13:45", "remaining_time": "12:42:12"}
|
| 167 |
+
{"current_steps": 334, "total_steps": 1638, "loss": 1.287142038345337, "lr": 1.862592596766763e-05, "epoch": 0.6117216117216118, "percentage": 20.39, "elapsed_time": "3:15:01", "remaining_time": "12:41:23"}
|
| 168 |
+
{"current_steps": 336, "total_steps": 1638, "loss": 0.9066182374954224, "lr": 1.8606954764260556e-05, "epoch": 0.6153846153846154, "percentage": 20.51, "elapsed_time": "3:15:59", "remaining_time": "12:39:27"}
|
| 169 |
+
{"current_steps": 338, "total_steps": 1638, "loss": 1.240350604057312, "lr": 1.8587864472577632e-05, "epoch": 0.6190476190476191, "percentage": 20.63, "elapsed_time": "3:17:13", "remaining_time": "12:38:31"}
|
| 170 |
+
{"current_steps": 340, "total_steps": 1638, "loss": 1.2407909631729126, "lr": 1.8568655391480882e-05, "epoch": 0.6227106227106227, "percentage": 20.76, "elapsed_time": "3:18:29", "remaining_time": "12:37:46"}
|
| 171 |
+
{"current_steps": 342, "total_steps": 1638, "loss": 0.5828521251678467, "lr": 1.8549327821692008e-05, "epoch": 0.6263736263736264, "percentage": 20.88, "elapsed_time": "3:19:37", "remaining_time": "12:36:26"}
|
| 172 |
+
{"current_steps": 344, "total_steps": 1638, "loss": 1.444503903388977, "lr": 1.852988206578767e-05, "epoch": 0.63003663003663, "percentage": 21.0, "elapsed_time": "3:20:52", "remaining_time": "12:35:38"}
|
| 173 |
+
{"current_steps": 346, "total_steps": 1638, "loss": 0.6921512484550476, "lr": 1.851031842819475e-05, "epoch": 0.6336996336996337, "percentage": 21.12, "elapsed_time": "3:22:01", "remaining_time": "12:34:22"}
|
| 174 |
+
{"current_steps": 348, "total_steps": 1638, "loss": 1.1690187454223633, "lr": 1.849063721518559e-05, "epoch": 0.6373626373626373, "percentage": 21.25, "elapsed_time": "3:23:05", "remaining_time": "12:32:49"}
|
| 175 |
+
{"current_steps": 350, "total_steps": 1638, "loss": 0.8881887197494507, "lr": 1.8470838734873205e-05, "epoch": 0.6410256410256411, "percentage": 21.37, "elapsed_time": "3:24:17", "remaining_time": "12:31:45"}
|
| 176 |
+
{"current_steps": 352, "total_steps": 1638, "loss": 0.9233137965202332, "lr": 1.8450923297206446e-05, "epoch": 0.6446886446886447, "percentage": 21.49, "elapsed_time": "3:25:25", "remaining_time": "12:30:28"}
|
| 177 |
+
{"current_steps": 354, "total_steps": 1638, "loss": 0.954558253288269, "lr": 1.8430891213965146e-05, "epoch": 0.6483516483516484, "percentage": 21.61, "elapsed_time": "3:26:36", "remaining_time": "12:29:23"}
|
| 178 |
+
{"current_steps": 356, "total_steps": 1638, "loss": 1.1792762279510498, "lr": 1.8410742798755255e-05, "epoch": 0.652014652014652, "percentage": 21.73, "elapsed_time": "3:27:43", "remaining_time": "12:28:03"}
|
| 179 |
+
{"current_steps": 358, "total_steps": 1638, "loss": 1.1631232500076294, "lr": 1.8390478367003922e-05, "epoch": 0.6556776556776557, "percentage": 21.86, "elapsed_time": "3:28:53", "remaining_time": "12:26:51"}
|
| 180 |
+
{"current_steps": 360, "total_steps": 1638, "loss": 0.6956652998924255, "lr": 1.8370098235954553e-05, "epoch": 0.6593406593406593, "percentage": 21.98, "elapsed_time": "3:30:02", "remaining_time": "12:25:39"}
|
| 181 |
+
{"current_steps": 362, "total_steps": 1638, "loss": 0.9578306078910828, "lr": 1.834960272466184e-05, "epoch": 0.663003663003663, "percentage": 22.1, "elapsed_time": "3:31:15", "remaining_time": "12:24:39"}
|
| 182 |
+
{"current_steps": 364, "total_steps": 1638, "loss": 0.9421581625938416, "lr": 1.832899215398679e-05, "epoch": 0.6666666666666666, "percentage": 22.22, "elapsed_time": "3:32:17", "remaining_time": "12:23:01"}
|
| 183 |
+
{"current_steps": 366, "total_steps": 1638, "loss": 1.2012219429016113, "lr": 1.8308266846591673e-05, "epoch": 0.6703296703296703, "percentage": 22.34, "elapsed_time": "3:33:33", "remaining_time": "12:22:12"}
|
| 184 |
+
{"current_steps": 368, "total_steps": 1638, "loss": 1.065047264099121, "lr": 1.828742712693499e-05, "epoch": 0.673992673992674, "percentage": 22.47, "elapsed_time": "3:34:48", "remaining_time": "12:21:20"}
|
| 185 |
+
{"current_steps": 370, "total_steps": 1638, "loss": 1.1004585027694702, "lr": 1.8266473321266385e-05, "epoch": 0.6776556776556777, "percentage": 22.59, "elapsed_time": "3:35:49", "remaining_time": "12:19:37"}
|
| 186 |
+
{"current_steps": 372, "total_steps": 1638, "loss": 1.1929394006729126, "lr": 1.824540575762154e-05, "epoch": 0.6813186813186813, "percentage": 22.71, "elapsed_time": "3:36:47", "remaining_time": "12:17:47"}
|
| 187 |
+
{"current_steps": 374, "total_steps": 1638, "loss": 1.217964768409729, "lr": 1.8224224765817033e-05, "epoch": 0.684981684981685, "percentage": 22.83, "elapsed_time": "3:38:04", "remaining_time": "12:16:59"}
|
| 188 |
+
{"current_steps": 376, "total_steps": 1638, "loss": 0.8868032097816467, "lr": 1.820293067744519e-05, "epoch": 0.6886446886446886, "percentage": 22.95, "elapsed_time": "3:39:02", "remaining_time": "12:15:12"}
|
| 189 |
+
{"current_steps": 378, "total_steps": 1638, "loss": 0.8352534174919128, "lr": 1.8181523825868882e-05, "epoch": 0.6923076923076923, "percentage": 23.08, "elapsed_time": "3:40:11", "remaining_time": "12:13:59"}
|
| 190 |
+
{"current_steps": 380, "total_steps": 1638, "loss": 1.06725013256073, "lr": 1.816000454621631e-05, "epoch": 0.6959706959706959, "percentage": 23.2, "elapsed_time": "3:41:23", "remaining_time": "12:12:56"}
|
| 191 |
+
{"current_steps": 382, "total_steps": 1638, "loss": 0.9851567149162292, "lr": 1.8138373175375744e-05, "epoch": 0.6996336996336996, "percentage": 23.32, "elapsed_time": "3:42:36", "remaining_time": "12:11:56"}
|
| 192 |
+
{"current_steps": 384, "total_steps": 1638, "loss": 1.1879215240478516, "lr": 1.8116630051990283e-05, "epoch": 0.7032967032967034, "percentage": 23.44, "elapsed_time": "3:43:50", "remaining_time": "12:10:57"}
|
| 193 |
+
{"current_steps": 386, "total_steps": 1638, "loss": 1.1000186204910278, "lr": 1.8094775516452522e-05, "epoch": 0.706959706959707, "percentage": 23.57, "elapsed_time": "3:45:00", "remaining_time": "12:09:47"}
|
| 194 |
+
{"current_steps": 388, "total_steps": 1638, "loss": 0.8919756412506104, "lr": 1.807280991089923e-05, "epoch": 0.7106227106227107, "percentage": 23.69, "elapsed_time": "3:46:12", "remaining_time": "12:08:46"}
|
| 195 |
+
{"current_steps": 390, "total_steps": 1638, "loss": 1.113328456878662, "lr": 1.8050733579206005e-05, "epoch": 0.7142857142857143, "percentage": 23.81, "elapsed_time": "3:47:26", "remaining_time": "12:07:47"}
|
| 196 |
+
{"current_steps": 392, "total_steps": 1638, "loss": 1.1803910732269287, "lr": 1.8028546866981875e-05, "epoch": 0.717948717948718, "percentage": 23.93, "elapsed_time": "3:48:37", "remaining_time": "12:06:41"}
|
| 197 |
+
{"current_steps": 394, "total_steps": 1638, "loss": 1.1312916278839111, "lr": 1.8006250121563903e-05, "epoch": 0.7216117216117216, "percentage": 24.05, "elapsed_time": "3:49:47", "remaining_time": "12:05:33"}
|
| 198 |
+
{"current_steps": 396, "total_steps": 1638, "loss": 1.2498586177825928, "lr": 1.798384369201174e-05, "epoch": 0.7252747252747253, "percentage": 24.18, "elapsed_time": "3:51:05", "remaining_time": "12:04:48"}
|
| 199 |
+
{"current_steps": 398, "total_steps": 1638, "loss": 0.92662513256073, "lr": 1.796132792910216e-05, "epoch": 0.7289377289377289, "percentage": 24.3, "elapsed_time": "3:52:06", "remaining_time": "12:03:09"}
|
| 200 |
+
{"current_steps": 400, "total_steps": 1638, "loss": 0.8566319942474365, "lr": 1.7938703185323575e-05, "epoch": 0.7326007326007326, "percentage": 24.42, "elapsed_time": "3:53:19", "remaining_time": "12:02:09"}
|
| 201 |
+
{"current_steps": 402, "total_steps": 1638, "loss": 1.2503591775894165, "lr": 1.7915969814870508e-05, "epoch": 0.7362637362637363, "percentage": 24.54, "elapsed_time": "3:54:31", "remaining_time": "12:01:05"}
|
| 202 |
+
{"current_steps": 404, "total_steps": 1638, "loss": 0.851728081703186, "lr": 1.789312817363805e-05, "epoch": 0.73992673992674, "percentage": 24.66, "elapsed_time": "3:55:38", "remaining_time": "11:59:45"}
|
| 203 |
+
{"current_steps": 406, "total_steps": 1638, "loss": 1.0317764282226562, "lr": 1.7870178619216304e-05, "epoch": 0.7435897435897436, "percentage": 24.79, "elapsed_time": "3:56:48", "remaining_time": "11:58:35"}
|
| 204 |
+
{"current_steps": 408, "total_steps": 1638, "loss": 1.029388666152954, "lr": 1.784712151088476e-05, "epoch": 0.7472527472527473, "percentage": 24.91, "elapsed_time": "3:57:57", "remaining_time": "11:57:22"}
|
| 205 |
+
{"current_steps": 410, "total_steps": 1638, "loss": 0.8808343410491943, "lr": 1.782395720960669e-05, "epoch": 0.7509157509157509, "percentage": 25.03, "elapsed_time": "3:58:44", "remaining_time": "11:55:03"}
|
| 206 |
+
{"current_steps": 412, "total_steps": 1638, "loss": 1.1766679286956787, "lr": 1.780068607802349e-05, "epoch": 0.7545787545787546, "percentage": 25.15, "elapsed_time": "4:00:00", "remaining_time": "11:54:10"}
|
| 207 |
+
{"current_steps": 414, "total_steps": 1638, "loss": 1.0107574462890625, "lr": 1.7777308480449006e-05, "epoch": 0.7582417582417582, "percentage": 25.27, "elapsed_time": "4:01:09", "remaining_time": "11:53:00"}
|
| 208 |
+
{"current_steps": 416, "total_steps": 1638, "loss": 1.28201425075531, "lr": 1.7753824782863827e-05, "epoch": 0.7619047619047619, "percentage": 25.4, "elapsed_time": "4:02:25", "remaining_time": "11:52:07"}
|
| 209 |
+
{"current_steps": 418, "total_steps": 1638, "loss": 0.6714475750923157, "lr": 1.773023535290956e-05, "epoch": 0.7655677655677655, "percentage": 25.52, "elapsed_time": "4:03:17", "remaining_time": "11:50:05"}
|
| 210 |
+
{"current_steps": 420, "total_steps": 1638, "loss": 1.245855450630188, "lr": 1.7706540559883066e-05, "epoch": 0.7692307692307693, "percentage": 25.64, "elapsed_time": "4:04:34", "remaining_time": "11:49:16"}
|
| 211 |
+
{"current_steps": 422, "total_steps": 1638, "loss": 0.9957686066627502, "lr": 1.7682740774730688e-05, "epoch": 0.7728937728937729, "percentage": 25.76, "elapsed_time": "4:05:45", "remaining_time": "11:48:08"}
|
| 212 |
+
{"current_steps": 424, "total_steps": 1638, "loss": 0.4911870062351227, "lr": 1.7658836370042443e-05, "epoch": 0.7765567765567766, "percentage": 25.89, "elapsed_time": "4:06:52", "remaining_time": "11:46:51"}
|
| 213 |
+
{"current_steps": 426, "total_steps": 1638, "loss": 0.7882061004638672, "lr": 1.7634827720046178e-05, "epoch": 0.7802197802197802, "percentage": 26.01, "elapsed_time": "4:07:52", "remaining_time": "11:45:12"}
|
| 214 |
+
{"current_steps": 428, "total_steps": 1638, "loss": 1.070567011833191, "lr": 1.7610715200601727e-05, "epoch": 0.7838827838827839, "percentage": 26.13, "elapsed_time": "4:08:50", "remaining_time": "11:43:30"}
|
| 215 |
+
{"current_steps": 430, "total_steps": 1638, "loss": 1.2160210609436035, "lr": 1.7586499189195016e-05, "epoch": 0.7875457875457875, "percentage": 26.25, "elapsed_time": "4:10:06", "remaining_time": "11:42:38"}
|
| 216 |
+
{"current_steps": 432, "total_steps": 1638, "loss": 1.3040070533752441, "lr": 1.7562180064932158e-05, "epoch": 0.7912087912087912, "percentage": 26.37, "elapsed_time": "4:11:19", "remaining_time": "11:41:38"}
|
| 217 |
+
{"current_steps": 434, "total_steps": 1638, "loss": 0.883861243724823, "lr": 1.7537758208533516e-05, "epoch": 0.7948717948717948, "percentage": 26.5, "elapsed_time": "4:12:31", "remaining_time": "11:40:32"}
|
| 218 |
+
{"current_steps": 436, "total_steps": 1638, "loss": 0.9705989956855774, "lr": 1.7513234002327738e-05, "epoch": 0.7985347985347986, "percentage": 26.62, "elapsed_time": "4:13:39", "remaining_time": "11:39:19"}
|
| 219 |
+
{"current_steps": 438, "total_steps": 1638, "loss": 0.8900943994522095, "lr": 1.748860783024579e-05, "epoch": 0.8021978021978022, "percentage": 26.74, "elapsed_time": "4:14:51", "remaining_time": "11:38:14"}
|
| 220 |
+
{"current_steps": 440, "total_steps": 1638, "loss": 1.3216501474380493, "lr": 1.746388007781492e-05, "epoch": 0.8058608058608059, "percentage": 26.86, "elapsed_time": "4:15:53", "remaining_time": "11:36:44"}
|
| 221 |
+
{"current_steps": 442, "total_steps": 1638, "loss": 1.203598976135254, "lr": 1.7439051132152644e-05, "epoch": 0.8095238095238095, "percentage": 26.98, "elapsed_time": "4:17:07", "remaining_time": "11:35:44"}
|
| 222 |
+
{"current_steps": 444, "total_steps": 1638, "loss": 1.205753207206726, "lr": 1.741412138196067e-05, "epoch": 0.8131868131868132, "percentage": 27.11, "elapsed_time": "4:18:24", "remaining_time": "11:34:55"}
|
| 223 |
+
{"current_steps": 446, "total_steps": 1638, "loss": 1.2312397956848145, "lr": 1.738909121751882e-05, "epoch": 0.8168498168498168, "percentage": 27.23, "elapsed_time": "4:19:37", "remaining_time": "11:33:51"}
|
| 224 |
+
{"current_steps": 448, "total_steps": 1638, "loss": 1.2268660068511963, "lr": 1.736396103067893e-05, "epoch": 0.8205128205128205, "percentage": 27.35, "elapsed_time": "4:21:01", "remaining_time": "11:33:22"}
|
| 225 |
+
{"current_steps": 450, "total_steps": 1638, "loss": 1.3640066385269165, "lr": 1.7338731214858688e-05, "epoch": 0.8241758241758241, "percentage": 27.47, "elapsed_time": "4:22:14", "remaining_time": "11:32:20"}
|
| 226 |
+
{"current_steps": 452, "total_steps": 1638, "loss": 1.0011274814605713, "lr": 1.7313402165035504e-05, "epoch": 0.8278388278388278, "percentage": 27.59, "elapsed_time": "4:23:17", "remaining_time": "11:30:50"}
|
| 227 |
+
{"current_steps": 454, "total_steps": 1638, "loss": 0.5261611938476562, "lr": 1.728797427774031e-05, "epoch": 0.8315018315018315, "percentage": 27.72, "elapsed_time": "4:24:00", "remaining_time": "11:28:30"}
|
| 228 |
+
{"current_steps": 456, "total_steps": 1638, "loss": 0.9001358151435852, "lr": 1.7262447951051366e-05, "epoch": 0.8351648351648352, "percentage": 27.84, "elapsed_time": "4:24:57", "remaining_time": "11:26:47"}
|
| 229 |
+
{"current_steps": 458, "total_steps": 1638, "loss": 0.8410728573799133, "lr": 1.7236823584587995e-05, "epoch": 0.8388278388278388, "percentage": 27.96, "elapsed_time": "4:25:56", "remaining_time": "11:25:11"}
|
| 230 |
+
{"current_steps": 460, "total_steps": 1638, "loss": 1.0333346128463745, "lr": 1.7211101579504382e-05, "epoch": 0.8424908424908425, "percentage": 28.08, "elapsed_time": "4:26:53", "remaining_time": "11:23:27"}
|
| 231 |
+
{"current_steps": 462, "total_steps": 1638, "loss": 1.230086326599121, "lr": 1.7185282338483243e-05, "epoch": 0.8461538461538461, "percentage": 28.21, "elapsed_time": "4:28:11", "remaining_time": "11:22:39"}
|
| 232 |
+
{"current_steps": 464, "total_steps": 1638, "loss": 1.1816717386245728, "lr": 1.7159366265729537e-05, "epoch": 0.8498168498168498, "percentage": 28.33, "elapsed_time": "4:29:30", "remaining_time": "11:21:54"}
|
| 233 |
+
{"current_steps": 466, "total_steps": 1638, "loss": 1.1999090909957886, "lr": 1.713335376696416e-05, "epoch": 0.8534798534798534, "percentage": 28.45, "elapsed_time": "4:30:31", "remaining_time": "11:20:21"}
|
| 234 |
+
{"current_steps": 468, "total_steps": 1638, "loss": 0.891631543636322, "lr": 1.7107245249417556e-05, "epoch": 0.8571428571428571, "percentage": 28.57, "elapsed_time": "4:31:40", "remaining_time": "11:19:11"}
|
| 235 |
+
{"current_steps": 470, "total_steps": 1638, "loss": 0.9280604720115662, "lr": 1.7081041121823375e-05, "epoch": 0.8608058608058609, "percentage": 28.69, "elapsed_time": "4:32:53", "remaining_time": "11:18:08"}
|
| 236 |
+
{"current_steps": 472, "total_steps": 1638, "loss": 1.1746187210083008, "lr": 1.705474179441205e-05, "epoch": 0.8644688644688645, "percentage": 28.82, "elapsed_time": "4:34:06", "remaining_time": "11:17:08"}
|
| 237 |
+
{"current_steps": 474, "total_steps": 1638, "loss": 0.8727067112922668, "lr": 1.7028347678904388e-05, "epoch": 0.8681318681318682, "percentage": 28.94, "elapsed_time": "4:35:01", "remaining_time": "11:15:21"}
|
| 238 |
+
{"current_steps": 476, "total_steps": 1638, "loss": 1.1008151769638062, "lr": 1.700185918850512e-05, "epoch": 0.8717948717948718, "percentage": 29.06, "elapsed_time": "4:36:09", "remaining_time": "11:14:08"}
|
| 239 |
+
{"current_steps": 478, "total_steps": 1638, "loss": 1.0441374778747559, "lr": 1.6975276737896443e-05, "epoch": 0.8754578754578755, "percentage": 29.18, "elapsed_time": "4:37:18", "remaining_time": "11:12:59"}
|
| 240 |
+
{"current_steps": 480, "total_steps": 1638, "loss": 1.0766282081604004, "lr": 1.69486007432315e-05, "epoch": 0.8791208791208791, "percentage": 29.3, "elapsed_time": "4:38:31", "remaining_time": "11:11:57"}
|
| 241 |
+
{"current_steps": 482, "total_steps": 1638, "loss": 1.1920297145843506, "lr": 1.6921831622127905e-05, "epoch": 0.8827838827838828, "percentage": 29.43, "elapsed_time": "4:39:48", "remaining_time": "11:11:05"}
|
| 242 |
+
{"current_steps": 484, "total_steps": 1638, "loss": 1.267449140548706, "lr": 1.6894969793661163e-05, "epoch": 0.8864468864468864, "percentage": 29.55, "elapsed_time": "4:41:05", "remaining_time": "11:10:12"}
|
| 243 |
+
{"current_steps": 486, "total_steps": 1638, "loss": 0.9071460366249084, "lr": 1.686801567835814e-05, "epoch": 0.8901098901098901, "percentage": 29.67, "elapsed_time": "4:42:04", "remaining_time": "11:08:37"}
|
| 244 |
+
{"current_steps": 488, "total_steps": 1638, "loss": 1.1632630825042725, "lr": 1.6840969698190467e-05, "epoch": 0.8937728937728938, "percentage": 29.79, "elapsed_time": "4:43:18", "remaining_time": "11:07:38"}
|
| 245 |
+
{"current_steps": 490, "total_steps": 1638, "loss": 1.1079881191253662, "lr": 1.6813832276567942e-05, "epoch": 0.8974358974358975, "percentage": 29.91, "elapsed_time": "4:44:29", "remaining_time": "11:06:30"}
|
| 246 |
+
{"current_steps": 492, "total_steps": 1638, "loss": 1.0516939163208008, "lr": 1.6786603838331894e-05, "epoch": 0.9010989010989011, "percentage": 30.04, "elapsed_time": "4:45:29", "remaining_time": "11:05:00"}
|
| 247 |
+
{"current_steps": 494, "total_steps": 1638, "loss": 0.5825961232185364, "lr": 1.6759284809748522e-05, "epoch": 0.9047619047619048, "percentage": 30.16, "elapsed_time": "4:46:25", "remaining_time": "11:03:18"}
|
| 248 |
+
{"current_steps": 496, "total_steps": 1638, "loss": 1.280515193939209, "lr": 1.673187561850225e-05, "epoch": 0.9084249084249084, "percentage": 30.28, "elapsed_time": "4:47:40", "remaining_time": "11:02:21"}
|
| 249 |
+
{"current_steps": 498, "total_steps": 1638, "loss": 1.1362996101379395, "lr": 1.6704376693689003e-05, "epoch": 0.9120879120879121, "percentage": 30.4, "elapsed_time": "4:48:41", "remaining_time": "11:00:51"}
|
| 250 |
+
{"current_steps": 500, "total_steps": 1638, "loss": 0.8131341338157654, "lr": 1.6676788465809506e-05, "epoch": 0.9157509157509157, "percentage": 30.53, "elapsed_time": "4:49:40", "remaining_time": "10:59:17"}
|
| 251 |
+
{"current_steps": 502, "total_steps": 1638, "loss": 0.8627150654792786, "lr": 1.6649111366762552e-05, "epoch": 0.9194139194139194, "percentage": 30.65, "elapsed_time": "4:50:39", "remaining_time": "10:57:45"}
|
| 252 |
+
{"current_steps": 504, "total_steps": 1638, "loss": 0.9442139267921448, "lr": 1.66213458298382e-05, "epoch": 0.9230769230769231, "percentage": 30.77, "elapsed_time": "4:51:51", "remaining_time": "10:56:40"}
|
README.md
ADDED
|
@@ -0,0 +1,66 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
library_name: peft
|
| 3 |
+
license: other
|
| 4 |
+
base_model: Qwen3.6-27B
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:/workspace/models/Qwen3.6-27B
|
| 7 |
+
- llama-factory
|
| 8 |
+
- lora
|
| 9 |
+
- transformers
|
| 10 |
+
pipeline_tag: text-generation
|
| 11 |
+
model-index:
|
| 12 |
+
- name: qwen35_caption_galore
|
| 13 |
+
results: []
|
| 14 |
+
---
|
| 15 |
+
|
| 16 |
+
<!-- This model card has been generated automatically according to the information the Trainer had access to. You
|
| 17 |
+
should probably proofread and complete it, then remove this comment. -->
|
| 18 |
+
|
| 19 |
+
# qwen35_caption_galore
|
| 20 |
+
|
| 21 |
+
This model is a fine-tuned version of [/workspace/models/Qwen3.6-27B](https://huggingface.co//workspace/models/Qwen3.6-27B) on the my_caption dataset.
|
| 22 |
+
|
| 23 |
+
## Model description
|
| 24 |
+
|
| 25 |
+
More information needed
|
| 26 |
+
|
| 27 |
+
## Intended uses & limitations
|
| 28 |
+
|
| 29 |
+
More information needed
|
| 30 |
+
|
| 31 |
+
## Training and evaluation data
|
| 32 |
+
|
| 33 |
+
More information needed
|
| 34 |
+
|
| 35 |
+
## Training procedure
|
| 36 |
+
|
| 37 |
+
### Training hyperparameters
|
| 38 |
+
|
| 39 |
+
The following hyperparameters were used during training:
|
| 40 |
+
- family_to_adamw_lr = {
|
| 41 |
+
"language": _fallback(getattr(training_args, "language_adamw_lr", 2e-5), language_lr),
|
| 42 |
+
"vision": _fallback(getattr(training_args, "vision_adamw_lr", 2e-5), vision_lr),
|
| 43 |
+
"merger": _fallback(getattr(training_args, "merger_adamw_lr", 4e-5), merger_lr),
|
| 44 |
+
}
|
| 45 |
+
- train_batch_size: 1
|
| 46 |
+
- eval_batch_size: 8
|
| 47 |
+
- seed: 42
|
| 48 |
+
- distributed_type: multi-GPU
|
| 49 |
+
- gradient_accumulation_steps: 24
|
| 50 |
+
- total_train_batch_size: 24
|
| 51 |
+
- optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
|
| 52 |
+
- lr_scheduler_type: cosine_with_min_lr
|
| 53 |
+
- lr_scheduler_warmup_steps: 0.03
|
| 54 |
+
- num_epochs: 3
|
| 55 |
+
|
| 56 |
+
### Training results
|
| 57 |
+
|
| 58 |
+
|
| 59 |
+
|
| 60 |
+
### Framework versions
|
| 61 |
+
|
| 62 |
+
- PEFT 0.18.1
|
| 63 |
+
- Transformers 5.5.3
|
| 64 |
+
- Pytorch 2.11.0+cu128
|
| 65 |
+
- Datasets 4.0.0
|
| 66 |
+
- Tokenizers 0.22.2
|
adapter_config.json
ADDED
|
@@ -0,0 +1,55 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "/workspace/models/Qwen3.6-27B",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"layer_replication": null,
|
| 16 |
+
"layers_pattern": null,
|
| 17 |
+
"layers_to_transform": null,
|
| 18 |
+
"loftq_config": {},
|
| 19 |
+
"lora_alpha": 256,
|
| 20 |
+
"lora_bias": false,
|
| 21 |
+
"lora_dropout": 0.05,
|
| 22 |
+
"megatron_config": null,
|
| 23 |
+
"megatron_core": "megatron.core",
|
| 24 |
+
"modules_to_save": null,
|
| 25 |
+
"peft_type": "LORA",
|
| 26 |
+
"peft_version": "0.18.1",
|
| 27 |
+
"qalora_group_size": 16,
|
| 28 |
+
"r": 256,
|
| 29 |
+
"rank_pattern": {},
|
| 30 |
+
"revision": null,
|
| 31 |
+
"target_modules": [
|
| 32 |
+
"down_proj",
|
| 33 |
+
"qkv",
|
| 34 |
+
"in_proj_z",
|
| 35 |
+
"v_proj",
|
| 36 |
+
"in_proj_qkv",
|
| 37 |
+
"gate_proj",
|
| 38 |
+
"out_proj",
|
| 39 |
+
"in_proj_b",
|
| 40 |
+
"o_proj",
|
| 41 |
+
"in_proj_a",
|
| 42 |
+
"linear_fc2",
|
| 43 |
+
"attn.proj",
|
| 44 |
+
"q_proj",
|
| 45 |
+
"linear_fc1",
|
| 46 |
+
"up_proj",
|
| 47 |
+
"k_proj"
|
| 48 |
+
],
|
| 49 |
+
"target_parameters": null,
|
| 50 |
+
"task_type": "CAUSAL_LM",
|
| 51 |
+
"trainable_token_indices": null,
|
| 52 |
+
"use_dora": false,
|
| 53 |
+
"use_qalora": false,
|
| 54 |
+
"use_rslora": false
|
| 55 |
+
}
|
adapter_model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:a43bc3fcd1012d1fe56cffbaa45b1b2e3ce1ed3194587b3a782b7772687ee956
|
| 3 |
+
size 7982962008
|
all_results.json
ADDED
|
@@ -0,0 +1,9 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"effective_tokens_per_sec": 877.2247045451046,
|
| 3 |
+
"epoch": 3.0,
|
| 4 |
+
"total_flos": 8.4482141520606e+18,
|
| 5 |
+
"train_loss": 1.0263564972723214,
|
| 6 |
+
"train_runtime": 57008.8117,
|
| 7 |
+
"train_samples_per_second": 0.69,
|
| 8 |
+
"train_steps_per_second": 0.029
|
| 9 |
+
}
|
chat_template.jinja
ADDED
|
@@ -0,0 +1,154 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- set image_count = namespace(value=0) %}
|
| 2 |
+
{%- set video_count = namespace(value=0) %}
|
| 3 |
+
{%- macro render_content(content, do_vision_count, is_system_content=false) %}
|
| 4 |
+
{%- if content is string %}
|
| 5 |
+
{{- content }}
|
| 6 |
+
{%- elif content is iterable and content is not mapping %}
|
| 7 |
+
{%- for item in content %}
|
| 8 |
+
{%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
|
| 9 |
+
{%- if is_system_content %}
|
| 10 |
+
{{- raise_exception('System message cannot contain images.') }}
|
| 11 |
+
{%- endif %}
|
| 12 |
+
{%- if do_vision_count %}
|
| 13 |
+
{%- set image_count.value = image_count.value + 1 %}
|
| 14 |
+
{%- endif %}
|
| 15 |
+
{%- if add_vision_id %}
|
| 16 |
+
{{- 'Picture ' ~ image_count.value ~ ': ' }}
|
| 17 |
+
{%- endif %}
|
| 18 |
+
{{- '<|vision_start|><|image_pad|><|vision_end|>' }}
|
| 19 |
+
{%- elif 'video' in item or item.type == 'video' %}
|
| 20 |
+
{%- if is_system_content %}
|
| 21 |
+
{{- raise_exception('System message cannot contain videos.') }}
|
| 22 |
+
{%- endif %}
|
| 23 |
+
{%- if do_vision_count %}
|
| 24 |
+
{%- set video_count.value = video_count.value + 1 %}
|
| 25 |
+
{%- endif %}
|
| 26 |
+
{%- if add_vision_id %}
|
| 27 |
+
{{- 'Video ' ~ video_count.value ~ ': ' }}
|
| 28 |
+
{%- endif %}
|
| 29 |
+
{{- '<|vision_start|><|video_pad|><|vision_end|>' }}
|
| 30 |
+
{%- elif 'text' in item %}
|
| 31 |
+
{{- item.text }}
|
| 32 |
+
{%- else %}
|
| 33 |
+
{{- raise_exception('Unexpected item type in content.') }}
|
| 34 |
+
{%- endif %}
|
| 35 |
+
{%- endfor %}
|
| 36 |
+
{%- elif content is none or content is undefined %}
|
| 37 |
+
{{- '' }}
|
| 38 |
+
{%- else %}
|
| 39 |
+
{{- raise_exception('Unexpected content type.') }}
|
| 40 |
+
{%- endif %}
|
| 41 |
+
{%- endmacro %}
|
| 42 |
+
{%- if not messages %}
|
| 43 |
+
{{- raise_exception('No messages provided.') }}
|
| 44 |
+
{%- endif %}
|
| 45 |
+
{%- if tools and tools is iterable and tools is not mapping %}
|
| 46 |
+
{{- '<|im_start|>system\n' }}
|
| 47 |
+
{{- "# Tools\n\nYou have access to the following functions:\n\n<tools>" }}
|
| 48 |
+
{%- for tool in tools %}
|
| 49 |
+
{{- "\n" }}
|
| 50 |
+
{{- tool | tojson }}
|
| 51 |
+
{%- endfor %}
|
| 52 |
+
{{- "\n</tools>" }}
|
| 53 |
+
{{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
|
| 54 |
+
{%- if messages[0].role == 'system' %}
|
| 55 |
+
{%- set content = render_content(messages[0].content, false, true)|trim %}
|
| 56 |
+
{%- if content %}
|
| 57 |
+
{{- '\n\n' + content }}
|
| 58 |
+
{%- endif %}
|
| 59 |
+
{%- endif %}
|
| 60 |
+
{{- '<|im_end|>\n' }}
|
| 61 |
+
{%- else %}
|
| 62 |
+
{%- if messages[0].role == 'system' %}
|
| 63 |
+
{%- set content = render_content(messages[0].content, false, true)|trim %}
|
| 64 |
+
{{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
|
| 65 |
+
{%- endif %}
|
| 66 |
+
{%- endif %}
|
| 67 |
+
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
| 68 |
+
{%- for message in messages[::-1] %}
|
| 69 |
+
{%- set index = (messages|length - 1) - loop.index0 %}
|
| 70 |
+
{%- if ns.multi_step_tool and message.role == "user" %}
|
| 71 |
+
{%- set content = render_content(message.content, false)|trim %}
|
| 72 |
+
{%- if not(content.startswith('<tool_response>') and content.endswith('</tool_response>')) %}
|
| 73 |
+
{%- set ns.multi_step_tool = false %}
|
| 74 |
+
{%- set ns.last_query_index = index %}
|
| 75 |
+
{%- endif %}
|
| 76 |
+
{%- endif %}
|
| 77 |
+
{%- endfor %}
|
| 78 |
+
{%- if ns.multi_step_tool %}
|
| 79 |
+
{{- raise_exception('No user query found in messages.') }}
|
| 80 |
+
{%- endif %}
|
| 81 |
+
{%- for message in messages %}
|
| 82 |
+
{%- set content = render_content(message.content, true)|trim %}
|
| 83 |
+
{%- if message.role == "system" %}
|
| 84 |
+
{%- if not loop.first %}
|
| 85 |
+
{{- raise_exception('System message must be at the beginning.') }}
|
| 86 |
+
{%- endif %}
|
| 87 |
+
{%- elif message.role == "user" %}
|
| 88 |
+
{{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
|
| 89 |
+
{%- elif message.role == "assistant" %}
|
| 90 |
+
{%- set reasoning_content = '' %}
|
| 91 |
+
{%- if message.reasoning_content is string %}
|
| 92 |
+
{%- set reasoning_content = message.reasoning_content %}
|
| 93 |
+
{%- else %}
|
| 94 |
+
{%- if '</think>' in content %}
|
| 95 |
+
{%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
| 96 |
+
{%- set content = content.split('</think>')[-1].lstrip('\n') %}
|
| 97 |
+
{%- endif %}
|
| 98 |
+
{%- endif %}
|
| 99 |
+
{%- set reasoning_content = reasoning_content|trim %}
|
| 100 |
+
{%- if (preserve_thinking is defined and preserve_thinking is true) or (loop.index0 > ns.last_query_index) %}
|
| 101 |
+
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }}
|
| 102 |
+
{%- else %}
|
| 103 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 104 |
+
{%- endif %}
|
| 105 |
+
{%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
|
| 106 |
+
{%- for tool_call in message.tool_calls %}
|
| 107 |
+
{%- if tool_call.function is defined %}
|
| 108 |
+
{%- set tool_call = tool_call.function %}
|
| 109 |
+
{%- endif %}
|
| 110 |
+
{%- if loop.first %}
|
| 111 |
+
{%- if content|trim %}
|
| 112 |
+
{{- '\n\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
|
| 113 |
+
{%- else %}
|
| 114 |
+
{{- '<tool_call>\n<function=' + tool_call.name + '>\n' }}
|
| 115 |
+
{%- endif %}
|
| 116 |
+
{%- else %}
|
| 117 |
+
{{- '\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
|
| 118 |
+
{%- endif %}
|
| 119 |
+
{%- if tool_call.arguments is defined %}
|
| 120 |
+
{%- for args_name, args_value in tool_call.arguments|items %}
|
| 121 |
+
{{- '<parameter=' + args_name + '>\n' }}
|
| 122 |
+
{%- set args_value = args_value | string if args_value is string else args_value | tojson | safe %}
|
| 123 |
+
{{- args_value }}
|
| 124 |
+
{{- '\n</parameter>\n' }}
|
| 125 |
+
{%- endfor %}
|
| 126 |
+
{%- endif %}
|
| 127 |
+
{{- '</function>\n</tool_call>' }}
|
| 128 |
+
{%- endfor %}
|
| 129 |
+
{%- endif %}
|
| 130 |
+
{{- '<|im_end|>\n' }}
|
| 131 |
+
{%- elif message.role == "tool" %}
|
| 132 |
+
{%- if loop.previtem and loop.previtem.role != "tool" %}
|
| 133 |
+
{{- '<|im_start|>user' }}
|
| 134 |
+
{%- endif %}
|
| 135 |
+
{{- '\n<tool_response>\n' }}
|
| 136 |
+
{{- content }}
|
| 137 |
+
{{- '\n</tool_response>' }}
|
| 138 |
+
{%- if not loop.last and loop.nextitem.role != "tool" %}
|
| 139 |
+
{{- '<|im_end|>\n' }}
|
| 140 |
+
{%- elif loop.last %}
|
| 141 |
+
{{- '<|im_end|>\n' }}
|
| 142 |
+
{%- endif %}
|
| 143 |
+
{%- else %}
|
| 144 |
+
{{- raise_exception('Unexpected message role.') }}
|
| 145 |
+
{%- endif %}
|
| 146 |
+
{%- endfor %}
|
| 147 |
+
{%- if add_generation_prompt %}
|
| 148 |
+
{{- '<|im_start|>assistant\n' }}
|
| 149 |
+
{%- if enable_thinking is defined and enable_thinking is false %}
|
| 150 |
+
{{- '<think>\n\n</think>\n\n' }}
|
| 151 |
+
{%- else %}
|
| 152 |
+
{{- '<think>\n' }}
|
| 153 |
+
{%- endif %}
|
| 154 |
+
{%- endif %}
|
processor_config.json
ADDED
|
@@ -0,0 +1,60 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"image_processor": {
|
| 3 |
+
"do_convert_rgb": true,
|
| 4 |
+
"do_normalize": true,
|
| 5 |
+
"do_rescale": true,
|
| 6 |
+
"do_resize": true,
|
| 7 |
+
"image_mean": [
|
| 8 |
+
0.5,
|
| 9 |
+
0.5,
|
| 10 |
+
0.5
|
| 11 |
+
],
|
| 12 |
+
"image_processor_type": "Qwen2VLImageProcessor",
|
| 13 |
+
"image_std": [
|
| 14 |
+
0.5,
|
| 15 |
+
0.5,
|
| 16 |
+
0.5
|
| 17 |
+
],
|
| 18 |
+
"merge_size": 2,
|
| 19 |
+
"patch_size": 16,
|
| 20 |
+
"resample": 3,
|
| 21 |
+
"rescale_factor": 0.00392156862745098,
|
| 22 |
+
"size": {
|
| 23 |
+
"longest_edge": 16777216,
|
| 24 |
+
"shortest_edge": 65536
|
| 25 |
+
},
|
| 26 |
+
"temporal_patch_size": 2
|
| 27 |
+
},
|
| 28 |
+
"processor_class": "Qwen3VLProcessor",
|
| 29 |
+
"video_processor": {
|
| 30 |
+
"do_convert_rgb": true,
|
| 31 |
+
"do_normalize": true,
|
| 32 |
+
"do_rescale": true,
|
| 33 |
+
"do_resize": true,
|
| 34 |
+
"do_sample_frames": true,
|
| 35 |
+
"fps": 2,
|
| 36 |
+
"image_mean": [
|
| 37 |
+
0.5,
|
| 38 |
+
0.5,
|
| 39 |
+
0.5
|
| 40 |
+
],
|
| 41 |
+
"image_std": [
|
| 42 |
+
0.5,
|
| 43 |
+
0.5,
|
| 44 |
+
0.5
|
| 45 |
+
],
|
| 46 |
+
"max_frames": 768,
|
| 47 |
+
"merge_size": 2,
|
| 48 |
+
"min_frames": 4,
|
| 49 |
+
"patch_size": 16,
|
| 50 |
+
"resample": 3,
|
| 51 |
+
"rescale_factor": 0.00392156862745098,
|
| 52 |
+
"return_metadata": false,
|
| 53 |
+
"size": {
|
| 54 |
+
"longest_edge": 25165824,
|
| 55 |
+
"shortest_edge": 4096
|
| 56 |
+
},
|
| 57 |
+
"temporal_patch_size": 2,
|
| 58 |
+
"video_processor_type": "Qwen3VLVideoProcessor"
|
| 59 |
+
}
|
| 60 |
+
}
|
tokenizer.json
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:639e352c0f904c1875d448ebed6f6faac005fd3eb58393b7f1fb3ff044e5ca03
|
| 3 |
+
size 19989510
|
tokenizer_config.json
ADDED
|
@@ -0,0 +1,37 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"audio_bos_token": "<|audio_start|>",
|
| 4 |
+
"audio_eos_token": "<|audio_end|>",
|
| 5 |
+
"audio_token": "<|audio_pad|>",
|
| 6 |
+
"backend": "tokenizers",
|
| 7 |
+
"bos_token": null,
|
| 8 |
+
"clean_up_tokenization_spaces": false,
|
| 9 |
+
"eos_token": "<|im_end|>",
|
| 10 |
+
"errors": "replace",
|
| 11 |
+
"image_token": "<|image_pad|>",
|
| 12 |
+
"is_local": true,
|
| 13 |
+
"local_files_only": false,
|
| 14 |
+
"max_length": null,
|
| 15 |
+
"model_max_length": 262144,
|
| 16 |
+
"model_specific_special_tokens": {
|
| 17 |
+
"audio_bos_token": "<|audio_start|>",
|
| 18 |
+
"audio_eos_token": "<|audio_end|>",
|
| 19 |
+
"audio_token": "<|audio_pad|>",
|
| 20 |
+
"image_token": "<|image_pad|>",
|
| 21 |
+
"video_token": "<|video_pad|>",
|
| 22 |
+
"vision_bos_token": "<|vision_start|>",
|
| 23 |
+
"vision_eos_token": "<|vision_end|>"
|
| 24 |
+
},
|
| 25 |
+
"pad_to_multiple_of": null,
|
| 26 |
+
"pad_token": "<|endoftext|>",
|
| 27 |
+
"pad_token_type_id": 0,
|
| 28 |
+
"padding_side": "right",
|
| 29 |
+
"pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
|
| 30 |
+
"processor_class": "Qwen3VLProcessor",
|
| 31 |
+
"split_special_tokens": false,
|
| 32 |
+
"tokenizer_class": "TokenizersBackend",
|
| 33 |
+
"unk_token": null,
|
| 34 |
+
"video_token": "<|video_pad|>",
|
| 35 |
+
"vision_bos_token": "<|vision_start|>",
|
| 36 |
+
"vision_eos_token": "<|vision_end|>"
|
| 37 |
+
}
|
train_results.json
ADDED
|
@@ -0,0 +1,9 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"effective_tokens_per_sec": 877.2247045451046,
|
| 3 |
+
"epoch": 3.0,
|
| 4 |
+
"total_flos": 8.4482141520606e+18,
|
| 5 |
+
"train_loss": 1.0263564972723214,
|
| 6 |
+
"train_runtime": 57008.8117,
|
| 7 |
+
"train_samples_per_second": 0.69,
|
| 8 |
+
"train_steps_per_second": 0.029
|
| 9 |
+
}
|
trainer_log.jsonl
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
trainer_state.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
training_args.bin
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:4445782897782eb6684c92fee23eb92728cc3f0c2c35c342df714dcfe682014f
|
| 3 |
+
size 5649
|
training_loss.png
ADDED
|