Add files using upload-large-folder tool
Browse filesThis view is limited to 50 files because it contains too many changes. Β
See raw diff
- wandb/run-20241030_010641-sozqwfy9/files/config.yaml +47 -0
- wandb/run-20241030_010641-sozqwfy9/files/output.log +4 -0
- wandb/run-20241030_010641-sozqwfy9/files/wandb-metadata.json +29 -0
- wandb/run-20241030_010641-sozqwfy9/files/wandb-summary.json +1 -0
- wandb/run-20241030_010641-sozqwfy9/logs/debug-internal.log +16 -0
- wandb/run-20241030_010641-sozqwfy9/logs/debug.log +27 -0
- wandb/run-20241030_010641-sozqwfy9/run-sozqwfy9.wandb +0 -0
- wandb/run-20241030_010759-0a5n2onp/files/config.yaml +47 -0
- wandb/run-20241030_010759-0a5n2onp/files/output.log +4 -0
- wandb/run-20241030_010759-0a5n2onp/files/wandb-metadata.json +97 -0
- wandb/run-20241030_010759-0a5n2onp/files/wandb-summary.json +1 -0
- wandb/run-20241030_010759-0a5n2onp/logs/debug-internal.log +16 -0
- wandb/run-20241030_010759-0a5n2onp/logs/debug.log +27 -0
- wandb/run-20241030_010759-0a5n2onp/run-0a5n2onp.wandb +0 -0
- wandb/run-20241030_010759-imp8g625/files/wandb-metadata.json +97 -0
- wandb/run-20241030_010759-imp8g625/files/wandb-summary.json +1 -0
- wandb/run-20241030_010759-imp8g625/logs/debug-internal.log +16 -0
- wandb/run-20241030_010759-imp8g625/logs/debug.log +27 -0
- wandb/run-20241030_011014-zmts6m10/files/config.yaml +47 -0
- wandb/run-20241030_011014-zmts6m10/files/output.log +6 -0
- wandb/run-20241030_011014-zmts6m10/files/wandb-metadata.json +97 -0
- wandb/run-20241030_011014-zmts6m10/files/wandb-summary.json +1 -0
- wandb/run-20241030_011014-zmts6m10/logs/debug-internal.log +16 -0
- wandb/run-20241030_011014-zmts6m10/logs/debug.log +27 -0
- wandb/run-20241030_011014-zmts6m10/run-zmts6m10.wandb +0 -0
- wandb/run-20241030_013339-dgadwxty/files/output.log +38 -0
- wandb/run-20241030_013339-dgadwxty/files/requirements.txt +147 -0
- wandb/run-20241030_013339-dgadwxty/files/wandb-metadata.json +97 -0
- wandb/run-20241030_013339-dgadwxty/logs/debug-internal.log +8 -0
- wandb/run-20241030_013339-dgadwxty/logs/debug.log +26 -0
- wandb/run-20241030_013339-s77qk5li/files/output.log +41 -0
- wandb/run-20241030_013339-s77qk5li/files/requirements.txt +147 -0
- wandb/run-20241030_013339-s77qk5li/files/wandb-metadata.json +97 -0
- wandb/run-20241030_013339-s77qk5li/logs/debug-internal.log +8 -0
- wandb/run-20241030_013339-s77qk5li/logs/debug.log +26 -0
- wandb/run-20241030_112700-jhzkfwvw/files/config.yaml +47 -0
- wandb/run-20241030_112700-jhzkfwvw/files/output.log +34 -0
- wandb/run-20241030_112700-jhzkfwvw/files/requirements.txt +147 -0
- wandb/run-20241030_112700-jhzkfwvw/files/wandb-metadata.json +97 -0
- wandb/run-20241030_112700-jhzkfwvw/files/wandb-summary.json +1 -0
- wandb/run-20241030_112700-jhzkfwvw/logs/debug-internal.log +11 -0
- wandb/run-20241030_112700-jhzkfwvw/logs/debug.log +27 -0
- wandb/run-20241030_112700-jhzkfwvw/run-jhzkfwvw.wandb +0 -0
- wandb/run-20241031_122113-8ldget07/files/config.yaml +50 -0
- wandb/run-20241031_122113-8ldget07/files/output.log +14 -0
- wandb/run-20241031_122113-8ldget07/files/wandb-metadata.json +97 -0
- wandb/run-20241031_122113-8ldget07/files/wandb-summary.json +1 -0
- wandb/run-20241031_122113-8ldget07/logs/debug-internal.log +18 -0
- wandb/run-20241031_122113-8ldget07/logs/debug.log +33 -0
- wandb/run-20241101_094656-b81aanqd/files/output.log +13 -0
wandb/run-20241030_010641-sozqwfy9/files/config.yaml
ADDED
|
@@ -0,0 +1,47 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
_wandb:
|
| 2 |
+
value:
|
| 3 |
+
cli_version: 0.18.5
|
| 4 |
+
m: []
|
| 5 |
+
python_version: 3.9.19
|
| 6 |
+
t:
|
| 7 |
+
"1":
|
| 8 |
+
- 1
|
| 9 |
+
- 5
|
| 10 |
+
- 11
|
| 11 |
+
- 49
|
| 12 |
+
- 51
|
| 13 |
+
- 53
|
| 14 |
+
- 55
|
| 15 |
+
- 71
|
| 16 |
+
- 98
|
| 17 |
+
"2":
|
| 18 |
+
- 1
|
| 19 |
+
- 5
|
| 20 |
+
- 11
|
| 21 |
+
- 49
|
| 22 |
+
- 51
|
| 23 |
+
- 53
|
| 24 |
+
- 55
|
| 25 |
+
- 71
|
| 26 |
+
- 98
|
| 27 |
+
"3":
|
| 28 |
+
- 13
|
| 29 |
+
- 23
|
| 30 |
+
- 55
|
| 31 |
+
"4": 3.9.19
|
| 32 |
+
"5": 0.18.5
|
| 33 |
+
"6": 4.45.1
|
| 34 |
+
"8":
|
| 35 |
+
- 5
|
| 36 |
+
"12": 0.18.5
|
| 37 |
+
"13": linux-x86_64
|
| 38 |
+
batch_size:
|
| 39 |
+
value: 3
|
| 40 |
+
epoch:
|
| 41 |
+
value: 7
|
| 42 |
+
perturbation:
|
| 43 |
+
value: reverse_control
|
| 44 |
+
seed:
|
| 45 |
+
value: 0
|
| 46 |
+
train_set:
|
| 47 |
+
value: 10M
|
wandb/run-20241030_010641-sozqwfy9/files/output.log
ADDED
|
@@ -0,0 +1,4 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Traceback (most recent call last):
|
| 2 |
+
File "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py", line 162, in <module>
|
| 3 |
+
dataset_name = f"babylm_{args.perturbation}_{args.train_zset}_seed{args.seed}"
|
| 4 |
+
AttributeError: 'Namespace' object has no attribute 'train_zset'
|
wandb/run-20241030_010641-sozqwfy9/files/wandb-metadata.json
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
|
| 3 |
+
"python": "3.9.19",
|
| 4 |
+
"startedAt": "2024-10-30T05:06:41.543355Z",
|
| 5 |
+
"args": [
|
| 6 |
+
"--perturbation",
|
| 7 |
+
"reverse_control",
|
| 8 |
+
"--train_set",
|
| 9 |
+
"10M",
|
| 10 |
+
"--batch_size",
|
| 11 |
+
"3",
|
| 12 |
+
"--epoch",
|
| 13 |
+
"7",
|
| 14 |
+
"--seed",
|
| 15 |
+
"0"
|
| 16 |
+
],
|
| 17 |
+
"program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
|
| 18 |
+
"codePath": "train/train_deep_wandb.py",
|
| 19 |
+
"git": {
|
| 20 |
+
"remote": "git@hf.co:Yaning1001/Impossible_llm.git",
|
| 21 |
+
"commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
|
| 22 |
+
},
|
| 23 |
+
"email": "yaning1001@gmail.com",
|
| 24 |
+
"root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
|
| 25 |
+
"host": "mms-large-2",
|
| 26 |
+
"username": "chunhui",
|
| 27 |
+
"executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
|
| 28 |
+
"codePathLocal": "train_deep_wandb.py"
|
| 29 |
+
}
|
wandb/run-20241030_010641-sozqwfy9/files/wandb-summary.json
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
{"_wandb":{"runtime":1}}
|
wandb/run-20241030_010641-sozqwfy9/logs/debug-internal.log
ADDED
|
@@ -0,0 +1,16 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"time":"2024-10-30T01:06:41.544820605-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
|
| 2 |
+
{"time":"2024-10-30T01:06:41.544828795-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_010641-sozqwfy9/logs/debug-core.log"}
|
| 3 |
+
{"time":"2024-10-30T01:06:41.651808474-04:00","level":"INFO","msg":"created new stream","id":"sozqwfy9"}
|
| 4 |
+
{"time":"2024-10-30T01:06:41.651840385-04:00","level":"INFO","msg":"stream: started","id":"sozqwfy9"}
|
| 5 |
+
{"time":"2024-10-30T01:06:41.651863335-04:00","level":"INFO","msg":"sender: started","stream_id":"sozqwfy9"}
|
| 6 |
+
{"time":"2024-10-30T01:06:41.651850035-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"sozqwfy9"}}
|
| 7 |
+
{"time":"2024-10-30T01:06:41.651862905-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"sozqwfy9"}}
|
| 8 |
+
{"time":"2024-10-30T01:06:43.032746197-04:00","level":"INFO","msg":"Starting system monitor"}
|
| 9 |
+
{"time":"2024-10-30T01:06:43.127763565-04:00","level":"INFO","msg":"stream: closing","id":"sozqwfy9"}
|
| 10 |
+
{"time":"2024-10-30T01:06:43.127845395-04:00","level":"INFO","msg":"Stopping system monitor"}
|
| 11 |
+
{"time":"2024-10-30T01:06:43.226914121-04:00","level":"INFO","msg":"Stopped system monitor"}
|
| 12 |
+
{"time":"2024-10-30T01:06:43.693295593-04:00","level":"INFO","msg":"fileTransfer: Close: file transfer manager closed"}
|
| 13 |
+
{"time":"2024-10-30T01:06:43.805986822-04:00","level":"INFO","msg":"handler: closed","stream_id":{"value":"sozqwfy9"}}
|
| 14 |
+
{"time":"2024-10-30T01:06:43.806037683-04:00","level":"INFO","msg":"writer: Close: closed","stream_id":{"value":"sozqwfy9"}}
|
| 15 |
+
{"time":"2024-10-30T01:06:43.806051923-04:00","level":"INFO","msg":"sender: closed","stream_id":"sozqwfy9"}
|
| 16 |
+
{"time":"2024-10-30T01:06:43.806121243-04:00","level":"INFO","msg":"stream: closed","id":"sozqwfy9"}
|
wandb/run-20241030_010641-sozqwfy9/logs/debug.log
ADDED
|
@@ -0,0 +1,27 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
2024-10-30 01:06:41,541 INFO MainThread:321598 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
|
| 2 |
+
2024-10-30 01:06:41,541 INFO MainThread:321598 [wandb_setup.py:_flush():79] Configure stats pid to 321598
|
| 3 |
+
2024-10-30 01:06:41,541 INFO MainThread:321598 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
|
| 4 |
+
2024-10-30 01:06:41,541 INFO MainThread:321598 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
|
| 5 |
+
2024-10-30 01:06:41,541 INFO MainThread:321598 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
|
| 6 |
+
2024-10-30 01:06:41,541 INFO MainThread:321598 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
|
| 7 |
+
2024-10-30 01:06:41,541 INFO MainThread:321598 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
|
| 8 |
+
2024-10-30 01:06:41,542 INFO MainThread:321598 [wandb_setup.py:_flush():79] Applying login settings: {}
|
| 9 |
+
2024-10-30 01:06:41,542 INFO MainThread:321598 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_010641-sozqwfy9/logs/debug.log
|
| 10 |
+
2024-10-30 01:06:41,542 INFO MainThread:321598 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_010641-sozqwfy9/logs/debug-internal.log
|
| 11 |
+
2024-10-30 01:06:41,542 INFO MainThread:321598 [wandb_init.py:init():621] calling init triggers
|
| 12 |
+
2024-10-30 01:06:41,542 INFO MainThread:321598 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
|
| 13 |
+
config: {}
|
| 14 |
+
2024-10-30 01:06:41,542 INFO MainThread:321598 [wandb_init.py:init():671] starting backend
|
| 15 |
+
2024-10-30 01:06:41,542 INFO MainThread:321598 [wandb_init.py:init():675] sending inform_init request
|
| 16 |
+
2024-10-30 01:06:41,542 INFO MainThread:321598 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
|
| 17 |
+
2024-10-30 01:06:41,543 INFO MainThread:321598 [wandb_init.py:init():688] backend started and connected
|
| 18 |
+
2024-10-30 01:06:41,545 INFO MainThread:321598 [wandb_init.py:init():783] updated telemetry
|
| 19 |
+
2024-10-30 01:06:41,569 INFO MainThread:321598 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
|
| 20 |
+
2024-10-30 01:06:43,029 INFO MainThread:321598 [wandb_init.py:init():867] starting run threads in backend
|
| 21 |
+
2024-10-30 01:06:43,124 INFO MainThread:321598 [wandb_run.py:_console_start():2463] atexit reg
|
| 22 |
+
2024-10-30 01:06:43,124 INFO MainThread:321598 [wandb_run.py:_redirect():2311] redirect: wrap_raw
|
| 23 |
+
2024-10-30 01:06:43,124 INFO MainThread:321598 [wandb_run.py:_redirect():2376] Wrapping output streams.
|
| 24 |
+
2024-10-30 01:06:43,124 INFO MainThread:321598 [wandb_run.py:_redirect():2401] Redirects installed.
|
| 25 |
+
2024-10-30 01:06:43,126 INFO MainThread:321598 [wandb_init.py:init():911] run started, returning control to user process
|
| 26 |
+
2024-10-30 01:06:43,126 INFO MainThread:321598 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'reverse_control', 'train_set': '10M', 'batch_size': 3, 'epoch': 7, 'seed': 0}
|
| 27 |
+
2024-10-30 01:06:43,127 WARNING MsgRouterThr:321598 [router.py:message_loop():77] message_loop has been closed
|
wandb/run-20241030_010641-sozqwfy9/run-sozqwfy9.wandb
ADDED
|
Binary file (1.57 kB). View file
|
|
|
wandb/run-20241030_010759-0a5n2onp/files/config.yaml
ADDED
|
@@ -0,0 +1,47 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
_wandb:
|
| 2 |
+
value:
|
| 3 |
+
cli_version: 0.18.5
|
| 4 |
+
m: []
|
| 5 |
+
python_version: 3.9.19
|
| 6 |
+
t:
|
| 7 |
+
"1":
|
| 8 |
+
- 1
|
| 9 |
+
- 5
|
| 10 |
+
- 11
|
| 11 |
+
- 49
|
| 12 |
+
- 51
|
| 13 |
+
- 53
|
| 14 |
+
- 55
|
| 15 |
+
- 71
|
| 16 |
+
- 98
|
| 17 |
+
"2":
|
| 18 |
+
- 1
|
| 19 |
+
- 5
|
| 20 |
+
- 11
|
| 21 |
+
- 49
|
| 22 |
+
- 51
|
| 23 |
+
- 53
|
| 24 |
+
- 55
|
| 25 |
+
- 71
|
| 26 |
+
- 98
|
| 27 |
+
"3":
|
| 28 |
+
- 13
|
| 29 |
+
- 23
|
| 30 |
+
- 55
|
| 31 |
+
"4": 3.9.19
|
| 32 |
+
"5": 0.18.5
|
| 33 |
+
"6": 4.45.1
|
| 34 |
+
"8":
|
| 35 |
+
- 5
|
| 36 |
+
"12": 0.18.5
|
| 37 |
+
"13": linux-x86_64
|
| 38 |
+
batch_size:
|
| 39 |
+
value: 3
|
| 40 |
+
epoch:
|
| 41 |
+
value: 7
|
| 42 |
+
perturbation:
|
| 43 |
+
value: reverse_control
|
| 44 |
+
seed:
|
| 45 |
+
value: 0
|
| 46 |
+
train_set:
|
| 47 |
+
value: 10M
|
wandb/run-20241030_010759-0a5n2onp/files/output.log
ADDED
|
@@ -0,0 +1,4 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Traceback (most recent call last):
|
| 2 |
+
File "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py", line 162, in <module>
|
| 3 |
+
dataset_name = f"babylm_{args.perturbation}_{args.train_zset}_seed{args.seed}"
|
| 4 |
+
AttributeError: 'Namespace' object has no attribute 'train_zset'
|
wandb/run-20241030_010759-0a5n2onp/files/wandb-metadata.json
ADDED
|
@@ -0,0 +1,97 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
|
| 3 |
+
"python": "3.9.19",
|
| 4 |
+
"startedAt": "2024-10-30T05:07:59.024551Z",
|
| 5 |
+
"args": [
|
| 6 |
+
"--perturbation",
|
| 7 |
+
"reverse_control",
|
| 8 |
+
"--train_set",
|
| 9 |
+
"10M",
|
| 10 |
+
"--batch_size",
|
| 11 |
+
"3",
|
| 12 |
+
"--epoch",
|
| 13 |
+
"7",
|
| 14 |
+
"--seed",
|
| 15 |
+
"0"
|
| 16 |
+
],
|
| 17 |
+
"program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
|
| 18 |
+
"codePath": "train/train_deep_wandb.py",
|
| 19 |
+
"git": {
|
| 20 |
+
"remote": "git@hf.co:Yaning1001/Impossible_llm.git",
|
| 21 |
+
"commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
|
| 22 |
+
},
|
| 23 |
+
"email": "yaning1001@gmail.com",
|
| 24 |
+
"root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
|
| 25 |
+
"host": "mms-large-2",
|
| 26 |
+
"username": "chunhui",
|
| 27 |
+
"executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
|
| 28 |
+
"codePathLocal": "train_deep_wandb.py",
|
| 29 |
+
"cpu_count": 32,
|
| 30 |
+
"cpu_count_logical": 64,
|
| 31 |
+
"gpu": "NVIDIA RTX A6000",
|
| 32 |
+
"gpu_count": 8,
|
| 33 |
+
"disk": {
|
| 34 |
+
"/": {
|
| 35 |
+
"total": "1888559353856",
|
| 36 |
+
"used": "1719200268288"
|
| 37 |
+
}
|
| 38 |
+
},
|
| 39 |
+
"memory": {
|
| 40 |
+
"total": "202617098240"
|
| 41 |
+
},
|
| 42 |
+
"cpu": {
|
| 43 |
+
"count": 32,
|
| 44 |
+
"countLogical": 64
|
| 45 |
+
},
|
| 46 |
+
"gpu_nvidia": [
|
| 47 |
+
{
|
| 48 |
+
"name": "NVIDIA RTX A6000",
|
| 49 |
+
"memoryTotal": "51527024640",
|
| 50 |
+
"cudaCores": 10752,
|
| 51 |
+
"architecture": "Ampere"
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"name": "NVIDIA RTX A6000",
|
| 55 |
+
"memoryTotal": "51527024640",
|
| 56 |
+
"cudaCores": 10752,
|
| 57 |
+
"architecture": "Ampere"
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"name": "NVIDIA RTX A6000",
|
| 61 |
+
"memoryTotal": "51527024640",
|
| 62 |
+
"cudaCores": 10752,
|
| 63 |
+
"architecture": "Ampere"
|
| 64 |
+
},
|
| 65 |
+
{
|
| 66 |
+
"name": "NVIDIA RTX A6000",
|
| 67 |
+
"memoryTotal": "51527024640",
|
| 68 |
+
"cudaCores": 10752,
|
| 69 |
+
"architecture": "Ampere"
|
| 70 |
+
},
|
| 71 |
+
{
|
| 72 |
+
"name": "NVIDIA RTX A6000",
|
| 73 |
+
"memoryTotal": "51527024640",
|
| 74 |
+
"cudaCores": 10752,
|
| 75 |
+
"architecture": "Ampere"
|
| 76 |
+
},
|
| 77 |
+
{
|
| 78 |
+
"name": "NVIDIA RTX A6000",
|
| 79 |
+
"memoryTotal": "51527024640",
|
| 80 |
+
"cudaCores": 10752,
|
| 81 |
+
"architecture": "Ampere"
|
| 82 |
+
},
|
| 83 |
+
{
|
| 84 |
+
"name": "NVIDIA RTX A6000",
|
| 85 |
+
"memoryTotal": "51527024640",
|
| 86 |
+
"cudaCores": 10752,
|
| 87 |
+
"architecture": "Ampere"
|
| 88 |
+
},
|
| 89 |
+
{
|
| 90 |
+
"name": "NVIDIA RTX A6000",
|
| 91 |
+
"memoryTotal": "51527024640",
|
| 92 |
+
"cudaCores": 10752,
|
| 93 |
+
"architecture": "Ampere"
|
| 94 |
+
}
|
| 95 |
+
],
|
| 96 |
+
"cudaVersion": "11.8"
|
| 97 |
+
}
|
wandb/run-20241030_010759-0a5n2onp/files/wandb-summary.json
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
{"_wandb":{"runtime":0}}
|
wandb/run-20241030_010759-0a5n2onp/logs/debug-internal.log
ADDED
|
@@ -0,0 +1,16 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"time":"2024-10-30T01:07:59.026431301-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
|
| 2 |
+
{"time":"2024-10-30T01:07:59.026442171-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_010759-0a5n2onp/logs/debug-core.log"}
|
| 3 |
+
{"time":"2024-10-30T01:07:59.133417319-04:00","level":"INFO","msg":"created new stream","id":"0a5n2onp"}
|
| 4 |
+
{"time":"2024-10-30T01:07:59.13345488-04:00","level":"INFO","msg":"stream: started","id":"0a5n2onp"}
|
| 5 |
+
{"time":"2024-10-30T01:07:59.13347872-04:00","level":"INFO","msg":"sender: started","stream_id":"0a5n2onp"}
|
| 6 |
+
{"time":"2024-10-30T01:07:59.13348111-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"0a5n2onp"}}
|
| 7 |
+
{"time":"2024-10-30T01:07:59.13347098-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"0a5n2onp"}}
|
| 8 |
+
{"time":"2024-10-30T01:07:59.301083731-04:00","level":"INFO","msg":"Starting system monitor"}
|
| 9 |
+
{"time":"2024-10-30T01:07:59.400210746-04:00","level":"INFO","msg":"stream: closing","id":"0a5n2onp"}
|
| 10 |
+
{"time":"2024-10-30T01:07:59.400243016-04:00","level":"INFO","msg":"Stopping system monitor"}
|
| 11 |
+
{"time":"2024-10-30T01:07:59.400545198-04:00","level":"INFO","msg":"Stopped system monitor"}
|
| 12 |
+
{"time":"2024-10-30T01:07:59.953692364-04:00","level":"INFO","msg":"fileTransfer: Close: file transfer manager closed"}
|
| 13 |
+
{"time":"2024-10-30T01:08:00.087291614-04:00","level":"INFO","msg":"handler: closed","stream_id":{"value":"0a5n2onp"}}
|
| 14 |
+
{"time":"2024-10-30T01:08:00.087360055-04:00","level":"INFO","msg":"sender: closed","stream_id":"0a5n2onp"}
|
| 15 |
+
{"time":"2024-10-30T01:08:00.087352894-04:00","level":"INFO","msg":"writer: Close: closed","stream_id":{"value":"0a5n2onp"}}
|
| 16 |
+
{"time":"2024-10-30T01:08:00.087440995-04:00","level":"INFO","msg":"stream: closed","id":"0a5n2onp"}
|
wandb/run-20241030_010759-0a5n2onp/logs/debug.log
ADDED
|
@@ -0,0 +1,27 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
2024-10-30 01:07:59,022 INFO MainThread:322458 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
|
| 2 |
+
2024-10-30 01:07:59,022 INFO MainThread:322458 [wandb_setup.py:_flush():79] Configure stats pid to 322458
|
| 3 |
+
2024-10-30 01:07:59,022 INFO MainThread:322458 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
|
| 4 |
+
2024-10-30 01:07:59,022 INFO MainThread:322458 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
|
| 5 |
+
2024-10-30 01:07:59,022 INFO MainThread:322458 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
|
| 6 |
+
2024-10-30 01:07:59,022 INFO MainThread:322458 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
|
| 7 |
+
2024-10-30 01:07:59,022 INFO MainThread:322458 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
|
| 8 |
+
2024-10-30 01:07:59,022 INFO MainThread:322458 [wandb_setup.py:_flush():79] Applying login settings: {}
|
| 9 |
+
2024-10-30 01:07:59,022 INFO MainThread:322458 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_010759-0a5n2onp/logs/debug.log
|
| 10 |
+
2024-10-30 01:07:59,023 INFO MainThread:322458 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_010759-0a5n2onp/logs/debug-internal.log
|
| 11 |
+
2024-10-30 01:07:59,023 INFO MainThread:322458 [wandb_init.py:init():621] calling init triggers
|
| 12 |
+
2024-10-30 01:07:59,023 INFO MainThread:322458 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
|
| 13 |
+
config: {}
|
| 14 |
+
2024-10-30 01:07:59,023 INFO MainThread:322458 [wandb_init.py:init():671] starting backend
|
| 15 |
+
2024-10-30 01:07:59,023 INFO MainThread:322458 [wandb_init.py:init():675] sending inform_init request
|
| 16 |
+
2024-10-30 01:07:59,024 INFO MainThread:322458 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
|
| 17 |
+
2024-10-30 01:07:59,024 INFO MainThread:322458 [wandb_init.py:init():688] backend started and connected
|
| 18 |
+
2024-10-30 01:07:59,027 INFO MainThread:322458 [wandb_init.py:init():783] updated telemetry
|
| 19 |
+
2024-10-30 01:07:59,064 INFO MainThread:322458 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
|
| 20 |
+
2024-10-30 01:07:59,297 INFO MainThread:322458 [wandb_init.py:init():867] starting run threads in backend
|
| 21 |
+
2024-10-30 01:07:59,396 INFO MainThread:322458 [wandb_run.py:_console_start():2463] atexit reg
|
| 22 |
+
2024-10-30 01:07:59,396 INFO MainThread:322458 [wandb_run.py:_redirect():2311] redirect: wrap_raw
|
| 23 |
+
2024-10-30 01:07:59,396 INFO MainThread:322458 [wandb_run.py:_redirect():2376] Wrapping output streams.
|
| 24 |
+
2024-10-30 01:07:59,396 INFO MainThread:322458 [wandb_run.py:_redirect():2401] Redirects installed.
|
| 25 |
+
2024-10-30 01:07:59,398 INFO MainThread:322458 [wandb_init.py:init():911] run started, returning control to user process
|
| 26 |
+
2024-10-30 01:07:59,399 INFO MainThread:322458 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'reverse_control', 'train_set': '10M', 'batch_size': 3, 'epoch': 7, 'seed': 0}
|
| 27 |
+
2024-10-30 01:07:59,400 WARNING MsgRouterThr:322458 [router.py:message_loop():77] message_loop has been closed
|
wandb/run-20241030_010759-0a5n2onp/run-0a5n2onp.wandb
ADDED
|
Binary file (1.6 kB). View file
|
|
|
wandb/run-20241030_010759-imp8g625/files/wandb-metadata.json
ADDED
|
@@ -0,0 +1,97 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
|
| 3 |
+
"python": "3.9.19",
|
| 4 |
+
"startedAt": "2024-10-30T05:07:59.128982Z",
|
| 5 |
+
"args": [
|
| 6 |
+
"--perturbation",
|
| 7 |
+
"reverse_control",
|
| 8 |
+
"--train_set",
|
| 9 |
+
"10M",
|
| 10 |
+
"--batch_size",
|
| 11 |
+
"3",
|
| 12 |
+
"--epoch",
|
| 13 |
+
"7",
|
| 14 |
+
"--seed",
|
| 15 |
+
"0"
|
| 16 |
+
],
|
| 17 |
+
"program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
|
| 18 |
+
"codePath": "train/train_deep_wandb.py",
|
| 19 |
+
"git": {
|
| 20 |
+
"remote": "git@hf.co:Yaning1001/Impossible_llm.git",
|
| 21 |
+
"commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
|
| 22 |
+
},
|
| 23 |
+
"email": "yaning1001@gmail.com",
|
| 24 |
+
"root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
|
| 25 |
+
"host": "mms-large-2",
|
| 26 |
+
"username": "chunhui",
|
| 27 |
+
"executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
|
| 28 |
+
"codePathLocal": "train_deep_wandb.py",
|
| 29 |
+
"cpu_count": 32,
|
| 30 |
+
"cpu_count_logical": 64,
|
| 31 |
+
"gpu": "NVIDIA RTX A6000",
|
| 32 |
+
"gpu_count": 8,
|
| 33 |
+
"disk": {
|
| 34 |
+
"/": {
|
| 35 |
+
"total": "1888559353856",
|
| 36 |
+
"used": "1719200272384"
|
| 37 |
+
}
|
| 38 |
+
},
|
| 39 |
+
"memory": {
|
| 40 |
+
"total": "202617098240"
|
| 41 |
+
},
|
| 42 |
+
"cpu": {
|
| 43 |
+
"count": 32,
|
| 44 |
+
"countLogical": 64
|
| 45 |
+
},
|
| 46 |
+
"gpu_nvidia": [
|
| 47 |
+
{
|
| 48 |
+
"name": "NVIDIA RTX A6000",
|
| 49 |
+
"memoryTotal": "51527024640",
|
| 50 |
+
"cudaCores": 10752,
|
| 51 |
+
"architecture": "Ampere"
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"name": "NVIDIA RTX A6000",
|
| 55 |
+
"memoryTotal": "51527024640",
|
| 56 |
+
"cudaCores": 10752,
|
| 57 |
+
"architecture": "Ampere"
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"name": "NVIDIA RTX A6000",
|
| 61 |
+
"memoryTotal": "51527024640",
|
| 62 |
+
"cudaCores": 10752,
|
| 63 |
+
"architecture": "Ampere"
|
| 64 |
+
},
|
| 65 |
+
{
|
| 66 |
+
"name": "NVIDIA RTX A6000",
|
| 67 |
+
"memoryTotal": "51527024640",
|
| 68 |
+
"cudaCores": 10752,
|
| 69 |
+
"architecture": "Ampere"
|
| 70 |
+
},
|
| 71 |
+
{
|
| 72 |
+
"name": "NVIDIA RTX A6000",
|
| 73 |
+
"memoryTotal": "51527024640",
|
| 74 |
+
"cudaCores": 10752,
|
| 75 |
+
"architecture": "Ampere"
|
| 76 |
+
},
|
| 77 |
+
{
|
| 78 |
+
"name": "NVIDIA RTX A6000",
|
| 79 |
+
"memoryTotal": "51527024640",
|
| 80 |
+
"cudaCores": 10752,
|
| 81 |
+
"architecture": "Ampere"
|
| 82 |
+
},
|
| 83 |
+
{
|
| 84 |
+
"name": "NVIDIA RTX A6000",
|
| 85 |
+
"memoryTotal": "51527024640",
|
| 86 |
+
"cudaCores": 10752,
|
| 87 |
+
"architecture": "Ampere"
|
| 88 |
+
},
|
| 89 |
+
{
|
| 90 |
+
"name": "NVIDIA RTX A6000",
|
| 91 |
+
"memoryTotal": "51527024640",
|
| 92 |
+
"cudaCores": 10752,
|
| 93 |
+
"architecture": "Ampere"
|
| 94 |
+
}
|
| 95 |
+
],
|
| 96 |
+
"cudaVersion": "11.8"
|
| 97 |
+
}
|
wandb/run-20241030_010759-imp8g625/files/wandb-summary.json
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
{"_wandb":{"runtime":0}}
|
wandb/run-20241030_010759-imp8g625/logs/debug-internal.log
ADDED
|
@@ -0,0 +1,16 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"time":"2024-10-30T01:07:59.130957192-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
|
| 2 |
+
{"time":"2024-10-30T01:07:59.130976003-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_010759-imp8g625/logs/debug-core.log"}
|
| 3 |
+
{"time":"2024-10-30T01:07:59.237319546-04:00","level":"INFO","msg":"created new stream","id":"imp8g625"}
|
| 4 |
+
{"time":"2024-10-30T01:07:59.237366547-04:00","level":"INFO","msg":"stream: started","id":"imp8g625"}
|
| 5 |
+
{"time":"2024-10-30T01:07:59.237426827-04:00","level":"INFO","msg":"sender: started","stream_id":"imp8g625"}
|
| 6 |
+
{"time":"2024-10-30T01:07:59.237392927-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"imp8g625"}}
|
| 7 |
+
{"time":"2024-10-30T01:07:59.237409567-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"imp8g625"}}
|
| 8 |
+
{"time":"2024-10-30T01:07:59.460104113-04:00","level":"INFO","msg":"Starting system monitor"}
|
| 9 |
+
{"time":"2024-10-30T01:07:59.567682276-04:00","level":"INFO","msg":"stream: closing","id":"imp8g625"}
|
| 10 |
+
{"time":"2024-10-30T01:07:59.567714696-04:00","level":"INFO","msg":"Stopping system monitor"}
|
| 11 |
+
{"time":"2024-10-30T01:07:59.568157119-04:00","level":"INFO","msg":"Stopped system monitor"}
|
| 12 |
+
{"time":"2024-10-30T01:08:00.146034424-04:00","level":"INFO","msg":"fileTransfer: Close: file transfer manager closed"}
|
| 13 |
+
{"time":"2024-10-30T01:08:00.269377913-04:00","level":"INFO","msg":"handler: closed","stream_id":{"value":"imp8g625"}}
|
| 14 |
+
{"time":"2024-10-30T01:08:00.269400854-04:00","level":"INFO","msg":"writer: Close: closed","stream_id":{"value":"imp8g625"}}
|
| 15 |
+
{"time":"2024-10-30T01:08:00.269409974-04:00","level":"INFO","msg":"sender: closed","stream_id":"imp8g625"}
|
| 16 |
+
{"time":"2024-10-30T01:08:00.269448294-04:00","level":"INFO","msg":"stream: closed","id":"imp8g625"}
|
wandb/run-20241030_010759-imp8g625/logs/debug.log
ADDED
|
@@ -0,0 +1,27 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
2024-10-30 01:07:59,126 INFO MainThread:322460 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
|
| 2 |
+
2024-10-30 01:07:59,126 INFO MainThread:322460 [wandb_setup.py:_flush():79] Configure stats pid to 322460
|
| 3 |
+
2024-10-30 01:07:59,126 INFO MainThread:322460 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
|
| 4 |
+
2024-10-30 01:07:59,126 INFO MainThread:322460 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
|
| 5 |
+
2024-10-30 01:07:59,126 INFO MainThread:322460 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
|
| 6 |
+
2024-10-30 01:07:59,126 INFO MainThread:322460 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
|
| 7 |
+
2024-10-30 01:07:59,126 INFO MainThread:322460 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
|
| 8 |
+
2024-10-30 01:07:59,126 INFO MainThread:322460 [wandb_setup.py:_flush():79] Applying login settings: {}
|
| 9 |
+
2024-10-30 01:07:59,126 INFO MainThread:322460 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_010759-imp8g625/logs/debug.log
|
| 10 |
+
2024-10-30 01:07:59,126 INFO MainThread:322460 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_010759-imp8g625/logs/debug-internal.log
|
| 11 |
+
2024-10-30 01:07:59,126 INFO MainThread:322460 [wandb_init.py:init():621] calling init triggers
|
| 12 |
+
2024-10-30 01:07:59,127 INFO MainThread:322460 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
|
| 13 |
+
config: {}
|
| 14 |
+
2024-10-30 01:07:59,127 INFO MainThread:322460 [wandb_init.py:init():671] starting backend
|
| 15 |
+
2024-10-30 01:07:59,127 INFO MainThread:322460 [wandb_init.py:init():675] sending inform_init request
|
| 16 |
+
2024-10-30 01:07:59,128 INFO MainThread:322460 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
|
| 17 |
+
2024-10-30 01:07:59,128 INFO MainThread:322460 [wandb_init.py:init():688] backend started and connected
|
| 18 |
+
2024-10-30 01:07:59,132 INFO MainThread:322460 [wandb_init.py:init():783] updated telemetry
|
| 19 |
+
2024-10-30 01:07:59,161 INFO MainThread:322460 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
|
| 20 |
+
2024-10-30 01:07:59,457 INFO MainThread:322460 [wandb_init.py:init():867] starting run threads in backend
|
| 21 |
+
2024-10-30 01:07:59,563 INFO MainThread:322460 [wandb_run.py:_console_start():2463] atexit reg
|
| 22 |
+
2024-10-30 01:07:59,563 INFO MainThread:322460 [wandb_run.py:_redirect():2311] redirect: wrap_raw
|
| 23 |
+
2024-10-30 01:07:59,563 INFO MainThread:322460 [wandb_run.py:_redirect():2376] Wrapping output streams.
|
| 24 |
+
2024-10-30 01:07:59,564 INFO MainThread:322460 [wandb_run.py:_redirect():2401] Redirects installed.
|
| 25 |
+
2024-10-30 01:07:59,566 INFO MainThread:322460 [wandb_init.py:init():911] run started, returning control to user process
|
| 26 |
+
2024-10-30 01:07:59,566 INFO MainThread:322460 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'reverse_control', 'train_set': '10M', 'batch_size': 3, 'epoch': 7, 'seed': 0}
|
| 27 |
+
2024-10-30 01:07:59,567 WARNING MsgRouterThr:322460 [router.py:message_loop():77] message_loop has been closed
|
wandb/run-20241030_011014-zmts6m10/files/config.yaml
ADDED
|
@@ -0,0 +1,47 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
_wandb:
|
| 2 |
+
value:
|
| 3 |
+
cli_version: 0.18.5
|
| 4 |
+
m: []
|
| 5 |
+
python_version: 3.9.19
|
| 6 |
+
t:
|
| 7 |
+
"1":
|
| 8 |
+
- 1
|
| 9 |
+
- 5
|
| 10 |
+
- 11
|
| 11 |
+
- 49
|
| 12 |
+
- 51
|
| 13 |
+
- 53
|
| 14 |
+
- 55
|
| 15 |
+
- 71
|
| 16 |
+
- 98
|
| 17 |
+
"2":
|
| 18 |
+
- 1
|
| 19 |
+
- 5
|
| 20 |
+
- 11
|
| 21 |
+
- 49
|
| 22 |
+
- 51
|
| 23 |
+
- 53
|
| 24 |
+
- 55
|
| 25 |
+
- 71
|
| 26 |
+
- 98
|
| 27 |
+
"3":
|
| 28 |
+
- 13
|
| 29 |
+
- 23
|
| 30 |
+
- 55
|
| 31 |
+
"4": 3.9.19
|
| 32 |
+
"5": 0.18.5
|
| 33 |
+
"6": 4.45.1
|
| 34 |
+
"8":
|
| 35 |
+
- 5
|
| 36 |
+
"12": 0.18.5
|
| 37 |
+
"13": linux-x86_64
|
| 38 |
+
batch_size:
|
| 39 |
+
value: 3
|
| 40 |
+
epoch:
|
| 41 |
+
value: 7
|
| 42 |
+
perturbation:
|
| 43 |
+
value: reverse_control
|
| 44 |
+
seed:
|
| 45 |
+
value: 0
|
| 46 |
+
train_set:
|
| 47 |
+
value: 10M
|
wandb/run-20241030_011014-zmts6m10/files/output.log
ADDED
|
@@ -0,0 +1,6 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Traceback (most recent call last):
|
| 2 |
+
File "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py", line 165, in <module>
|
| 3 |
+
valid_dataset = dataset['validation']
|
| 4 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/datasets/dataset_dict.py", line 72, in __getitem__
|
| 5 |
+
return super().__getitem__(k)
|
| 6 |
+
KeyError: 'validation'
|
wandb/run-20241030_011014-zmts6m10/files/wandb-metadata.json
ADDED
|
@@ -0,0 +1,97 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
|
| 3 |
+
"python": "3.9.19",
|
| 4 |
+
"startedAt": "2024-10-30T05:10:14.058527Z",
|
| 5 |
+
"args": [
|
| 6 |
+
"--perturbation",
|
| 7 |
+
"reverse_control",
|
| 8 |
+
"--train_set",
|
| 9 |
+
"10M",
|
| 10 |
+
"--batch_size",
|
| 11 |
+
"3",
|
| 12 |
+
"--epoch",
|
| 13 |
+
"7",
|
| 14 |
+
"--seed",
|
| 15 |
+
"0"
|
| 16 |
+
],
|
| 17 |
+
"program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
|
| 18 |
+
"codePath": "train/train_deep_wandb.py",
|
| 19 |
+
"git": {
|
| 20 |
+
"remote": "git@hf.co:Yaning1001/Impossible_llm.git",
|
| 21 |
+
"commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
|
| 22 |
+
},
|
| 23 |
+
"email": "yaning1001@gmail.com",
|
| 24 |
+
"root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
|
| 25 |
+
"host": "mms-large-2",
|
| 26 |
+
"username": "chunhui",
|
| 27 |
+
"executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
|
| 28 |
+
"codePathLocal": "train_deep_wandb.py",
|
| 29 |
+
"cpu_count": 32,
|
| 30 |
+
"cpu_count_logical": 64,
|
| 31 |
+
"gpu": "NVIDIA RTX A6000",
|
| 32 |
+
"gpu_count": 8,
|
| 33 |
+
"disk": {
|
| 34 |
+
"/": {
|
| 35 |
+
"total": "1888559353856",
|
| 36 |
+
"used": "1719200362496"
|
| 37 |
+
}
|
| 38 |
+
},
|
| 39 |
+
"memory": {
|
| 40 |
+
"total": "202617098240"
|
| 41 |
+
},
|
| 42 |
+
"cpu": {
|
| 43 |
+
"count": 32,
|
| 44 |
+
"countLogical": 64
|
| 45 |
+
},
|
| 46 |
+
"gpu_nvidia": [
|
| 47 |
+
{
|
| 48 |
+
"name": "NVIDIA RTX A6000",
|
| 49 |
+
"memoryTotal": "51527024640",
|
| 50 |
+
"cudaCores": 10752,
|
| 51 |
+
"architecture": "Ampere"
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"name": "NVIDIA RTX A6000",
|
| 55 |
+
"memoryTotal": "51527024640",
|
| 56 |
+
"cudaCores": 10752,
|
| 57 |
+
"architecture": "Ampere"
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"name": "NVIDIA RTX A6000",
|
| 61 |
+
"memoryTotal": "51527024640",
|
| 62 |
+
"cudaCores": 10752,
|
| 63 |
+
"architecture": "Ampere"
|
| 64 |
+
},
|
| 65 |
+
{
|
| 66 |
+
"name": "NVIDIA RTX A6000",
|
| 67 |
+
"memoryTotal": "51527024640",
|
| 68 |
+
"cudaCores": 10752,
|
| 69 |
+
"architecture": "Ampere"
|
| 70 |
+
},
|
| 71 |
+
{
|
| 72 |
+
"name": "NVIDIA RTX A6000",
|
| 73 |
+
"memoryTotal": "51527024640",
|
| 74 |
+
"cudaCores": 10752,
|
| 75 |
+
"architecture": "Ampere"
|
| 76 |
+
},
|
| 77 |
+
{
|
| 78 |
+
"name": "NVIDIA RTX A6000",
|
| 79 |
+
"memoryTotal": "51527024640",
|
| 80 |
+
"cudaCores": 10752,
|
| 81 |
+
"architecture": "Ampere"
|
| 82 |
+
},
|
| 83 |
+
{
|
| 84 |
+
"name": "NVIDIA RTX A6000",
|
| 85 |
+
"memoryTotal": "51527024640",
|
| 86 |
+
"cudaCores": 10752,
|
| 87 |
+
"architecture": "Ampere"
|
| 88 |
+
},
|
| 89 |
+
{
|
| 90 |
+
"name": "NVIDIA RTX A6000",
|
| 91 |
+
"memoryTotal": "51527024640",
|
| 92 |
+
"cudaCores": 10752,
|
| 93 |
+
"architecture": "Ampere"
|
| 94 |
+
}
|
| 95 |
+
],
|
| 96 |
+
"cudaVersion": "11.8"
|
| 97 |
+
}
|
wandb/run-20241030_011014-zmts6m10/files/wandb-summary.json
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
{"_wandb":{"runtime":7}}
|
wandb/run-20241030_011014-zmts6m10/logs/debug-internal.log
ADDED
|
@@ -0,0 +1,16 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"time":"2024-10-30T01:10:14.060931522-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
|
| 2 |
+
{"time":"2024-10-30T01:10:14.060945603-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_011014-zmts6m10/logs/debug-core.log"}
|
| 3 |
+
{"time":"2024-10-30T01:10:14.16965664-04:00","level":"INFO","msg":"created new stream","id":"zmts6m10"}
|
| 4 |
+
{"time":"2024-10-30T01:10:14.169719361-04:00","level":"INFO","msg":"stream: started","id":"zmts6m10"}
|
| 5 |
+
{"time":"2024-10-30T01:10:14.169812672-04:00","level":"INFO","msg":"sender: started","stream_id":"zmts6m10"}
|
| 6 |
+
{"time":"2024-10-30T01:10:14.169748791-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"zmts6m10"}}
|
| 7 |
+
{"time":"2024-10-30T01:10:14.169800151-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"zmts6m10"}}
|
| 8 |
+
{"time":"2024-10-30T01:10:14.564909463-04:00","level":"INFO","msg":"Starting system monitor"}
|
| 9 |
+
{"time":"2024-10-30T01:10:22.006513638-04:00","level":"INFO","msg":"stream: closing","id":"zmts6m10"}
|
| 10 |
+
{"time":"2024-10-30T01:10:22.006567608-04:00","level":"INFO","msg":"Stopping system monitor"}
|
| 11 |
+
{"time":"2024-10-30T01:10:22.007193673-04:00","level":"INFO","msg":"Stopped system monitor"}
|
| 12 |
+
{"time":"2024-10-30T01:10:22.414493317-04:00","level":"INFO","msg":"fileTransfer: Close: file transfer manager closed"}
|
| 13 |
+
{"time":"2024-10-30T01:10:22.536976149-04:00","level":"INFO","msg":"handler: closed","stream_id":{"value":"zmts6m10"}}
|
| 14 |
+
{"time":"2024-10-30T01:10:22.537005739-04:00","level":"INFO","msg":"writer: Close: closed","stream_id":{"value":"zmts6m10"}}
|
| 15 |
+
{"time":"2024-10-30T01:10:22.537036749-04:00","level":"INFO","msg":"sender: closed","stream_id":"zmts6m10"}
|
| 16 |
+
{"time":"2024-10-30T01:10:22.53708566-04:00","level":"INFO","msg":"stream: closed","id":"zmts6m10"}
|
wandb/run-20241030_011014-zmts6m10/logs/debug.log
ADDED
|
@@ -0,0 +1,27 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
2024-10-30 01:10:14,056 INFO MainThread:323567 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
|
| 2 |
+
2024-10-30 01:10:14,056 INFO MainThread:323567 [wandb_setup.py:_flush():79] Configure stats pid to 323567
|
| 3 |
+
2024-10-30 01:10:14,056 INFO MainThread:323567 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
|
| 4 |
+
2024-10-30 01:10:14,056 INFO MainThread:323567 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
|
| 5 |
+
2024-10-30 01:10:14,057 INFO MainThread:323567 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
|
| 6 |
+
2024-10-30 01:10:14,057 INFO MainThread:323567 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
|
| 7 |
+
2024-10-30 01:10:14,057 INFO MainThread:323567 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
|
| 8 |
+
2024-10-30 01:10:14,057 INFO MainThread:323567 [wandb_setup.py:_flush():79] Applying login settings: {}
|
| 9 |
+
2024-10-30 01:10:14,057 INFO MainThread:323567 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_011014-zmts6m10/logs/debug.log
|
| 10 |
+
2024-10-30 01:10:14,057 INFO MainThread:323567 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_011014-zmts6m10/logs/debug-internal.log
|
| 11 |
+
2024-10-30 01:10:14,057 INFO MainThread:323567 [wandb_init.py:init():621] calling init triggers
|
| 12 |
+
2024-10-30 01:10:14,057 INFO MainThread:323567 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
|
| 13 |
+
config: {}
|
| 14 |
+
2024-10-30 01:10:14,057 INFO MainThread:323567 [wandb_init.py:init():671] starting backend
|
| 15 |
+
2024-10-30 01:10:14,057 INFO MainThread:323567 [wandb_init.py:init():675] sending inform_init request
|
| 16 |
+
2024-10-30 01:10:14,058 INFO MainThread:323567 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
|
| 17 |
+
2024-10-30 01:10:14,058 INFO MainThread:323567 [wandb_init.py:init():688] backend started and connected
|
| 18 |
+
2024-10-30 01:10:14,061 INFO MainThread:323567 [wandb_init.py:init():783] updated telemetry
|
| 19 |
+
2024-10-30 01:10:14,093 INFO MainThread:323567 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
|
| 20 |
+
2024-10-30 01:10:14,561 INFO MainThread:323567 [wandb_init.py:init():867] starting run threads in backend
|
| 21 |
+
2024-10-30 01:10:14,660 INFO MainThread:323567 [wandb_run.py:_console_start():2463] atexit reg
|
| 22 |
+
2024-10-30 01:10:14,660 INFO MainThread:323567 [wandb_run.py:_redirect():2311] redirect: wrap_raw
|
| 23 |
+
2024-10-30 01:10:14,660 INFO MainThread:323567 [wandb_run.py:_redirect():2376] Wrapping output streams.
|
| 24 |
+
2024-10-30 01:10:14,660 INFO MainThread:323567 [wandb_run.py:_redirect():2401] Redirects installed.
|
| 25 |
+
2024-10-30 01:10:14,661 INFO MainThread:323567 [wandb_init.py:init():911] run started, returning control to user process
|
| 26 |
+
2024-10-30 01:10:14,662 INFO MainThread:323567 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'reverse_control', 'train_set': '10M', 'batch_size': 3, 'epoch': 7, 'seed': 0}
|
| 27 |
+
2024-10-30 01:10:22,006 WARNING MsgRouterThr:323567 [router.py:message_loop():77] message_loop has been closed
|
wandb/run-20241030_011014-zmts6m10/run-zmts6m10.wandb
ADDED
|
Binary file (1.8 kB). View file
|
|
|
wandb/run-20241030_013339-dgadwxty/files/output.log
ADDED
|
@@ -0,0 +1,38 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Downloading shards: 100%|ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ| 2/2 [02:32<00:00, 76.19s/it]
|
| 2 |
+
Loading checkpoint shards: 100%|βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ| 2/2 [00:05<00:00, 2.52s/it]
|
| 3 |
+
Map: 100%|βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ| 17519/17519 [00:58<00:00, 301.89 examples/s]
|
| 4 |
+
Map: 100%|βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ| 18140/18140 [00:56<00:00, 322.79 examples/s]
|
| 5 |
+
tokenized_valid: Dataset({
|
| 6 |
+
features: ['input_ids', 'attention_mask'],
|
| 7 |
+
num_rows: 600
|
| 8 |
+
})
|
| 9 |
+
/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/training_args.py:1545: FutureWarning: `evaluation_strategy` is deprecated and will be removed in version 4.46 of π€ Transformers. Use `eval_strategy` instead
|
| 10 |
+
warnings.warn(
|
| 11 |
+
[2024-10-30 01:38:13,521] [INFO] [real_accelerator.py:219:get_accelerator] Setting ds_accelerator to cuda (auto detect)
|
| 12 |
+
[2024-10-30 01:38:21,110] [INFO] [comm.py:652:init_distributed] cdb=None
|
| 13 |
+
Installed CUDA version 11.8 does not match the version torch was compiled with 11.7 but since the APIs are compatible, accepting this combination
|
| 14 |
+
Using /home/chunhui/.cache/torch_extensions/py39_cu117 as PyTorch extensions root...
|
| 15 |
+
Loading extension module cpu_adam...
|
| 16 |
+
Time to load cpu_adam op: 4.589682579040527 seconds
|
| 17 |
+
Traceback (most recent call last):
|
| 18 |
+
File "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py", line 219, in <module>
|
| 19 |
+
trainer.train()
|
| 20 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/trainer.py", line 2052, in train
|
| 21 |
+
return inner_training_loop(
|
| 22 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/trainer.py", line 2467, in _inner_training_loop
|
| 23 |
+
self._maybe_log_save_evaluate(tr_loss, grad_norm, model, trial, epoch, ignore_keys_for_eval)
|
| 24 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/trainer.py", line 2918, in _maybe_log_save_evaluate
|
| 25 |
+
self._save_checkpoint(model, trial, metrics=metrics)
|
| 26 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/trainer.py", line 3008, in _save_checkpoint
|
| 27 |
+
self.save_model(output_dir, _internal_call=True)
|
| 28 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/trainer.py", line 3608, in save_model
|
| 29 |
+
state_dict = self.accelerator.get_state_dict(self.deepspeed)
|
| 30 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/accelerate/accelerator.py", line 3382, in get_state_dict
|
| 31 |
+
state_dict = clone_tensors_for_torch_save(self.unwrap_model(model).state_dict())
|
| 32 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/deepspeed/checkpoint/utils.py", line 60, in clone_tensors_for_torch_save
|
| 33 |
+
return type(item)({k: clone_tensors_for_torch_save(v, device) for k, v in item.items()})
|
| 34 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/deepspeed/checkpoint/utils.py", line 60, in <dictcomp>
|
| 35 |
+
return type(item)({k: clone_tensors_for_torch_save(v, device) for k, v in item.items()})
|
| 36 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/deepspeed/checkpoint/utils.py", line 54, in clone_tensors_for_torch_save
|
| 37 |
+
return item.detach().clone().to(device)
|
| 38 |
+
KeyboardInterrupt
|
wandb/run-20241030_013339-dgadwxty/files/requirements.txt
ADDED
|
@@ -0,0 +1,147 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
funcsigs==1.0.2
|
| 2 |
+
sentry-sdk==2.17.0
|
| 3 |
+
multiprocess==0.70.16
|
| 4 |
+
numpy==1.26.2
|
| 5 |
+
pluralizer==1.2.0
|
| 6 |
+
debugpy==1.6.7
|
| 7 |
+
nvidia-cudnn-cu11==8.5.0.96
|
| 8 |
+
deepspeed==0.15.2
|
| 9 |
+
data==0.4
|
| 10 |
+
pandas==2.1.3
|
| 11 |
+
tomli==2.0.1
|
| 12 |
+
charset-normalizer==3.3.2
|
| 13 |
+
attrs==24.2.0
|
| 14 |
+
aiosignal==1.3.1
|
| 15 |
+
fsspec==2023.10.0
|
| 16 |
+
nvidia-cusparse-cu11==11.7.4.91
|
| 17 |
+
zipp==3.12.0
|
| 18 |
+
mypy-extensions==1.0.0
|
| 19 |
+
datasets==3.0.1
|
| 20 |
+
joblib==1.3.2
|
| 21 |
+
hjson==3.1.0
|
| 22 |
+
traitlets==5.7.1
|
| 23 |
+
stack-data==0.6.0
|
| 24 |
+
transformers==4.45.1
|
| 25 |
+
sympy==1.11.1
|
| 26 |
+
Pygments==2.15.0
|
| 27 |
+
docker-pycreds==0.4.0
|
| 28 |
+
dill==0.3.8
|
| 29 |
+
wheel==0.44.0
|
| 30 |
+
prompt-toolkit==3.0.30
|
| 31 |
+
parso==0.8.3
|
| 32 |
+
ipykernel==6.23.1
|
| 33 |
+
pyarrow==17.0.0
|
| 34 |
+
certifi==2023.11.17
|
| 35 |
+
nvidia-cufft-cu11==10.9.0.58
|
| 36 |
+
six==1.16.0
|
| 37 |
+
pydantic==2.9.2
|
| 38 |
+
click==8.1.7
|
| 39 |
+
nest-asyncio==1.5.6
|
| 40 |
+
gmpy2==2.1.0
|
| 41 |
+
matplotlib==3.8.2
|
| 42 |
+
scipy==1.11.4
|
| 43 |
+
typing_extensions==4.12.2
|
| 44 |
+
statsmodels==0.14.0
|
| 45 |
+
huggingface-hub==0.25.0
|
| 46 |
+
frozenlist==1.4.1
|
| 47 |
+
gpustat==1.1.1
|
| 48 |
+
nvidia-nvtx-cu11==11.7.91
|
| 49 |
+
safetensors==0.4.5
|
| 50 |
+
stanza==1.9.2
|
| 51 |
+
decorator==5.1.1
|
| 52 |
+
seaborn==0.13.0
|
| 53 |
+
sentencepiece==0.2.0
|
| 54 |
+
PyYAML==6.0.1
|
| 55 |
+
black==24.8.0
|
| 56 |
+
protobuf==4.25.1
|
| 57 |
+
pickleshare==0.7.5
|
| 58 |
+
peft==0.13.0
|
| 59 |
+
triton==2.0.0
|
| 60 |
+
nvidia-cuda-runtime-cu11==11.7.99
|
| 61 |
+
Jinja2==3.1.2
|
| 62 |
+
nvidia-cusolver-cu11==11.4.0.1
|
| 63 |
+
executing==1.2.0
|
| 64 |
+
jupyter_client==8.1.0
|
| 65 |
+
pluggy==1.3.0
|
| 66 |
+
cmake==3.30.3
|
| 67 |
+
pytz==2023.3.post1
|
| 68 |
+
aiohappyeyeballs==2.4.2
|
| 69 |
+
kiwisolver==1.4.5
|
| 70 |
+
py-cpuinfo==9.0.0
|
| 71 |
+
Pillow==10.1.0
|
| 72 |
+
ptyprocess==0.7.0
|
| 73 |
+
importlib_resources==6.4.5
|
| 74 |
+
GitPython==3.1.43
|
| 75 |
+
importlib-metadata==6.0.0
|
| 76 |
+
iniconfig==2.0.0
|
| 77 |
+
scikit-learn==1.3.2
|
| 78 |
+
exceptiongroup==1.1.0
|
| 79 |
+
networkx==2.8.6
|
| 80 |
+
accelerate==1.0.0
|
| 81 |
+
nltk==3.8.1
|
| 82 |
+
shutilwhich==1.1.0
|
| 83 |
+
fonttools==4.45.1
|
| 84 |
+
future==0.18.3
|
| 85 |
+
aiohttp==3.10.6
|
| 86 |
+
wcwidth==0.2.5
|
| 87 |
+
idna==3.6
|
| 88 |
+
filelock==3.12.2
|
| 89 |
+
pathspec==0.12.1
|
| 90 |
+
jupyter_core==5.1.0
|
| 91 |
+
lit==18.1.8
|
| 92 |
+
nvidia-curand-cu11==10.2.10.91
|
| 93 |
+
nvidia-cublas-cu11==11.10.3.66
|
| 94 |
+
nvidia-ml-py==12.560.30
|
| 95 |
+
msgpack==1.1.0
|
| 96 |
+
python-dateutil==2.8.2
|
| 97 |
+
blessed==1.20.0
|
| 98 |
+
packaging==23.0
|
| 99 |
+
gitdb==4.0.11
|
| 100 |
+
yarl==1.13.0
|
| 101 |
+
emoji==2.8.0
|
| 102 |
+
tzdata==2023.3
|
| 103 |
+
cycler==0.12.1
|
| 104 |
+
tornado==6.2
|
| 105 |
+
backcall==0.2.0
|
| 106 |
+
plotnine==0.12.4
|
| 107 |
+
ninja==1.11.1.1
|
| 108 |
+
latex==0.7.0
|
| 109 |
+
wandb==0.18.5
|
| 110 |
+
setproctitle==1.3.3
|
| 111 |
+
threadpoolctl==3.2.0
|
| 112 |
+
requests==2.32.3
|
| 113 |
+
pyparsing==3.1.1
|
| 114 |
+
smmap==5.0.1
|
| 115 |
+
pyzmq==23.0.0
|
| 116 |
+
async-timeout==4.0.3
|
| 117 |
+
annotated-types==0.7.0
|
| 118 |
+
matplotlib-inline==0.1.6
|
| 119 |
+
latexcodec==1.0.0
|
| 120 |
+
ipython==8.0.0
|
| 121 |
+
patsy==0.5.3
|
| 122 |
+
contourpy==1.2.0
|
| 123 |
+
multidict==6.1.0
|
| 124 |
+
mizani==0.9.3
|
| 125 |
+
urllib3==2.1.0
|
| 126 |
+
tokenizers==0.20.0
|
| 127 |
+
MarkupSafe==2.1.2
|
| 128 |
+
pip==24.2
|
| 129 |
+
pexpect==4.8.0
|
| 130 |
+
tqdm==4.66.5
|
| 131 |
+
jedi==0.18.2
|
| 132 |
+
pydantic_core==2.23.4
|
| 133 |
+
tempdir==0.7.1
|
| 134 |
+
mpmath==1.2.1
|
| 135 |
+
setuptools==72.1.0
|
| 136 |
+
pytest==7.4.3
|
| 137 |
+
pure-eval==0.2.2
|
| 138 |
+
psutil==5.9.1
|
| 139 |
+
comm==0.1.2
|
| 140 |
+
nvidia-cuda-cupti-cu11==11.7.101
|
| 141 |
+
nvidia-cuda-nvrtc-cu11==11.7.99
|
| 142 |
+
regex==2023.10.3
|
| 143 |
+
platformdirs==2.5.2
|
| 144 |
+
asttokens==2.2.1
|
| 145 |
+
torch==2.0.0
|
| 146 |
+
nvidia-nccl-cu11==2.14.3
|
| 147 |
+
xxhash==3.5.0
|
wandb/run-20241030_013339-dgadwxty/files/wandb-metadata.json
ADDED
|
@@ -0,0 +1,97 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
|
| 3 |
+
"python": "3.9.19",
|
| 4 |
+
"startedAt": "2024-10-30T05:33:39.878858Z",
|
| 5 |
+
"args": [
|
| 6 |
+
"--perturbation",
|
| 7 |
+
"reverse_full",
|
| 8 |
+
"--train_set",
|
| 9 |
+
"10M",
|
| 10 |
+
"--batch_size",
|
| 11 |
+
"3",
|
| 12 |
+
"--epoch",
|
| 13 |
+
"7",
|
| 14 |
+
"--seed",
|
| 15 |
+
"0"
|
| 16 |
+
],
|
| 17 |
+
"program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
|
| 18 |
+
"codePath": "train/train_deep_wandb.py",
|
| 19 |
+
"git": {
|
| 20 |
+
"remote": "git@hf.co:Yaning1001/Impossible_llm.git",
|
| 21 |
+
"commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
|
| 22 |
+
},
|
| 23 |
+
"email": "yaning1001@gmail.com",
|
| 24 |
+
"root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
|
| 25 |
+
"host": "mms-large-2",
|
| 26 |
+
"username": "chunhui",
|
| 27 |
+
"executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
|
| 28 |
+
"codePathLocal": "train_deep_wandb.py",
|
| 29 |
+
"cpu_count": 32,
|
| 30 |
+
"cpu_count_logical": 64,
|
| 31 |
+
"gpu": "NVIDIA RTX A6000",
|
| 32 |
+
"gpu_count": 8,
|
| 33 |
+
"disk": {
|
| 34 |
+
"/": {
|
| 35 |
+
"total": "1888559353856",
|
| 36 |
+
"used": "1710081835008"
|
| 37 |
+
}
|
| 38 |
+
},
|
| 39 |
+
"memory": {
|
| 40 |
+
"total": "202617098240"
|
| 41 |
+
},
|
| 42 |
+
"cpu": {
|
| 43 |
+
"count": 32,
|
| 44 |
+
"countLogical": 64
|
| 45 |
+
},
|
| 46 |
+
"gpu_nvidia": [
|
| 47 |
+
{
|
| 48 |
+
"name": "NVIDIA RTX A6000",
|
| 49 |
+
"memoryTotal": "51527024640",
|
| 50 |
+
"cudaCores": 10752,
|
| 51 |
+
"architecture": "Ampere"
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"name": "NVIDIA RTX A6000",
|
| 55 |
+
"memoryTotal": "51527024640",
|
| 56 |
+
"cudaCores": 10752,
|
| 57 |
+
"architecture": "Ampere"
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"name": "NVIDIA RTX A6000",
|
| 61 |
+
"memoryTotal": "51527024640",
|
| 62 |
+
"cudaCores": 10752,
|
| 63 |
+
"architecture": "Ampere"
|
| 64 |
+
},
|
| 65 |
+
{
|
| 66 |
+
"name": "NVIDIA RTX A6000",
|
| 67 |
+
"memoryTotal": "51527024640",
|
| 68 |
+
"cudaCores": 10752,
|
| 69 |
+
"architecture": "Ampere"
|
| 70 |
+
},
|
| 71 |
+
{
|
| 72 |
+
"name": "NVIDIA RTX A6000",
|
| 73 |
+
"memoryTotal": "51527024640",
|
| 74 |
+
"cudaCores": 10752,
|
| 75 |
+
"architecture": "Ampere"
|
| 76 |
+
},
|
| 77 |
+
{
|
| 78 |
+
"name": "NVIDIA RTX A6000",
|
| 79 |
+
"memoryTotal": "51527024640",
|
| 80 |
+
"cudaCores": 10752,
|
| 81 |
+
"architecture": "Ampere"
|
| 82 |
+
},
|
| 83 |
+
{
|
| 84 |
+
"name": "NVIDIA RTX A6000",
|
| 85 |
+
"memoryTotal": "51527024640",
|
| 86 |
+
"cudaCores": 10752,
|
| 87 |
+
"architecture": "Ampere"
|
| 88 |
+
},
|
| 89 |
+
{
|
| 90 |
+
"name": "NVIDIA RTX A6000",
|
| 91 |
+
"memoryTotal": "51527024640",
|
| 92 |
+
"cudaCores": 10752,
|
| 93 |
+
"architecture": "Ampere"
|
| 94 |
+
}
|
| 95 |
+
],
|
| 96 |
+
"cudaVersion": "11.8"
|
| 97 |
+
}
|
wandb/run-20241030_013339-dgadwxty/logs/debug-internal.log
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"time":"2024-10-30T01:33:39.881038903-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
|
| 2 |
+
{"time":"2024-10-30T01:33:39.881052453-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_013339-dgadwxty/logs/debug-core.log"}
|
| 3 |
+
{"time":"2024-10-30T01:33:39.990088507-04:00","level":"INFO","msg":"created new stream","id":"dgadwxty"}
|
| 4 |
+
{"time":"2024-10-30T01:33:39.990129677-04:00","level":"INFO","msg":"stream: started","id":"dgadwxty"}
|
| 5 |
+
{"time":"2024-10-30T01:33:39.990203308-04:00","level":"INFO","msg":"sender: started","stream_id":"dgadwxty"}
|
| 6 |
+
{"time":"2024-10-30T01:33:39.990163337-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"dgadwxty"}}
|
| 7 |
+
{"time":"2024-10-30T01:33:39.990200188-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"dgadwxty"}}
|
| 8 |
+
{"time":"2024-10-30T01:33:40.18682614-04:00","level":"INFO","msg":"Starting system monitor"}
|
wandb/run-20241030_013339-dgadwxty/logs/debug.log
ADDED
|
@@ -0,0 +1,26 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
2024-10-30 01:33:39,874 INFO MainThread:337257 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
|
| 2 |
+
2024-10-30 01:33:39,874 INFO MainThread:337257 [wandb_setup.py:_flush():79] Configure stats pid to 337257
|
| 3 |
+
2024-10-30 01:33:39,875 INFO MainThread:337257 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
|
| 4 |
+
2024-10-30 01:33:39,875 INFO MainThread:337257 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
|
| 5 |
+
2024-10-30 01:33:39,875 INFO MainThread:337257 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
|
| 6 |
+
2024-10-30 01:33:39,875 INFO MainThread:337257 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
|
| 7 |
+
2024-10-30 01:33:39,875 INFO MainThread:337257 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
|
| 8 |
+
2024-10-30 01:33:39,875 INFO MainThread:337257 [wandb_setup.py:_flush():79] Applying login settings: {}
|
| 9 |
+
2024-10-30 01:33:39,875 INFO MainThread:337257 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_013339-dgadwxty/logs/debug.log
|
| 10 |
+
2024-10-30 01:33:39,875 INFO MainThread:337257 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_013339-dgadwxty/logs/debug-internal.log
|
| 11 |
+
2024-10-30 01:33:39,875 INFO MainThread:337257 [wandb_init.py:init():621] calling init triggers
|
| 12 |
+
2024-10-30 01:33:39,875 INFO MainThread:337257 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
|
| 13 |
+
config: {}
|
| 14 |
+
2024-10-30 01:33:39,875 INFO MainThread:337257 [wandb_init.py:init():671] starting backend
|
| 15 |
+
2024-10-30 01:33:39,875 INFO MainThread:337257 [wandb_init.py:init():675] sending inform_init request
|
| 16 |
+
2024-10-30 01:33:39,878 INFO MainThread:337257 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
|
| 17 |
+
2024-10-30 01:33:39,878 INFO MainThread:337257 [wandb_init.py:init():688] backend started and connected
|
| 18 |
+
2024-10-30 01:33:39,882 INFO MainThread:337257 [wandb_init.py:init():783] updated telemetry
|
| 19 |
+
2024-10-30 01:33:39,911 INFO MainThread:337257 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
|
| 20 |
+
2024-10-30 01:33:40,183 INFO MainThread:337257 [wandb_init.py:init():867] starting run threads in backend
|
| 21 |
+
2024-10-30 01:33:40,280 INFO MainThread:337257 [wandb_run.py:_console_start():2463] atexit reg
|
| 22 |
+
2024-10-30 01:33:40,281 INFO MainThread:337257 [wandb_run.py:_redirect():2311] redirect: wrap_raw
|
| 23 |
+
2024-10-30 01:33:40,281 INFO MainThread:337257 [wandb_run.py:_redirect():2376] Wrapping output streams.
|
| 24 |
+
2024-10-30 01:33:40,281 INFO MainThread:337257 [wandb_run.py:_redirect():2401] Redirects installed.
|
| 25 |
+
2024-10-30 01:33:40,282 INFO MainThread:337257 [wandb_init.py:init():911] run started, returning control to user process
|
| 26 |
+
2024-10-30 01:33:40,282 INFO MainThread:337257 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'reverse_full', 'train_set': '10M', 'batch_size': 3, 'epoch': 7, 'seed': 0}
|
wandb/run-20241030_013339-s77qk5li/files/output.log
ADDED
|
@@ -0,0 +1,41 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Downloading shards: 100%|ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ| 2/2 [02:32<00:00, 76.21s/it]
|
| 2 |
+
Loading checkpoint shards: 100%|βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ| 2/2 [00:04<00:00, 2.25s/it]
|
| 3 |
+
Map: 100%|βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ| 17519/17519 [00:56<00:00, 308.15 examples/s]
|
| 4 |
+
Map: 100%|βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ| 18140/18140 [00:57<00:00, 313.18 examples/s]
|
| 5 |
+
tokenized_valid: Dataset({
|
| 6 |
+
features: ['input_ids', 'attention_mask'],
|
| 7 |
+
num_rows: 600
|
| 8 |
+
})
|
| 9 |
+
/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/training_args.py:1545: FutureWarning: `evaluation_strategy` is deprecated and will be removed in version 4.46 of π€ Transformers. Use `eval_strategy` instead
|
| 10 |
+
warnings.warn(
|
| 11 |
+
[2024-10-30 01:38:13,602] [INFO] [real_accelerator.py:219:get_accelerator] Setting ds_accelerator to cuda (auto detect)
|
| 12 |
+
[2024-10-30 01:38:21,321] [INFO] [comm.py:652:init_distributed] cdb=None
|
| 13 |
+
Installed CUDA version 11.8 does not match the version torch was compiled with 11.7 but since the APIs are compatible, accepting this combination
|
| 14 |
+
Using /home/chunhui/.cache/torch_extensions/py39_cu117 as PyTorch extensions root...
|
| 15 |
+
Emitting ninja build file /home/chunhui/.cache/torch_extensions/py39_cu117/cpu_adam/build.ninja...
|
| 16 |
+
Building extension module cpu_adam...
|
| 17 |
+
Allowing ninja to set a default number of workers... (overridable by setting the environment variable MAX_JOBS=N)
|
| 18 |
+
Loading extension module cpu_adam...
|
| 19 |
+
Time to load cpu_adam op: 4.4626922607421875 seconds
|
| 20 |
+
Traceback (most recent call last):
|
| 21 |
+
File "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py", line 219, in <module>
|
| 22 |
+
trainer.train()
|
| 23 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/trainer.py", line 2052, in train
|
| 24 |
+
return inner_training_loop(
|
| 25 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/trainer.py", line 2467, in _inner_training_loop
|
| 26 |
+
self._maybe_log_save_evaluate(tr_loss, grad_norm, model, trial, epoch, ignore_keys_for_eval)
|
| 27 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/trainer.py", line 2918, in _maybe_log_save_evaluate
|
| 28 |
+
self._save_checkpoint(model, trial, metrics=metrics)
|
| 29 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/trainer.py", line 3008, in _save_checkpoint
|
| 30 |
+
self.save_model(output_dir, _internal_call=True)
|
| 31 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/trainer.py", line 3608, in save_model
|
| 32 |
+
state_dict = self.accelerator.get_state_dict(self.deepspeed)
|
| 33 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/accelerate/accelerator.py", line 3382, in get_state_dict
|
| 34 |
+
state_dict = clone_tensors_for_torch_save(self.unwrap_model(model).state_dict())
|
| 35 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/deepspeed/checkpoint/utils.py", line 60, in clone_tensors_for_torch_save
|
| 36 |
+
return type(item)({k: clone_tensors_for_torch_save(v, device) for k, v in item.items()})
|
| 37 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/deepspeed/checkpoint/utils.py", line 60, in <dictcomp>
|
| 38 |
+
return type(item)({k: clone_tensors_for_torch_save(v, device) for k, v in item.items()})
|
| 39 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/deepspeed/checkpoint/utils.py", line 54, in clone_tensors_for_torch_save
|
| 40 |
+
return item.detach().clone().to(device)
|
| 41 |
+
KeyboardInterrupt
|
wandb/run-20241030_013339-s77qk5li/files/requirements.txt
ADDED
|
@@ -0,0 +1,147 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
funcsigs==1.0.2
|
| 2 |
+
sentry-sdk==2.17.0
|
| 3 |
+
multiprocess==0.70.16
|
| 4 |
+
numpy==1.26.2
|
| 5 |
+
pluralizer==1.2.0
|
| 6 |
+
debugpy==1.6.7
|
| 7 |
+
nvidia-cudnn-cu11==8.5.0.96
|
| 8 |
+
deepspeed==0.15.2
|
| 9 |
+
data==0.4
|
| 10 |
+
pandas==2.1.3
|
| 11 |
+
tomli==2.0.1
|
| 12 |
+
charset-normalizer==3.3.2
|
| 13 |
+
attrs==24.2.0
|
| 14 |
+
aiosignal==1.3.1
|
| 15 |
+
fsspec==2023.10.0
|
| 16 |
+
nvidia-cusparse-cu11==11.7.4.91
|
| 17 |
+
zipp==3.12.0
|
| 18 |
+
mypy-extensions==1.0.0
|
| 19 |
+
datasets==3.0.1
|
| 20 |
+
joblib==1.3.2
|
| 21 |
+
hjson==3.1.0
|
| 22 |
+
traitlets==5.7.1
|
| 23 |
+
stack-data==0.6.0
|
| 24 |
+
transformers==4.45.1
|
| 25 |
+
sympy==1.11.1
|
| 26 |
+
Pygments==2.15.0
|
| 27 |
+
docker-pycreds==0.4.0
|
| 28 |
+
dill==0.3.8
|
| 29 |
+
wheel==0.44.0
|
| 30 |
+
prompt-toolkit==3.0.30
|
| 31 |
+
parso==0.8.3
|
| 32 |
+
ipykernel==6.23.1
|
| 33 |
+
pyarrow==17.0.0
|
| 34 |
+
certifi==2023.11.17
|
| 35 |
+
nvidia-cufft-cu11==10.9.0.58
|
| 36 |
+
six==1.16.0
|
| 37 |
+
pydantic==2.9.2
|
| 38 |
+
click==8.1.7
|
| 39 |
+
nest-asyncio==1.5.6
|
| 40 |
+
gmpy2==2.1.0
|
| 41 |
+
matplotlib==3.8.2
|
| 42 |
+
scipy==1.11.4
|
| 43 |
+
typing_extensions==4.12.2
|
| 44 |
+
statsmodels==0.14.0
|
| 45 |
+
huggingface-hub==0.25.0
|
| 46 |
+
frozenlist==1.4.1
|
| 47 |
+
gpustat==1.1.1
|
| 48 |
+
nvidia-nvtx-cu11==11.7.91
|
| 49 |
+
safetensors==0.4.5
|
| 50 |
+
stanza==1.9.2
|
| 51 |
+
decorator==5.1.1
|
| 52 |
+
seaborn==0.13.0
|
| 53 |
+
sentencepiece==0.2.0
|
| 54 |
+
PyYAML==6.0.1
|
| 55 |
+
black==24.8.0
|
| 56 |
+
protobuf==4.25.1
|
| 57 |
+
pickleshare==0.7.5
|
| 58 |
+
peft==0.13.0
|
| 59 |
+
triton==2.0.0
|
| 60 |
+
nvidia-cuda-runtime-cu11==11.7.99
|
| 61 |
+
Jinja2==3.1.2
|
| 62 |
+
nvidia-cusolver-cu11==11.4.0.1
|
| 63 |
+
executing==1.2.0
|
| 64 |
+
jupyter_client==8.1.0
|
| 65 |
+
pluggy==1.3.0
|
| 66 |
+
cmake==3.30.3
|
| 67 |
+
pytz==2023.3.post1
|
| 68 |
+
aiohappyeyeballs==2.4.2
|
| 69 |
+
kiwisolver==1.4.5
|
| 70 |
+
py-cpuinfo==9.0.0
|
| 71 |
+
Pillow==10.1.0
|
| 72 |
+
ptyprocess==0.7.0
|
| 73 |
+
importlib_resources==6.4.5
|
| 74 |
+
GitPython==3.1.43
|
| 75 |
+
importlib-metadata==6.0.0
|
| 76 |
+
iniconfig==2.0.0
|
| 77 |
+
scikit-learn==1.3.2
|
| 78 |
+
exceptiongroup==1.1.0
|
| 79 |
+
networkx==2.8.6
|
| 80 |
+
accelerate==1.0.0
|
| 81 |
+
nltk==3.8.1
|
| 82 |
+
shutilwhich==1.1.0
|
| 83 |
+
fonttools==4.45.1
|
| 84 |
+
future==0.18.3
|
| 85 |
+
aiohttp==3.10.6
|
| 86 |
+
wcwidth==0.2.5
|
| 87 |
+
idna==3.6
|
| 88 |
+
filelock==3.12.2
|
| 89 |
+
pathspec==0.12.1
|
| 90 |
+
jupyter_core==5.1.0
|
| 91 |
+
lit==18.1.8
|
| 92 |
+
nvidia-curand-cu11==10.2.10.91
|
| 93 |
+
nvidia-cublas-cu11==11.10.3.66
|
| 94 |
+
nvidia-ml-py==12.560.30
|
| 95 |
+
msgpack==1.1.0
|
| 96 |
+
python-dateutil==2.8.2
|
| 97 |
+
blessed==1.20.0
|
| 98 |
+
packaging==23.0
|
| 99 |
+
gitdb==4.0.11
|
| 100 |
+
yarl==1.13.0
|
| 101 |
+
emoji==2.8.0
|
| 102 |
+
tzdata==2023.3
|
| 103 |
+
cycler==0.12.1
|
| 104 |
+
tornado==6.2
|
| 105 |
+
backcall==0.2.0
|
| 106 |
+
plotnine==0.12.4
|
| 107 |
+
ninja==1.11.1.1
|
| 108 |
+
latex==0.7.0
|
| 109 |
+
wandb==0.18.5
|
| 110 |
+
setproctitle==1.3.3
|
| 111 |
+
threadpoolctl==3.2.0
|
| 112 |
+
requests==2.32.3
|
| 113 |
+
pyparsing==3.1.1
|
| 114 |
+
smmap==5.0.1
|
| 115 |
+
pyzmq==23.0.0
|
| 116 |
+
async-timeout==4.0.3
|
| 117 |
+
annotated-types==0.7.0
|
| 118 |
+
matplotlib-inline==0.1.6
|
| 119 |
+
latexcodec==1.0.0
|
| 120 |
+
ipython==8.0.0
|
| 121 |
+
patsy==0.5.3
|
| 122 |
+
contourpy==1.2.0
|
| 123 |
+
multidict==6.1.0
|
| 124 |
+
mizani==0.9.3
|
| 125 |
+
urllib3==2.1.0
|
| 126 |
+
tokenizers==0.20.0
|
| 127 |
+
MarkupSafe==2.1.2
|
| 128 |
+
pip==24.2
|
| 129 |
+
pexpect==4.8.0
|
| 130 |
+
tqdm==4.66.5
|
| 131 |
+
jedi==0.18.2
|
| 132 |
+
pydantic_core==2.23.4
|
| 133 |
+
tempdir==0.7.1
|
| 134 |
+
mpmath==1.2.1
|
| 135 |
+
setuptools==72.1.0
|
| 136 |
+
pytest==7.4.3
|
| 137 |
+
pure-eval==0.2.2
|
| 138 |
+
psutil==5.9.1
|
| 139 |
+
comm==0.1.2
|
| 140 |
+
nvidia-cuda-cupti-cu11==11.7.101
|
| 141 |
+
nvidia-cuda-nvrtc-cu11==11.7.99
|
| 142 |
+
regex==2023.10.3
|
| 143 |
+
platformdirs==2.5.2
|
| 144 |
+
asttokens==2.2.1
|
| 145 |
+
torch==2.0.0
|
| 146 |
+
nvidia-nccl-cu11==2.14.3
|
| 147 |
+
xxhash==3.5.0
|
wandb/run-20241030_013339-s77qk5li/files/wandb-metadata.json
ADDED
|
@@ -0,0 +1,97 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
|
| 3 |
+
"python": "3.9.19",
|
| 4 |
+
"startedAt": "2024-10-30T05:33:39.943616Z",
|
| 5 |
+
"args": [
|
| 6 |
+
"--perturbation",
|
| 7 |
+
"reverse_full",
|
| 8 |
+
"--train_set",
|
| 9 |
+
"10M",
|
| 10 |
+
"--batch_size",
|
| 11 |
+
"3",
|
| 12 |
+
"--epoch",
|
| 13 |
+
"7",
|
| 14 |
+
"--seed",
|
| 15 |
+
"0"
|
| 16 |
+
],
|
| 17 |
+
"program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
|
| 18 |
+
"codePath": "train/train_deep_wandb.py",
|
| 19 |
+
"git": {
|
| 20 |
+
"remote": "git@hf.co:Yaning1001/Impossible_llm.git",
|
| 21 |
+
"commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
|
| 22 |
+
},
|
| 23 |
+
"email": "yaning1001@gmail.com",
|
| 24 |
+
"root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
|
| 25 |
+
"host": "mms-large-2",
|
| 26 |
+
"username": "chunhui",
|
| 27 |
+
"executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
|
| 28 |
+
"codePathLocal": "train_deep_wandb.py",
|
| 29 |
+
"cpu_count": 32,
|
| 30 |
+
"cpu_count_logical": 64,
|
| 31 |
+
"gpu": "NVIDIA RTX A6000",
|
| 32 |
+
"gpu_count": 8,
|
| 33 |
+
"disk": {
|
| 34 |
+
"/": {
|
| 35 |
+
"total": "1888559353856",
|
| 36 |
+
"used": "1710081839104"
|
| 37 |
+
}
|
| 38 |
+
},
|
| 39 |
+
"memory": {
|
| 40 |
+
"total": "202617098240"
|
| 41 |
+
},
|
| 42 |
+
"cpu": {
|
| 43 |
+
"count": 32,
|
| 44 |
+
"countLogical": 64
|
| 45 |
+
},
|
| 46 |
+
"gpu_nvidia": [
|
| 47 |
+
{
|
| 48 |
+
"name": "NVIDIA RTX A6000",
|
| 49 |
+
"memoryTotal": "51527024640",
|
| 50 |
+
"cudaCores": 10752,
|
| 51 |
+
"architecture": "Ampere"
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"name": "NVIDIA RTX A6000",
|
| 55 |
+
"memoryTotal": "51527024640",
|
| 56 |
+
"cudaCores": 10752,
|
| 57 |
+
"architecture": "Ampere"
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"name": "NVIDIA RTX A6000",
|
| 61 |
+
"memoryTotal": "51527024640",
|
| 62 |
+
"cudaCores": 10752,
|
| 63 |
+
"architecture": "Ampere"
|
| 64 |
+
},
|
| 65 |
+
{
|
| 66 |
+
"name": "NVIDIA RTX A6000",
|
| 67 |
+
"memoryTotal": "51527024640",
|
| 68 |
+
"cudaCores": 10752,
|
| 69 |
+
"architecture": "Ampere"
|
| 70 |
+
},
|
| 71 |
+
{
|
| 72 |
+
"name": "NVIDIA RTX A6000",
|
| 73 |
+
"memoryTotal": "51527024640",
|
| 74 |
+
"cudaCores": 10752,
|
| 75 |
+
"architecture": "Ampere"
|
| 76 |
+
},
|
| 77 |
+
{
|
| 78 |
+
"name": "NVIDIA RTX A6000",
|
| 79 |
+
"memoryTotal": "51527024640",
|
| 80 |
+
"cudaCores": 10752,
|
| 81 |
+
"architecture": "Ampere"
|
| 82 |
+
},
|
| 83 |
+
{
|
| 84 |
+
"name": "NVIDIA RTX A6000",
|
| 85 |
+
"memoryTotal": "51527024640",
|
| 86 |
+
"cudaCores": 10752,
|
| 87 |
+
"architecture": "Ampere"
|
| 88 |
+
},
|
| 89 |
+
{
|
| 90 |
+
"name": "NVIDIA RTX A6000",
|
| 91 |
+
"memoryTotal": "51527024640",
|
| 92 |
+
"cudaCores": 10752,
|
| 93 |
+
"architecture": "Ampere"
|
| 94 |
+
}
|
| 95 |
+
],
|
| 96 |
+
"cudaVersion": "11.8"
|
| 97 |
+
}
|
wandb/run-20241030_013339-s77qk5li/logs/debug-internal.log
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"time":"2024-10-30T01:33:39.945853369-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
|
| 2 |
+
{"time":"2024-10-30T01:33:39.945865199-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_013339-s77qk5li/logs/debug-core.log"}
|
| 3 |
+
{"time":"2024-10-30T01:33:40.053085091-04:00","level":"INFO","msg":"created new stream","id":"s77qk5li"}
|
| 4 |
+
{"time":"2024-10-30T01:33:40.053121041-04:00","level":"INFO","msg":"stream: started","id":"s77qk5li"}
|
| 5 |
+
{"time":"2024-10-30T01:33:40.053172051-04:00","level":"INFO","msg":"sender: started","stream_id":"s77qk5li"}
|
| 6 |
+
{"time":"2024-10-30T01:33:40.053132691-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"s77qk5li"}}
|
| 7 |
+
{"time":"2024-10-30T01:33:40.053160001-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"s77qk5li"}}
|
| 8 |
+
{"time":"2024-10-30T01:33:40.209477562-04:00","level":"INFO","msg":"Starting system monitor"}
|
wandb/run-20241030_013339-s77qk5li/logs/debug.log
ADDED
|
@@ -0,0 +1,26 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
2024-10-30 01:33:39,941 INFO MainThread:337256 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
|
| 2 |
+
2024-10-30 01:33:39,941 INFO MainThread:337256 [wandb_setup.py:_flush():79] Configure stats pid to 337256
|
| 3 |
+
2024-10-30 01:33:39,941 INFO MainThread:337256 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
|
| 4 |
+
2024-10-30 01:33:39,941 INFO MainThread:337256 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
|
| 5 |
+
2024-10-30 01:33:39,942 INFO MainThread:337256 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
|
| 6 |
+
2024-10-30 01:33:39,942 INFO MainThread:337256 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
|
| 7 |
+
2024-10-30 01:33:39,942 INFO MainThread:337256 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
|
| 8 |
+
2024-10-30 01:33:39,942 INFO MainThread:337256 [wandb_setup.py:_flush():79] Applying login settings: {}
|
| 9 |
+
2024-10-30 01:33:39,942 INFO MainThread:337256 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_013339-s77qk5li/logs/debug.log
|
| 10 |
+
2024-10-30 01:33:39,942 INFO MainThread:337256 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_013339-s77qk5li/logs/debug-internal.log
|
| 11 |
+
2024-10-30 01:33:39,942 INFO MainThread:337256 [wandb_init.py:init():621] calling init triggers
|
| 12 |
+
2024-10-30 01:33:39,942 INFO MainThread:337256 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
|
| 13 |
+
config: {}
|
| 14 |
+
2024-10-30 01:33:39,942 INFO MainThread:337256 [wandb_init.py:init():671] starting backend
|
| 15 |
+
2024-10-30 01:33:39,942 INFO MainThread:337256 [wandb_init.py:init():675] sending inform_init request
|
| 16 |
+
2024-10-30 01:33:39,943 INFO MainThread:337256 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
|
| 17 |
+
2024-10-30 01:33:39,943 INFO MainThread:337256 [wandb_init.py:init():688] backend started and connected
|
| 18 |
+
2024-10-30 01:33:39,946 INFO MainThread:337256 [wandb_init.py:init():783] updated telemetry
|
| 19 |
+
2024-10-30 01:33:39,977 INFO MainThread:337256 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
|
| 20 |
+
2024-10-30 01:33:40,206 INFO MainThread:337256 [wandb_init.py:init():867] starting run threads in backend
|
| 21 |
+
2024-10-30 01:33:40,298 INFO MainThread:337256 [wandb_run.py:_console_start():2463] atexit reg
|
| 22 |
+
2024-10-30 01:33:40,298 INFO MainThread:337256 [wandb_run.py:_redirect():2311] redirect: wrap_raw
|
| 23 |
+
2024-10-30 01:33:40,298 INFO MainThread:337256 [wandb_run.py:_redirect():2376] Wrapping output streams.
|
| 24 |
+
2024-10-30 01:33:40,298 INFO MainThread:337256 [wandb_run.py:_redirect():2401] Redirects installed.
|
| 25 |
+
2024-10-30 01:33:40,300 INFO MainThread:337256 [wandb_init.py:init():911] run started, returning control to user process
|
| 26 |
+
2024-10-30 01:33:40,300 INFO MainThread:337256 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'reverse_full', 'train_set': '10M', 'batch_size': 3, 'epoch': 7, 'seed': 0}
|
wandb/run-20241030_112700-jhzkfwvw/files/config.yaml
ADDED
|
@@ -0,0 +1,47 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
_wandb:
|
| 2 |
+
value:
|
| 3 |
+
cli_version: 0.18.5
|
| 4 |
+
m: []
|
| 5 |
+
python_version: 3.9.19
|
| 6 |
+
t:
|
| 7 |
+
"1":
|
| 8 |
+
- 1
|
| 9 |
+
- 5
|
| 10 |
+
- 11
|
| 11 |
+
- 49
|
| 12 |
+
- 51
|
| 13 |
+
- 53
|
| 14 |
+
- 55
|
| 15 |
+
- 71
|
| 16 |
+
- 98
|
| 17 |
+
"2":
|
| 18 |
+
- 1
|
| 19 |
+
- 5
|
| 20 |
+
- 11
|
| 21 |
+
- 49
|
| 22 |
+
- 51
|
| 23 |
+
- 53
|
| 24 |
+
- 55
|
| 25 |
+
- 71
|
| 26 |
+
- 98
|
| 27 |
+
"3":
|
| 28 |
+
- 13
|
| 29 |
+
- 23
|
| 30 |
+
- 55
|
| 31 |
+
"4": 3.9.19
|
| 32 |
+
"5": 0.18.5
|
| 33 |
+
"6": 4.45.1
|
| 34 |
+
"8":
|
| 35 |
+
- 5
|
| 36 |
+
"12": 0.18.5
|
| 37 |
+
"13": linux-x86_64
|
| 38 |
+
batch_size:
|
| 39 |
+
value: 3
|
| 40 |
+
epoch:
|
| 41 |
+
value: 3
|
| 42 |
+
perturbation:
|
| 43 |
+
value: reverse_control
|
| 44 |
+
seed:
|
| 45 |
+
value: 0
|
| 46 |
+
train_set:
|
| 47 |
+
value: 10M
|
wandb/run-20241030_112700-jhzkfwvw/files/output.log
ADDED
|
@@ -0,0 +1,34 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Downloading shards: 0%| | 0/2 [01:32<?, ?it/s]
|
| 2 |
+
Error in sys.excepthook:
|
| 3 |
+
Traceback (most recent call last):
|
| 4 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/wandb/sdk/lib/exit_hooks.py", line 41, in exc_handler
|
| 5 |
+
def exc_handler(
|
| 6 |
+
KeyboardInterrupt
|
| 7 |
+
|
| 8 |
+
Original exception was:
|
| 9 |
+
Traceback (most recent call last):
|
| 10 |
+
File "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py", line 172, in <module>
|
| 11 |
+
model = AutoModelForCausalLM.from_pretrained(model_name,
|
| 12 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/models/auto/auto_factory.py", line 564, in from_pretrained
|
| 13 |
+
return model_class.from_pretrained(
|
| 14 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/modeling_utils.py", line 3769, in from_pretrained
|
| 15 |
+
resolved_archive_file, sharded_metadata = get_checkpoint_shard_files(
|
| 16 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/utils/hub.py", line 1098, in get_checkpoint_shard_files
|
| 17 |
+
cached_filename = cached_file(
|
| 18 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/utils/hub.py", line 403, in cached_file
|
| 19 |
+
resolved_file = hf_hub_download(
|
| 20 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/utils/_deprecation.py", line 101, in inner_f
|
| 21 |
+
return f(*args, **kwargs)
|
| 22 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/utils/_validators.py", line 114, in _inner_fn
|
| 23 |
+
return fn(*args, **kwargs)
|
| 24 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/file_download.py", line 1232, in hf_hub_download
|
| 25 |
+
return _hf_hub_download_to_cache_dir(
|
| 26 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/file_download.py", line 1380, in _hf_hub_download_to_cache_dir
|
| 27 |
+
with WeakFileLock(lock_path):
|
| 28 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/contextlib.py", line 119, in __enter__
|
| 29 |
+
return next(self.gen)
|
| 30 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/utils/_fixes.py", line 98, in WeakFileLock
|
| 31 |
+
lock.acquire()
|
| 32 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/filelock/_api.py", line 225, in acquire
|
| 33 |
+
time.sleep(poll_interval)
|
| 34 |
+
KeyboardInterrupt
|
wandb/run-20241030_112700-jhzkfwvw/files/requirements.txt
ADDED
|
@@ -0,0 +1,147 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
funcsigs==1.0.2
|
| 2 |
+
sentry-sdk==2.17.0
|
| 3 |
+
multiprocess==0.70.16
|
| 4 |
+
numpy==1.26.2
|
| 5 |
+
pluralizer==1.2.0
|
| 6 |
+
debugpy==1.6.7
|
| 7 |
+
nvidia-cudnn-cu11==8.5.0.96
|
| 8 |
+
deepspeed==0.15.2
|
| 9 |
+
data==0.4
|
| 10 |
+
pandas==2.1.3
|
| 11 |
+
tomli==2.0.1
|
| 12 |
+
charset-normalizer==3.3.2
|
| 13 |
+
attrs==24.2.0
|
| 14 |
+
aiosignal==1.3.1
|
| 15 |
+
fsspec==2023.10.0
|
| 16 |
+
nvidia-cusparse-cu11==11.7.4.91
|
| 17 |
+
zipp==3.12.0
|
| 18 |
+
mypy-extensions==1.0.0
|
| 19 |
+
datasets==3.0.1
|
| 20 |
+
joblib==1.3.2
|
| 21 |
+
hjson==3.1.0
|
| 22 |
+
traitlets==5.7.1
|
| 23 |
+
stack-data==0.6.0
|
| 24 |
+
transformers==4.45.1
|
| 25 |
+
sympy==1.11.1
|
| 26 |
+
Pygments==2.15.0
|
| 27 |
+
docker-pycreds==0.4.0
|
| 28 |
+
dill==0.3.8
|
| 29 |
+
wheel==0.44.0
|
| 30 |
+
prompt-toolkit==3.0.30
|
| 31 |
+
parso==0.8.3
|
| 32 |
+
ipykernel==6.23.1
|
| 33 |
+
pyarrow==17.0.0
|
| 34 |
+
certifi==2023.11.17
|
| 35 |
+
nvidia-cufft-cu11==10.9.0.58
|
| 36 |
+
six==1.16.0
|
| 37 |
+
pydantic==2.9.2
|
| 38 |
+
click==8.1.7
|
| 39 |
+
nest-asyncio==1.5.6
|
| 40 |
+
gmpy2==2.1.0
|
| 41 |
+
matplotlib==3.8.2
|
| 42 |
+
scipy==1.11.4
|
| 43 |
+
typing_extensions==4.12.2
|
| 44 |
+
statsmodels==0.14.0
|
| 45 |
+
huggingface-hub==0.25.0
|
| 46 |
+
frozenlist==1.4.1
|
| 47 |
+
gpustat==1.1.1
|
| 48 |
+
nvidia-nvtx-cu11==11.7.91
|
| 49 |
+
safetensors==0.4.5
|
| 50 |
+
stanza==1.9.2
|
| 51 |
+
decorator==5.1.1
|
| 52 |
+
seaborn==0.13.0
|
| 53 |
+
sentencepiece==0.2.0
|
| 54 |
+
PyYAML==6.0.1
|
| 55 |
+
black==24.8.0
|
| 56 |
+
protobuf==4.25.1
|
| 57 |
+
pickleshare==0.7.5
|
| 58 |
+
peft==0.13.0
|
| 59 |
+
triton==2.0.0
|
| 60 |
+
nvidia-cuda-runtime-cu11==11.7.99
|
| 61 |
+
Jinja2==3.1.2
|
| 62 |
+
nvidia-cusolver-cu11==11.4.0.1
|
| 63 |
+
executing==1.2.0
|
| 64 |
+
jupyter_client==8.1.0
|
| 65 |
+
pluggy==1.3.0
|
| 66 |
+
cmake==3.30.3
|
| 67 |
+
pytz==2023.3.post1
|
| 68 |
+
aiohappyeyeballs==2.4.2
|
| 69 |
+
kiwisolver==1.4.5
|
| 70 |
+
py-cpuinfo==9.0.0
|
| 71 |
+
Pillow==10.1.0
|
| 72 |
+
ptyprocess==0.7.0
|
| 73 |
+
importlib_resources==6.4.5
|
| 74 |
+
GitPython==3.1.43
|
| 75 |
+
importlib-metadata==6.0.0
|
| 76 |
+
iniconfig==2.0.0
|
| 77 |
+
scikit-learn==1.3.2
|
| 78 |
+
exceptiongroup==1.1.0
|
| 79 |
+
networkx==2.8.6
|
| 80 |
+
accelerate==1.0.0
|
| 81 |
+
nltk==3.8.1
|
| 82 |
+
shutilwhich==1.1.0
|
| 83 |
+
fonttools==4.45.1
|
| 84 |
+
future==0.18.3
|
| 85 |
+
aiohttp==3.10.6
|
| 86 |
+
wcwidth==0.2.5
|
| 87 |
+
idna==3.6
|
| 88 |
+
filelock==3.12.2
|
| 89 |
+
pathspec==0.12.1
|
| 90 |
+
jupyter_core==5.1.0
|
| 91 |
+
lit==18.1.8
|
| 92 |
+
nvidia-curand-cu11==10.2.10.91
|
| 93 |
+
nvidia-cublas-cu11==11.10.3.66
|
| 94 |
+
nvidia-ml-py==12.560.30
|
| 95 |
+
msgpack==1.1.0
|
| 96 |
+
python-dateutil==2.8.2
|
| 97 |
+
blessed==1.20.0
|
| 98 |
+
packaging==23.0
|
| 99 |
+
gitdb==4.0.11
|
| 100 |
+
yarl==1.13.0
|
| 101 |
+
emoji==2.8.0
|
| 102 |
+
tzdata==2023.3
|
| 103 |
+
cycler==0.12.1
|
| 104 |
+
tornado==6.2
|
| 105 |
+
backcall==0.2.0
|
| 106 |
+
plotnine==0.12.4
|
| 107 |
+
ninja==1.11.1.1
|
| 108 |
+
latex==0.7.0
|
| 109 |
+
wandb==0.18.5
|
| 110 |
+
setproctitle==1.3.3
|
| 111 |
+
threadpoolctl==3.2.0
|
| 112 |
+
requests==2.32.3
|
| 113 |
+
pyparsing==3.1.1
|
| 114 |
+
smmap==5.0.1
|
| 115 |
+
pyzmq==23.0.0
|
| 116 |
+
async-timeout==4.0.3
|
| 117 |
+
annotated-types==0.7.0
|
| 118 |
+
matplotlib-inline==0.1.6
|
| 119 |
+
latexcodec==1.0.0
|
| 120 |
+
ipython==8.0.0
|
| 121 |
+
patsy==0.5.3
|
| 122 |
+
contourpy==1.2.0
|
| 123 |
+
multidict==6.1.0
|
| 124 |
+
mizani==0.9.3
|
| 125 |
+
urllib3==2.1.0
|
| 126 |
+
tokenizers==0.20.0
|
| 127 |
+
MarkupSafe==2.1.2
|
| 128 |
+
pip==24.2
|
| 129 |
+
pexpect==4.8.0
|
| 130 |
+
tqdm==4.66.5
|
| 131 |
+
jedi==0.18.2
|
| 132 |
+
pydantic_core==2.23.4
|
| 133 |
+
tempdir==0.7.1
|
| 134 |
+
mpmath==1.2.1
|
| 135 |
+
setuptools==72.1.0
|
| 136 |
+
pytest==7.4.3
|
| 137 |
+
pure-eval==0.2.2
|
| 138 |
+
psutil==5.9.1
|
| 139 |
+
comm==0.1.2
|
| 140 |
+
nvidia-cuda-cupti-cu11==11.7.101
|
| 141 |
+
nvidia-cuda-nvrtc-cu11==11.7.99
|
| 142 |
+
regex==2023.10.3
|
| 143 |
+
platformdirs==2.5.2
|
| 144 |
+
asttokens==2.2.1
|
| 145 |
+
torch==2.0.0
|
| 146 |
+
nvidia-nccl-cu11==2.14.3
|
| 147 |
+
xxhash==3.5.0
|
wandb/run-20241030_112700-jhzkfwvw/files/wandb-metadata.json
ADDED
|
@@ -0,0 +1,97 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
|
| 3 |
+
"python": "3.9.19",
|
| 4 |
+
"startedAt": "2024-10-30T15:27:00.613781Z",
|
| 5 |
+
"args": [
|
| 6 |
+
"--perturbation",
|
| 7 |
+
"reverse_control",
|
| 8 |
+
"--train_set",
|
| 9 |
+
"10M",
|
| 10 |
+
"--batch_size",
|
| 11 |
+
"3",
|
| 12 |
+
"--epoch",
|
| 13 |
+
"3",
|
| 14 |
+
"--seed",
|
| 15 |
+
"0"
|
| 16 |
+
],
|
| 17 |
+
"program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
|
| 18 |
+
"codePath": "train/train_deep_wandb.py",
|
| 19 |
+
"git": {
|
| 20 |
+
"remote": "git@hf.co:Yaning1001/Impossible_llm.git",
|
| 21 |
+
"commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
|
| 22 |
+
},
|
| 23 |
+
"email": "yaning1001@gmail.com",
|
| 24 |
+
"root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
|
| 25 |
+
"host": "mms-large-2",
|
| 26 |
+
"username": "chunhui",
|
| 27 |
+
"executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
|
| 28 |
+
"codePathLocal": "train_deep_wandb.py",
|
| 29 |
+
"cpu_count": 32,
|
| 30 |
+
"cpu_count_logical": 64,
|
| 31 |
+
"gpu": "NVIDIA RTX A6000",
|
| 32 |
+
"gpu_count": 8,
|
| 33 |
+
"disk": {
|
| 34 |
+
"/": {
|
| 35 |
+
"total": "1888559353856",
|
| 36 |
+
"used": "1710831087616"
|
| 37 |
+
}
|
| 38 |
+
},
|
| 39 |
+
"memory": {
|
| 40 |
+
"total": "202617098240"
|
| 41 |
+
},
|
| 42 |
+
"cpu": {
|
| 43 |
+
"count": 32,
|
| 44 |
+
"countLogical": 64
|
| 45 |
+
},
|
| 46 |
+
"gpu_nvidia": [
|
| 47 |
+
{
|
| 48 |
+
"name": "NVIDIA RTX A6000",
|
| 49 |
+
"memoryTotal": "51527024640",
|
| 50 |
+
"cudaCores": 10752,
|
| 51 |
+
"architecture": "Ampere"
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"name": "NVIDIA RTX A6000",
|
| 55 |
+
"memoryTotal": "51527024640",
|
| 56 |
+
"cudaCores": 10752,
|
| 57 |
+
"architecture": "Ampere"
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"name": "NVIDIA RTX A6000",
|
| 61 |
+
"memoryTotal": "51527024640",
|
| 62 |
+
"cudaCores": 10752,
|
| 63 |
+
"architecture": "Ampere"
|
| 64 |
+
},
|
| 65 |
+
{
|
| 66 |
+
"name": "NVIDIA RTX A6000",
|
| 67 |
+
"memoryTotal": "51527024640",
|
| 68 |
+
"cudaCores": 10752,
|
| 69 |
+
"architecture": "Ampere"
|
| 70 |
+
},
|
| 71 |
+
{
|
| 72 |
+
"name": "NVIDIA RTX A6000",
|
| 73 |
+
"memoryTotal": "51527024640",
|
| 74 |
+
"cudaCores": 10752,
|
| 75 |
+
"architecture": "Ampere"
|
| 76 |
+
},
|
| 77 |
+
{
|
| 78 |
+
"name": "NVIDIA RTX A6000",
|
| 79 |
+
"memoryTotal": "51527024640",
|
| 80 |
+
"cudaCores": 10752,
|
| 81 |
+
"architecture": "Ampere"
|
| 82 |
+
},
|
| 83 |
+
{
|
| 84 |
+
"name": "NVIDIA RTX A6000",
|
| 85 |
+
"memoryTotal": "51527024640",
|
| 86 |
+
"cudaCores": 10752,
|
| 87 |
+
"architecture": "Ampere"
|
| 88 |
+
},
|
| 89 |
+
{
|
| 90 |
+
"name": "NVIDIA RTX A6000",
|
| 91 |
+
"memoryTotal": "51527024640",
|
| 92 |
+
"cudaCores": 10752,
|
| 93 |
+
"architecture": "Ampere"
|
| 94 |
+
}
|
| 95 |
+
],
|
| 96 |
+
"cudaVersion": "11.8"
|
| 97 |
+
}
|
wandb/run-20241030_112700-jhzkfwvw/files/wandb-summary.json
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
{"_wandb":{"runtime":93}}
|
wandb/run-20241030_112700-jhzkfwvw/logs/debug-internal.log
ADDED
|
@@ -0,0 +1,11 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"time":"2024-10-30T11:27:00.617213638-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
|
| 2 |
+
{"time":"2024-10-30T11:27:00.617233858-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_112700-jhzkfwvw/logs/debug-core.log"}
|
| 3 |
+
{"time":"2024-10-30T11:27:00.727987118-04:00","level":"INFO","msg":"created new stream","id":"jhzkfwvw"}
|
| 4 |
+
{"time":"2024-10-30T11:27:00.728057478-04:00","level":"INFO","msg":"stream: started","id":"jhzkfwvw"}
|
| 5 |
+
{"time":"2024-10-30T11:27:00.728115208-04:00","level":"INFO","msg":"sender: started","stream_id":"jhzkfwvw"}
|
| 6 |
+
{"time":"2024-10-30T11:27:00.728102078-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"jhzkfwvw"}}
|
| 7 |
+
{"time":"2024-10-30T11:27:00.728090568-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"jhzkfwvw"}}
|
| 8 |
+
{"time":"2024-10-30T11:27:00.93145688-04:00","level":"INFO","msg":"Starting system monitor"}
|
| 9 |
+
{"time":"2024-10-30T11:28:34.16243391-04:00","level":"INFO","msg":"stream: closing","id":"jhzkfwvw"}
|
| 10 |
+
{"time":"2024-10-30T11:28:34.16248596-04:00","level":"INFO","msg":"Stopping system monitor"}
|
| 11 |
+
{"time":"2024-10-30T11:28:34.163735668-04:00","level":"INFO","msg":"Stopped system monitor"}
|
wandb/run-20241030_112700-jhzkfwvw/logs/debug.log
ADDED
|
@@ -0,0 +1,27 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
2024-10-30 11:27:00,611 INFO MainThread:366800 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
|
| 2 |
+
2024-10-30 11:27:00,611 INFO MainThread:366800 [wandb_setup.py:_flush():79] Configure stats pid to 366800
|
| 3 |
+
2024-10-30 11:27:00,611 INFO MainThread:366800 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
|
| 4 |
+
2024-10-30 11:27:00,611 INFO MainThread:366800 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
|
| 5 |
+
2024-10-30 11:27:00,611 INFO MainThread:366800 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
|
| 6 |
+
2024-10-30 11:27:00,611 INFO MainThread:366800 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
|
| 7 |
+
2024-10-30 11:27:00,611 INFO MainThread:366800 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
|
| 8 |
+
2024-10-30 11:27:00,611 INFO MainThread:366800 [wandb_setup.py:_flush():79] Applying login settings: {}
|
| 9 |
+
2024-10-30 11:27:00,611 INFO MainThread:366800 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_112700-jhzkfwvw/logs/debug.log
|
| 10 |
+
2024-10-30 11:27:00,611 INFO MainThread:366800 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_112700-jhzkfwvw/logs/debug-internal.log
|
| 11 |
+
2024-10-30 11:27:00,611 INFO MainThread:366800 [wandb_init.py:init():621] calling init triggers
|
| 12 |
+
2024-10-30 11:27:00,611 INFO MainThread:366800 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
|
| 13 |
+
config: {}
|
| 14 |
+
2024-10-30 11:27:00,611 INFO MainThread:366800 [wandb_init.py:init():671] starting backend
|
| 15 |
+
2024-10-30 11:27:00,611 INFO MainThread:366800 [wandb_init.py:init():675] sending inform_init request
|
| 16 |
+
2024-10-30 11:27:00,613 INFO MainThread:366800 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
|
| 17 |
+
2024-10-30 11:27:00,613 INFO MainThread:366800 [wandb_init.py:init():688] backend started and connected
|
| 18 |
+
2024-10-30 11:27:00,616 INFO MainThread:366800 [wandb_init.py:init():783] updated telemetry
|
| 19 |
+
2024-10-30 11:27:00,651 INFO MainThread:366800 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
|
| 20 |
+
2024-10-30 11:27:00,928 INFO MainThread:366800 [wandb_init.py:init():867] starting run threads in backend
|
| 21 |
+
2024-10-30 11:27:01,055 INFO MainThread:366800 [wandb_run.py:_console_start():2463] atexit reg
|
| 22 |
+
2024-10-30 11:27:01,056 INFO MainThread:366800 [wandb_run.py:_redirect():2311] redirect: wrap_raw
|
| 23 |
+
2024-10-30 11:27:01,056 INFO MainThread:366800 [wandb_run.py:_redirect():2376] Wrapping output streams.
|
| 24 |
+
2024-10-30 11:27:01,056 INFO MainThread:366800 [wandb_run.py:_redirect():2401] Redirects installed.
|
| 25 |
+
2024-10-30 11:27:01,057 INFO MainThread:366800 [wandb_init.py:init():911] run started, returning control to user process
|
| 26 |
+
2024-10-30 11:27:01,058 INFO MainThread:366800 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'reverse_control', 'train_set': '10M', 'batch_size': 3, 'epoch': 3, 'seed': 0}
|
| 27 |
+
2024-10-30 11:28:34,162 WARNING MsgRouterThr:366800 [router.py:message_loop():77] message_loop has been closed
|
wandb/run-20241030_112700-jhzkfwvw/run-jhzkfwvw.wandb
ADDED
|
Binary file (32.8 kB). View file
|
|
|
wandb/run-20241031_122113-8ldget07/files/config.yaml
ADDED
|
@@ -0,0 +1,50 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
_wandb:
|
| 2 |
+
value:
|
| 3 |
+
cli_version: 0.18.5
|
| 4 |
+
m: []
|
| 5 |
+
python_version: 3.9.19
|
| 6 |
+
t:
|
| 7 |
+
"1":
|
| 8 |
+
- 1
|
| 9 |
+
- 5
|
| 10 |
+
- 11
|
| 11 |
+
- 49
|
| 12 |
+
- 51
|
| 13 |
+
- 53
|
| 14 |
+
- 55
|
| 15 |
+
- 71
|
| 16 |
+
- 98
|
| 17 |
+
"2":
|
| 18 |
+
- 1
|
| 19 |
+
- 5
|
| 20 |
+
- 11
|
| 21 |
+
- 49
|
| 22 |
+
- 51
|
| 23 |
+
- 53
|
| 24 |
+
- 55
|
| 25 |
+
- 71
|
| 26 |
+
- 98
|
| 27 |
+
"3":
|
| 28 |
+
- 2
|
| 29 |
+
- 13
|
| 30 |
+
- 23
|
| 31 |
+
- 55
|
| 32 |
+
"4": 3.9.19
|
| 33 |
+
"5": 0.18.5
|
| 34 |
+
"6": 4.45.1
|
| 35 |
+
"8":
|
| 36 |
+
- 5
|
| 37 |
+
"12": 0.18.5
|
| 38 |
+
"13": linux-x86_64
|
| 39 |
+
batch_size:
|
| 40 |
+
value: 3
|
| 41 |
+
epoch:
|
| 42 |
+
value: 6
|
| 43 |
+
lr:
|
| 44 |
+
value: 5e-06
|
| 45 |
+
perturbation:
|
| 46 |
+
value: reverse_full
|
| 47 |
+
seed:
|
| 48 |
+
value: 0
|
| 49 |
+
train_set:
|
| 50 |
+
value: 10M
|
wandb/run-20241031_122113-8ldget07/files/output.log
ADDED
|
@@ -0,0 +1,14 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Downloading shards: 100%|ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ| 2/2 [02:09<00:00, 64.61s/it]
|
| 2 |
+
Loading checkpoint shards: 100%|βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ| 2/2 [00:05<00:00, 2.61s/it]
|
| 3 |
+
tokenized_valid: Dataset({
|
| 4 |
+
features: ['input_ids', 'attention_mask'],
|
| 5 |
+
num_rows: 600
|
| 6 |
+
})
|
| 7 |
+
/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/training_args.py:1545: FutureWarning: `evaluation_strategy` is deprecated and will be removed in version 4.46 of π€ Transformers. Use `eval_strategy` instead
|
| 8 |
+
warnings.warn(
|
| 9 |
+
[2024-10-31 12:23:30,565] [INFO] [real_accelerator.py:219:get_accelerator] Setting ds_accelerator to cuda (auto detect)
|
| 10 |
+
[2024-10-31 12:23:39,013] [INFO] [comm.py:652:init_distributed] cdb=None
|
| 11 |
+
Installed CUDA version 11.8 does not match the version torch was compiled with 11.7 but since the APIs are compatible, accepting this combination
|
| 12 |
+
Using /home/chunhui/.cache/torch_extensions/py39_cu117 as PyTorch extensions root...
|
| 13 |
+
Loading extension module cpu_adam...
|
| 14 |
+
Time to load cpu_adam op: 5.0257627964019775 seconds
|
wandb/run-20241031_122113-8ldget07/files/wandb-metadata.json
ADDED
|
@@ -0,0 +1,97 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
|
| 3 |
+
"python": "3.9.19",
|
| 4 |
+
"startedAt": "2024-10-31T16:21:13.909771Z",
|
| 5 |
+
"args": [
|
| 6 |
+
"--perturbation",
|
| 7 |
+
"reverse_full",
|
| 8 |
+
"--train_set",
|
| 9 |
+
"10M",
|
| 10 |
+
"--batch_size",
|
| 11 |
+
"3",
|
| 12 |
+
"--epoch",
|
| 13 |
+
"6",
|
| 14 |
+
"--seed",
|
| 15 |
+
"0"
|
| 16 |
+
],
|
| 17 |
+
"program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
|
| 18 |
+
"codePath": "train/train_deep_wandb.py",
|
| 19 |
+
"git": {
|
| 20 |
+
"remote": "git@hf.co:Yaning1001/Impossible_llm.git",
|
| 21 |
+
"commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
|
| 22 |
+
},
|
| 23 |
+
"email": "yaning1001@gmail.com",
|
| 24 |
+
"root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
|
| 25 |
+
"host": "mms-large-2",
|
| 26 |
+
"username": "chunhui",
|
| 27 |
+
"executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
|
| 28 |
+
"codePathLocal": "train_deep_wandb.py",
|
| 29 |
+
"cpu_count": 32,
|
| 30 |
+
"cpu_count_logical": 64,
|
| 31 |
+
"gpu": "NVIDIA RTX A6000",
|
| 32 |
+
"gpu_count": 8,
|
| 33 |
+
"disk": {
|
| 34 |
+
"/": {
|
| 35 |
+
"total": "1888559353856",
|
| 36 |
+
"used": "1753159962624"
|
| 37 |
+
}
|
| 38 |
+
},
|
| 39 |
+
"memory": {
|
| 40 |
+
"total": "202617098240"
|
| 41 |
+
},
|
| 42 |
+
"cpu": {
|
| 43 |
+
"count": 32,
|
| 44 |
+
"countLogical": 64
|
| 45 |
+
},
|
| 46 |
+
"gpu_nvidia": [
|
| 47 |
+
{
|
| 48 |
+
"name": "NVIDIA RTX A6000",
|
| 49 |
+
"memoryTotal": "51527024640",
|
| 50 |
+
"cudaCores": 10752,
|
| 51 |
+
"architecture": "Ampere"
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"name": "NVIDIA RTX A6000",
|
| 55 |
+
"memoryTotal": "51527024640",
|
| 56 |
+
"cudaCores": 10752,
|
| 57 |
+
"architecture": "Ampere"
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"name": "NVIDIA RTX A6000",
|
| 61 |
+
"memoryTotal": "51527024640",
|
| 62 |
+
"cudaCores": 10752,
|
| 63 |
+
"architecture": "Ampere"
|
| 64 |
+
},
|
| 65 |
+
{
|
| 66 |
+
"name": "NVIDIA RTX A6000",
|
| 67 |
+
"memoryTotal": "51527024640",
|
| 68 |
+
"cudaCores": 10752,
|
| 69 |
+
"architecture": "Ampere"
|
| 70 |
+
},
|
| 71 |
+
{
|
| 72 |
+
"name": "NVIDIA RTX A6000",
|
| 73 |
+
"memoryTotal": "51527024640",
|
| 74 |
+
"cudaCores": 10752,
|
| 75 |
+
"architecture": "Ampere"
|
| 76 |
+
},
|
| 77 |
+
{
|
| 78 |
+
"name": "NVIDIA RTX A6000",
|
| 79 |
+
"memoryTotal": "51527024640",
|
| 80 |
+
"cudaCores": 10752,
|
| 81 |
+
"architecture": "Ampere"
|
| 82 |
+
},
|
| 83 |
+
{
|
| 84 |
+
"name": "NVIDIA RTX A6000",
|
| 85 |
+
"memoryTotal": "51527024640",
|
| 86 |
+
"cudaCores": 10752,
|
| 87 |
+
"architecture": "Ampere"
|
| 88 |
+
},
|
| 89 |
+
{
|
| 90 |
+
"name": "NVIDIA RTX A6000",
|
| 91 |
+
"memoryTotal": "51527024640",
|
| 92 |
+
"cudaCores": 10752,
|
| 93 |
+
"architecture": "Ampere"
|
| 94 |
+
}
|
| 95 |
+
],
|
| 96 |
+
"cudaVersion": "11.8"
|
| 97 |
+
}
|
wandb/run-20241031_122113-8ldget07/files/wandb-summary.json
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
{"_wandb":{"runtime":32015}}
|
wandb/run-20241031_122113-8ldget07/logs/debug-internal.log
ADDED
|
@@ -0,0 +1,18 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"time":"2024-10-31T12:21:13.912869182-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
|
| 2 |
+
{"time":"2024-10-31T12:21:13.912891512-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241031_122113-8ldget07/logs/debug-core.log"}
|
| 3 |
+
{"time":"2024-10-31T12:21:14.021742616-04:00","level":"INFO","msg":"created new stream","id":"8ldget07"}
|
| 4 |
+
{"time":"2024-10-31T12:21:14.021795536-04:00","level":"INFO","msg":"stream: started","id":"8ldget07"}
|
| 5 |
+
{"time":"2024-10-31T12:21:14.021865226-04:00","level":"INFO","msg":"sender: started","stream_id":"8ldget07"}
|
| 6 |
+
{"time":"2024-10-31T12:21:14.021867856-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"8ldget07"}}
|
| 7 |
+
{"time":"2024-10-31T12:21:14.021843416-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"8ldget07"}}
|
| 8 |
+
{"time":"2024-10-31T12:21:14.317036864-04:00","level":"INFO","msg":"Starting system monitor"}
|
| 9 |
+
{"time":"2024-10-31T17:51:23.39906822-04:00","level":"INFO","msg":"api: retrying error","error":"Post \"https://api.wandb.ai/graphql\": net/http: request canceled (Client.Timeout exceeded while awaiting headers)"}
|
| 10 |
+
{"time":"2024-10-31T21:14:49.439325218-04:00","level":"INFO","msg":"Stopping system monitor"}
|
| 11 |
+
{"time":"2024-10-31T21:14:49.51060772-04:00","level":"INFO","msg":"Stopped system monitor"}
|
| 12 |
+
{"time":"2024-10-31T21:14:50.239504017-04:00","level":"INFO","msg":"fileTransfer: Close: file transfer manager closed"}
|
| 13 |
+
{"time":"2024-10-31T21:14:50.358731013-04:00","level":"INFO","msg":"handler: operation stats","stats":{}}
|
| 14 |
+
{"time":"2024-10-31T21:14:51.422549145-04:00","level":"INFO","msg":"stream: closing","id":"8ldget07"}
|
| 15 |
+
{"time":"2024-10-31T21:14:51.422578535-04:00","level":"INFO","msg":"handler: closed","stream_id":{"value":"8ldget07"}}
|
| 16 |
+
{"time":"2024-10-31T21:14:51.422600615-04:00","level":"INFO","msg":"sender: closed","stream_id":"8ldget07"}
|
| 17 |
+
{"time":"2024-10-31T21:14:51.422595475-04:00","level":"INFO","msg":"writer: Close: closed","stream_id":{"value":"8ldget07"}}
|
| 18 |
+
{"time":"2024-10-31T21:14:51.422696086-04:00","level":"INFO","msg":"stream: closed","id":"8ldget07"}
|
wandb/run-20241031_122113-8ldget07/logs/debug.log
ADDED
|
@@ -0,0 +1,33 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
2024-10-31 12:21:13,907 INFO MainThread:558432 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
|
| 2 |
+
2024-10-31 12:21:13,907 INFO MainThread:558432 [wandb_setup.py:_flush():79] Configure stats pid to 558432
|
| 3 |
+
2024-10-31 12:21:13,907 INFO MainThread:558432 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
|
| 4 |
+
2024-10-31 12:21:13,907 INFO MainThread:558432 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
|
| 5 |
+
2024-10-31 12:21:13,907 INFO MainThread:558432 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
|
| 6 |
+
2024-10-31 12:21:13,907 INFO MainThread:558432 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
|
| 7 |
+
2024-10-31 12:21:13,907 INFO MainThread:558432 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
|
| 8 |
+
2024-10-31 12:21:13,907 INFO MainThread:558432 [wandb_setup.py:_flush():79] Applying login settings: {}
|
| 9 |
+
2024-10-31 12:21:13,907 INFO MainThread:558432 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241031_122113-8ldget07/logs/debug.log
|
| 10 |
+
2024-10-31 12:21:13,907 INFO MainThread:558432 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241031_122113-8ldget07/logs/debug-internal.log
|
| 11 |
+
2024-10-31 12:21:13,907 INFO MainThread:558432 [wandb_init.py:init():621] calling init triggers
|
| 12 |
+
2024-10-31 12:21:13,907 INFO MainThread:558432 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
|
| 13 |
+
config: {}
|
| 14 |
+
2024-10-31 12:21:13,907 INFO MainThread:558432 [wandb_init.py:init():671] starting backend
|
| 15 |
+
2024-10-31 12:21:13,907 INFO MainThread:558432 [wandb_init.py:init():675] sending inform_init request
|
| 16 |
+
2024-10-31 12:21:13,909 INFO MainThread:558432 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
|
| 17 |
+
2024-10-31 12:21:13,909 INFO MainThread:558432 [wandb_init.py:init():688] backend started and connected
|
| 18 |
+
2024-10-31 12:21:13,913 INFO MainThread:558432 [wandb_init.py:init():783] updated telemetry
|
| 19 |
+
2024-10-31 12:21:13,951 INFO MainThread:558432 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
|
| 20 |
+
2024-10-31 12:21:14,314 INFO MainThread:558432 [wandb_init.py:init():867] starting run threads in backend
|
| 21 |
+
2024-10-31 12:21:14,400 INFO MainThread:558432 [wandb_run.py:_console_start():2463] atexit reg
|
| 22 |
+
2024-10-31 12:21:14,400 INFO MainThread:558432 [wandb_run.py:_redirect():2311] redirect: wrap_raw
|
| 23 |
+
2024-10-31 12:21:14,400 INFO MainThread:558432 [wandb_run.py:_redirect():2376] Wrapping output streams.
|
| 24 |
+
2024-10-31 12:21:14,400 INFO MainThread:558432 [wandb_run.py:_redirect():2401] Redirects installed.
|
| 25 |
+
2024-10-31 12:21:14,402 INFO MainThread:558432 [wandb_init.py:init():911] run started, returning control to user process
|
| 26 |
+
2024-10-31 12:21:14,402 INFO MainThread:558432 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'reverse_full', 'train_set': '10M', 'batch_size': 3, 'epoch': 6, 'seed': 0, 'lr': 5e-06}
|
| 27 |
+
2024-10-31 21:14:49,288 INFO MainThread:558432 [wandb_run.py:_finish():2158] finishing run yaning1001-dartmouth-college/impossible_llm_reverse/8ldget07
|
| 28 |
+
2024-10-31 21:14:49,400 INFO MainThread:558432 [wandb_run.py:_atexit_cleanup():2426] got exitcode: 0
|
| 29 |
+
2024-10-31 21:14:49,400 INFO MainThread:558432 [wandb_run.py:_restore():2408] restore
|
| 30 |
+
2024-10-31 21:14:49,400 INFO MainThread:558432 [wandb_run.py:_restore():2414] restore done
|
| 31 |
+
2024-10-31 21:14:51,361 INFO MainThread:558432 [wandb_run.py:_footer_history_summary_info():3975] rendering history
|
| 32 |
+
2024-10-31 21:14:51,361 INFO MainThread:558432 [wandb_run.py:_footer_history_summary_info():4007] rendering summary
|
| 33 |
+
2024-10-31 21:14:51,421 INFO MainThread:558432 [wandb_run.py:_footer_sync_info():3934] logging synced files
|
wandb/run-20241101_094656-b81aanqd/files/output.log
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Loading checkpoint shards: 100%|ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ| 2/2 [00:05<00:00, 2.54s/it]
|
| 2 |
+
tokenized_valid: Dataset({
|
| 3 |
+
features: ['input_ids', 'attention_mask'],
|
| 4 |
+
num_rows: 600
|
| 5 |
+
})
|
| 6 |
+
/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/training_args.py:1545: FutureWarning: `evaluation_strategy` is deprecated and will be removed in version 4.46 of π€ Transformers. Use `eval_strategy` instead
|
| 7 |
+
warnings.warn(
|
| 8 |
+
[2024-11-01 09:47:03,559] [INFO] [real_accelerator.py:219:get_accelerator] Setting ds_accelerator to cuda (auto detect)
|
| 9 |
+
[2024-11-01 09:47:12,684] [INFO] [comm.py:652:init_distributed] cdb=None
|
| 10 |
+
Installed CUDA version 11.8 does not match the version torch was compiled with 11.7 but since the APIs are compatible, accepting this combination
|
| 11 |
+
Using /home/chunhui/.cache/torch_extensions/py39_cu117 as PyTorch extensions root...
|
| 12 |
+
Loading extension module cpu_adam...
|
| 13 |
+
Time to load cpu_adam op: 4.639373540878296 seconds
|