Add files using upload-large-folder tool
Browse filesThis view is limited to 50 files because it contains too many changes.
See raw diff
- wandb/run-20241030_010759-dim9v1es/files/config.yaml +47 -0
- wandb/run-20241030_010759-dim9v1es/files/output.log +4 -0
- wandb/run-20241030_010759-dim9v1es/run-dim9v1es.wandb +0 -0
- wandb/run-20241030_011509-46hc4g2h/logs/debug-internal.log +8 -0
- wandb/run-20241030_011509-46hc4g2h/logs/debug.log +26 -0
- wandb/run-20241030_012617-0h15y3p4/logs/debug-internal.log +11 -0
- wandb/run-20241030_112700-j5l8vh9z/files/config.yaml +47 -0
- wandb/run-20241030_112700-j5l8vh9z/files/requirements.txt +147 -0
- wandb/run-20241030_112700-j5l8vh9z/files/wandb-summary.json +1 -0
- wandb/run-20241030_225833-frh96rd1/files/wandb-metadata.json +97 -0
- wandb/run-20241030_225833-frh96rd1/logs/debug-internal.log +8 -0
- wandb/run-20241030_225833-frh96rd1/logs/debug.log +26 -0
- wandb/run-20241031_000839-cu7972v5/files/output.log +13 -0
- wandb/run-20241031_000839-cu7972v5/files/requirements.txt +147 -0
- wandb/run-20241031_000839-cu7972v5/files/wandb-metadata.json +97 -0
- wandb/run-20241031_000839-cu7972v5/logs/debug-internal.log +8 -0
- wandb/run-20241031_000839-cu7972v5/logs/debug.log +26 -0
- wandb/run-20241031_000839-cu7972v5/run-cu7972v5.wandb +0 -0
- wandb/run-20241031_001055-32u9qnul/logs/debug-internal.log +8 -0
- wandb/run-20241031_001055-32u9qnul/logs/debug.log +26 -0
- wandb/run-20241031_002020-u516mysu/files/output.log +0 -0
- wandb/run-20241031_114700-jx2hqvx3/files/wandb-metadata.json +97 -0
- wandb/run-20241101_012438-61w48leq/files/config.yaml +49 -0
- wandb/run-20241101_012438-61w48leq/files/output.log +12 -0
- wandb/run-20241101_012438-61w48leq/files/wandb-metadata.json +97 -0
- wandb/run-20241101_012438-61w48leq/files/wandb-summary.json +1 -0
- wandb/run-20241101_012438-61w48leq/run-61w48leq.wandb +0 -0
- wandb/run-20241101_012734-m18lsdzn/files/wandb-metadata.json +97 -0
- wandb/run-20241101_012734-m18lsdzn/logs/debug-internal.log +8 -0
- wandb/run-20241101_092804-qhsuxbxe/run-qhsuxbxe.wandb +0 -0
- wandb/run-20241101_200517-77b12390/run-77b12390.wandb +0 -0
- wandb/run-20241101_200517-iopieyi0/files/config.yaml +49 -0
- wandb/run-20241101_200517-iopieyi0/files/output.log +42 -0
- wandb/run-20241101_200517-iopieyi0/files/requirements.txt +147 -0
- wandb/run-20241101_200517-iopieyi0/files/wandb-metadata.json +97 -0
- wandb/run-20241101_200517-iopieyi0/files/wandb-summary.json +1 -0
- wandb/run-20241101_200517-iopieyi0/logs/debug-internal.log +11 -0
- wandb/run-20241101_201927-8tmqrwpx/logs/debug-internal.log +16 -0
- wandb/run-20241105_160059-czoj7ear/files/wandb-summary.json +1 -0
- wandb/run-20241105_160059-czoj7ear/logs/debug-internal.log +17 -0
- wandb/run-20241105_160059-czoj7ear/run-czoj7ear.wandb +0 -0
- wandb/run-20241105_161832-c18fx9uc/files/requirements.txt +147 -0
- wandb/run-20241105_162858-6py0unak/files/config.yaml +49 -0
- wandb/run-20241105_162858-6py0unak/files/output.log +34 -0
- wandb/run-20241105_162858-6py0unak/files/requirements.txt +147 -0
- wandb/run-20241105_162858-6py0unak/files/wandb-metadata.json +97 -0
- wandb/run-20241105_162858-6py0unak/files/wandb-summary.json +1 -0
- wandb/run-20241105_162858-6py0unak/logs/debug-internal.log +12 -0
- wandb/run-20241105_162858-6py0unak/logs/debug.log +27 -0
- wandb/run-20241105_163001-vvohahtj/run-vvohahtj.wandb +0 -0
wandb/run-20241030_010759-dim9v1es/files/config.yaml
ADDED
|
@@ -0,0 +1,47 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
_wandb:
|
| 2 |
+
value:
|
| 3 |
+
cli_version: 0.18.5
|
| 4 |
+
m: []
|
| 5 |
+
python_version: 3.9.19
|
| 6 |
+
t:
|
| 7 |
+
"1":
|
| 8 |
+
- 1
|
| 9 |
+
- 5
|
| 10 |
+
- 11
|
| 11 |
+
- 49
|
| 12 |
+
- 51
|
| 13 |
+
- 53
|
| 14 |
+
- 55
|
| 15 |
+
- 71
|
| 16 |
+
- 98
|
| 17 |
+
"2":
|
| 18 |
+
- 1
|
| 19 |
+
- 5
|
| 20 |
+
- 11
|
| 21 |
+
- 49
|
| 22 |
+
- 51
|
| 23 |
+
- 53
|
| 24 |
+
- 55
|
| 25 |
+
- 71
|
| 26 |
+
- 98
|
| 27 |
+
"3":
|
| 28 |
+
- 13
|
| 29 |
+
- 23
|
| 30 |
+
- 55
|
| 31 |
+
"4": 3.9.19
|
| 32 |
+
"5": 0.18.5
|
| 33 |
+
"6": 4.45.1
|
| 34 |
+
"8":
|
| 35 |
+
- 5
|
| 36 |
+
"12": 0.18.5
|
| 37 |
+
"13": linux-x86_64
|
| 38 |
+
batch_size:
|
| 39 |
+
value: 3
|
| 40 |
+
epoch:
|
| 41 |
+
value: 7
|
| 42 |
+
perturbation:
|
| 43 |
+
value: reverse_control
|
| 44 |
+
seed:
|
| 45 |
+
value: 0
|
| 46 |
+
train_set:
|
| 47 |
+
value: 10M
|
wandb/run-20241030_010759-dim9v1es/files/output.log
ADDED
|
@@ -0,0 +1,4 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Traceback (most recent call last):
|
| 2 |
+
File "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py", line 162, in <module>
|
| 3 |
+
dataset_name = f"babylm_{args.perturbation}_{args.train_zset}_seed{args.seed}"
|
| 4 |
+
AttributeError: 'Namespace' object has no attribute 'train_zset'
|
wandb/run-20241030_010759-dim9v1es/run-dim9v1es.wandb
ADDED
|
Binary file (1.6 kB). View file
|
|
|
wandb/run-20241030_011509-46hc4g2h/logs/debug-internal.log
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"time":"2024-10-30T01:15:09.556689441-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
|
| 2 |
+
{"time":"2024-10-30T01:15:09.556707271-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_011509-46hc4g2h/logs/debug-core.log"}
|
| 3 |
+
{"time":"2024-10-30T01:15:09.662171735-04:00","level":"INFO","msg":"created new stream","id":"46hc4g2h"}
|
| 4 |
+
{"time":"2024-10-30T01:15:09.662202115-04:00","level":"INFO","msg":"stream: started","id":"46hc4g2h"}
|
| 5 |
+
{"time":"2024-10-30T01:15:09.662224815-04:00","level":"INFO","msg":"sender: started","stream_id":"46hc4g2h"}
|
| 6 |
+
{"time":"2024-10-30T01:15:09.662216935-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"46hc4g2h"}}
|
| 7 |
+
{"time":"2024-10-30T01:15:09.662203965-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"46hc4g2h"}}
|
| 8 |
+
{"time":"2024-10-30T01:15:09.829266444-04:00","level":"INFO","msg":"Starting system monitor"}
|
wandb/run-20241030_011509-46hc4g2h/logs/debug.log
ADDED
|
@@ -0,0 +1,26 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
2024-10-30 01:15:09,552 INFO MainThread:324928 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
|
| 2 |
+
2024-10-30 01:15:09,552 INFO MainThread:324928 [wandb_setup.py:_flush():79] Configure stats pid to 324928
|
| 3 |
+
2024-10-30 01:15:09,552 INFO MainThread:324928 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
|
| 4 |
+
2024-10-30 01:15:09,552 INFO MainThread:324928 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
|
| 5 |
+
2024-10-30 01:15:09,552 INFO MainThread:324928 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
|
| 6 |
+
2024-10-30 01:15:09,552 INFO MainThread:324928 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
|
| 7 |
+
2024-10-30 01:15:09,552 INFO MainThread:324928 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
|
| 8 |
+
2024-10-30 01:15:09,552 INFO MainThread:324928 [wandb_setup.py:_flush():79] Applying login settings: {}
|
| 9 |
+
2024-10-30 01:15:09,552 INFO MainThread:324928 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_011509-46hc4g2h/logs/debug.log
|
| 10 |
+
2024-10-30 01:15:09,552 INFO MainThread:324928 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_011509-46hc4g2h/logs/debug-internal.log
|
| 11 |
+
2024-10-30 01:15:09,552 INFO MainThread:324928 [wandb_init.py:init():621] calling init triggers
|
| 12 |
+
2024-10-30 01:15:09,552 INFO MainThread:324928 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
|
| 13 |
+
config: {}
|
| 14 |
+
2024-10-30 01:15:09,552 INFO MainThread:324928 [wandb_init.py:init():671] starting backend
|
| 15 |
+
2024-10-30 01:15:09,552 INFO MainThread:324928 [wandb_init.py:init():675] sending inform_init request
|
| 16 |
+
2024-10-30 01:15:09,554 INFO MainThread:324928 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
|
| 17 |
+
2024-10-30 01:15:09,554 INFO MainThread:324928 [wandb_init.py:init():688] backend started and connected
|
| 18 |
+
2024-10-30 01:15:09,558 INFO MainThread:324928 [wandb_init.py:init():783] updated telemetry
|
| 19 |
+
2024-10-30 01:15:09,594 INFO MainThread:324928 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
|
| 20 |
+
2024-10-30 01:15:09,825 INFO MainThread:324928 [wandb_init.py:init():867] starting run threads in backend
|
| 21 |
+
2024-10-30 01:15:09,928 INFO MainThread:324928 [wandb_run.py:_console_start():2463] atexit reg
|
| 22 |
+
2024-10-30 01:15:09,928 INFO MainThread:324928 [wandb_run.py:_redirect():2311] redirect: wrap_raw
|
| 23 |
+
2024-10-30 01:15:09,929 INFO MainThread:324928 [wandb_run.py:_redirect():2376] Wrapping output streams.
|
| 24 |
+
2024-10-30 01:15:09,929 INFO MainThread:324928 [wandb_run.py:_redirect():2401] Redirects installed.
|
| 25 |
+
2024-10-30 01:15:09,931 INFO MainThread:324928 [wandb_init.py:init():911] run started, returning control to user process
|
| 26 |
+
2024-10-30 01:15:09,931 INFO MainThread:324928 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'reverse_control', 'train_set': '10M', 'batch_size': 3, 'epoch': 7, 'seed': 0}
|
wandb/run-20241030_012617-0h15y3p4/logs/debug-internal.log
ADDED
|
@@ -0,0 +1,11 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"time":"2024-10-30T01:26:17.390480352-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
|
| 2 |
+
{"time":"2024-10-30T01:26:17.390492292-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_012617-0h15y3p4/logs/debug-core.log"}
|
| 3 |
+
{"time":"2024-10-30T01:26:17.497198871-04:00","level":"INFO","msg":"created new stream","id":"0h15y3p4"}
|
| 4 |
+
{"time":"2024-10-30T01:26:17.497231081-04:00","level":"INFO","msg":"stream: started","id":"0h15y3p4"}
|
| 5 |
+
{"time":"2024-10-30T01:26:17.498386319-04:00","level":"INFO","msg":"sender: started","stream_id":"0h15y3p4"}
|
| 6 |
+
{"time":"2024-10-30T01:26:17.498411899-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"0h15y3p4"}}
|
| 7 |
+
{"time":"2024-10-30T01:26:17.498470659-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"0h15y3p4"}}
|
| 8 |
+
{"time":"2024-10-30T01:26:17.697892892-04:00","level":"INFO","msg":"Starting system monitor"}
|
| 9 |
+
{"time":"2024-10-30T01:26:32.054557486-04:00","level":"INFO","msg":"stream: closing","id":"0h15y3p4"}
|
| 10 |
+
{"time":"2024-10-30T01:26:32.054592926-04:00","level":"INFO","msg":"Stopping system monitor"}
|
| 11 |
+
{"time":"2024-10-30T01:26:32.055079549-04:00","level":"INFO","msg":"Stopped system monitor"}
|
wandb/run-20241030_112700-j5l8vh9z/files/config.yaml
ADDED
|
@@ -0,0 +1,47 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
_wandb:
|
| 2 |
+
value:
|
| 3 |
+
cli_version: 0.18.5
|
| 4 |
+
m: []
|
| 5 |
+
python_version: 3.9.19
|
| 6 |
+
t:
|
| 7 |
+
"1":
|
| 8 |
+
- 1
|
| 9 |
+
- 5
|
| 10 |
+
- 11
|
| 11 |
+
- 49
|
| 12 |
+
- 51
|
| 13 |
+
- 53
|
| 14 |
+
- 55
|
| 15 |
+
- 71
|
| 16 |
+
- 98
|
| 17 |
+
"2":
|
| 18 |
+
- 1
|
| 19 |
+
- 5
|
| 20 |
+
- 11
|
| 21 |
+
- 49
|
| 22 |
+
- 51
|
| 23 |
+
- 53
|
| 24 |
+
- 55
|
| 25 |
+
- 71
|
| 26 |
+
- 98
|
| 27 |
+
"3":
|
| 28 |
+
- 13
|
| 29 |
+
- 23
|
| 30 |
+
- 55
|
| 31 |
+
"4": 3.9.19
|
| 32 |
+
"5": 0.18.5
|
| 33 |
+
"6": 4.45.1
|
| 34 |
+
"8":
|
| 35 |
+
- 5
|
| 36 |
+
"12": 0.18.5
|
| 37 |
+
"13": linux-x86_64
|
| 38 |
+
batch_size:
|
| 39 |
+
value: 3
|
| 40 |
+
epoch:
|
| 41 |
+
value: 3
|
| 42 |
+
perturbation:
|
| 43 |
+
value: reverse_control
|
| 44 |
+
seed:
|
| 45 |
+
value: 0
|
| 46 |
+
train_set:
|
| 47 |
+
value: 10M
|
wandb/run-20241030_112700-j5l8vh9z/files/requirements.txt
ADDED
|
@@ -0,0 +1,147 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
funcsigs==1.0.2
|
| 2 |
+
sentry-sdk==2.17.0
|
| 3 |
+
multiprocess==0.70.16
|
| 4 |
+
numpy==1.26.2
|
| 5 |
+
pluralizer==1.2.0
|
| 6 |
+
debugpy==1.6.7
|
| 7 |
+
nvidia-cudnn-cu11==8.5.0.96
|
| 8 |
+
deepspeed==0.15.2
|
| 9 |
+
data==0.4
|
| 10 |
+
pandas==2.1.3
|
| 11 |
+
tomli==2.0.1
|
| 12 |
+
charset-normalizer==3.3.2
|
| 13 |
+
attrs==24.2.0
|
| 14 |
+
aiosignal==1.3.1
|
| 15 |
+
fsspec==2023.10.0
|
| 16 |
+
nvidia-cusparse-cu11==11.7.4.91
|
| 17 |
+
zipp==3.12.0
|
| 18 |
+
mypy-extensions==1.0.0
|
| 19 |
+
datasets==3.0.1
|
| 20 |
+
joblib==1.3.2
|
| 21 |
+
hjson==3.1.0
|
| 22 |
+
traitlets==5.7.1
|
| 23 |
+
stack-data==0.6.0
|
| 24 |
+
transformers==4.45.1
|
| 25 |
+
sympy==1.11.1
|
| 26 |
+
Pygments==2.15.0
|
| 27 |
+
docker-pycreds==0.4.0
|
| 28 |
+
dill==0.3.8
|
| 29 |
+
wheel==0.44.0
|
| 30 |
+
prompt-toolkit==3.0.30
|
| 31 |
+
parso==0.8.3
|
| 32 |
+
ipykernel==6.23.1
|
| 33 |
+
pyarrow==17.0.0
|
| 34 |
+
certifi==2023.11.17
|
| 35 |
+
nvidia-cufft-cu11==10.9.0.58
|
| 36 |
+
six==1.16.0
|
| 37 |
+
pydantic==2.9.2
|
| 38 |
+
click==8.1.7
|
| 39 |
+
nest-asyncio==1.5.6
|
| 40 |
+
gmpy2==2.1.0
|
| 41 |
+
matplotlib==3.8.2
|
| 42 |
+
scipy==1.11.4
|
| 43 |
+
typing_extensions==4.12.2
|
| 44 |
+
statsmodels==0.14.0
|
| 45 |
+
huggingface-hub==0.25.0
|
| 46 |
+
frozenlist==1.4.1
|
| 47 |
+
gpustat==1.1.1
|
| 48 |
+
nvidia-nvtx-cu11==11.7.91
|
| 49 |
+
safetensors==0.4.5
|
| 50 |
+
stanza==1.9.2
|
| 51 |
+
decorator==5.1.1
|
| 52 |
+
seaborn==0.13.0
|
| 53 |
+
sentencepiece==0.2.0
|
| 54 |
+
PyYAML==6.0.1
|
| 55 |
+
black==24.8.0
|
| 56 |
+
protobuf==4.25.1
|
| 57 |
+
pickleshare==0.7.5
|
| 58 |
+
peft==0.13.0
|
| 59 |
+
triton==2.0.0
|
| 60 |
+
nvidia-cuda-runtime-cu11==11.7.99
|
| 61 |
+
Jinja2==3.1.2
|
| 62 |
+
nvidia-cusolver-cu11==11.4.0.1
|
| 63 |
+
executing==1.2.0
|
| 64 |
+
jupyter_client==8.1.0
|
| 65 |
+
pluggy==1.3.0
|
| 66 |
+
cmake==3.30.3
|
| 67 |
+
pytz==2023.3.post1
|
| 68 |
+
aiohappyeyeballs==2.4.2
|
| 69 |
+
kiwisolver==1.4.5
|
| 70 |
+
py-cpuinfo==9.0.0
|
| 71 |
+
Pillow==10.1.0
|
| 72 |
+
ptyprocess==0.7.0
|
| 73 |
+
importlib_resources==6.4.5
|
| 74 |
+
GitPython==3.1.43
|
| 75 |
+
importlib-metadata==6.0.0
|
| 76 |
+
iniconfig==2.0.0
|
| 77 |
+
scikit-learn==1.3.2
|
| 78 |
+
exceptiongroup==1.1.0
|
| 79 |
+
networkx==2.8.6
|
| 80 |
+
accelerate==1.0.0
|
| 81 |
+
nltk==3.8.1
|
| 82 |
+
shutilwhich==1.1.0
|
| 83 |
+
fonttools==4.45.1
|
| 84 |
+
future==0.18.3
|
| 85 |
+
aiohttp==3.10.6
|
| 86 |
+
wcwidth==0.2.5
|
| 87 |
+
idna==3.6
|
| 88 |
+
filelock==3.12.2
|
| 89 |
+
pathspec==0.12.1
|
| 90 |
+
jupyter_core==5.1.0
|
| 91 |
+
lit==18.1.8
|
| 92 |
+
nvidia-curand-cu11==10.2.10.91
|
| 93 |
+
nvidia-cublas-cu11==11.10.3.66
|
| 94 |
+
nvidia-ml-py==12.560.30
|
| 95 |
+
msgpack==1.1.0
|
| 96 |
+
python-dateutil==2.8.2
|
| 97 |
+
blessed==1.20.0
|
| 98 |
+
packaging==23.0
|
| 99 |
+
gitdb==4.0.11
|
| 100 |
+
yarl==1.13.0
|
| 101 |
+
emoji==2.8.0
|
| 102 |
+
tzdata==2023.3
|
| 103 |
+
cycler==0.12.1
|
| 104 |
+
tornado==6.2
|
| 105 |
+
backcall==0.2.0
|
| 106 |
+
plotnine==0.12.4
|
| 107 |
+
ninja==1.11.1.1
|
| 108 |
+
latex==0.7.0
|
| 109 |
+
wandb==0.18.5
|
| 110 |
+
setproctitle==1.3.3
|
| 111 |
+
threadpoolctl==3.2.0
|
| 112 |
+
requests==2.32.3
|
| 113 |
+
pyparsing==3.1.1
|
| 114 |
+
smmap==5.0.1
|
| 115 |
+
pyzmq==23.0.0
|
| 116 |
+
async-timeout==4.0.3
|
| 117 |
+
annotated-types==0.7.0
|
| 118 |
+
matplotlib-inline==0.1.6
|
| 119 |
+
latexcodec==1.0.0
|
| 120 |
+
ipython==8.0.0
|
| 121 |
+
patsy==0.5.3
|
| 122 |
+
contourpy==1.2.0
|
| 123 |
+
multidict==6.1.0
|
| 124 |
+
mizani==0.9.3
|
| 125 |
+
urllib3==2.1.0
|
| 126 |
+
tokenizers==0.20.0
|
| 127 |
+
MarkupSafe==2.1.2
|
| 128 |
+
pip==24.2
|
| 129 |
+
pexpect==4.8.0
|
| 130 |
+
tqdm==4.66.5
|
| 131 |
+
jedi==0.18.2
|
| 132 |
+
pydantic_core==2.23.4
|
| 133 |
+
tempdir==0.7.1
|
| 134 |
+
mpmath==1.2.1
|
| 135 |
+
setuptools==72.1.0
|
| 136 |
+
pytest==7.4.3
|
| 137 |
+
pure-eval==0.2.2
|
| 138 |
+
psutil==5.9.1
|
| 139 |
+
comm==0.1.2
|
| 140 |
+
nvidia-cuda-cupti-cu11==11.7.101
|
| 141 |
+
nvidia-cuda-nvrtc-cu11==11.7.99
|
| 142 |
+
regex==2023.10.3
|
| 143 |
+
platformdirs==2.5.2
|
| 144 |
+
asttokens==2.2.1
|
| 145 |
+
torch==2.0.0
|
| 146 |
+
nvidia-nccl-cu11==2.14.3
|
| 147 |
+
xxhash==3.5.0
|
wandb/run-20241030_112700-j5l8vh9z/files/wandb-summary.json
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
{"_wandb":{"runtime":93}}
|
wandb/run-20241030_225833-frh96rd1/files/wandb-metadata.json
ADDED
|
@@ -0,0 +1,97 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
|
| 3 |
+
"python": "3.9.19",
|
| 4 |
+
"startedAt": "2024-10-31T02:58:33.401365Z",
|
| 5 |
+
"args": [
|
| 6 |
+
"--perturbation",
|
| 7 |
+
"reverse_full",
|
| 8 |
+
"--train_set",
|
| 9 |
+
"10M",
|
| 10 |
+
"--batch_size",
|
| 11 |
+
"3",
|
| 12 |
+
"--epoch",
|
| 13 |
+
"3",
|
| 14 |
+
"--seed",
|
| 15 |
+
"0"
|
| 16 |
+
],
|
| 17 |
+
"program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
|
| 18 |
+
"codePath": "train/train_deep_wandb.py",
|
| 19 |
+
"git": {
|
| 20 |
+
"remote": "git@hf.co:Yaning1001/Impossible_llm.git",
|
| 21 |
+
"commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
|
| 22 |
+
},
|
| 23 |
+
"email": "yaning1001@gmail.com",
|
| 24 |
+
"root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
|
| 25 |
+
"host": "mms-large-2",
|
| 26 |
+
"username": "chunhui",
|
| 27 |
+
"executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
|
| 28 |
+
"codePathLocal": "train_deep_wandb.py",
|
| 29 |
+
"cpu_count": 32,
|
| 30 |
+
"cpu_count_logical": 64,
|
| 31 |
+
"gpu": "NVIDIA RTX A6000",
|
| 32 |
+
"gpu_count": 8,
|
| 33 |
+
"disk": {
|
| 34 |
+
"/": {
|
| 35 |
+
"total": "1888559353856",
|
| 36 |
+
"used": "1710970519552"
|
| 37 |
+
}
|
| 38 |
+
},
|
| 39 |
+
"memory": {
|
| 40 |
+
"total": "202617098240"
|
| 41 |
+
},
|
| 42 |
+
"cpu": {
|
| 43 |
+
"count": 32,
|
| 44 |
+
"countLogical": 64
|
| 45 |
+
},
|
| 46 |
+
"gpu_nvidia": [
|
| 47 |
+
{
|
| 48 |
+
"name": "NVIDIA RTX A6000",
|
| 49 |
+
"memoryTotal": "51527024640",
|
| 50 |
+
"cudaCores": 10752,
|
| 51 |
+
"architecture": "Ampere"
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"name": "NVIDIA RTX A6000",
|
| 55 |
+
"memoryTotal": "51527024640",
|
| 56 |
+
"cudaCores": 10752,
|
| 57 |
+
"architecture": "Ampere"
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"name": "NVIDIA RTX A6000",
|
| 61 |
+
"memoryTotal": "51527024640",
|
| 62 |
+
"cudaCores": 10752,
|
| 63 |
+
"architecture": "Ampere"
|
| 64 |
+
},
|
| 65 |
+
{
|
| 66 |
+
"name": "NVIDIA RTX A6000",
|
| 67 |
+
"memoryTotal": "51527024640",
|
| 68 |
+
"cudaCores": 10752,
|
| 69 |
+
"architecture": "Ampere"
|
| 70 |
+
},
|
| 71 |
+
{
|
| 72 |
+
"name": "NVIDIA RTX A6000",
|
| 73 |
+
"memoryTotal": "51527024640",
|
| 74 |
+
"cudaCores": 10752,
|
| 75 |
+
"architecture": "Ampere"
|
| 76 |
+
},
|
| 77 |
+
{
|
| 78 |
+
"name": "NVIDIA RTX A6000",
|
| 79 |
+
"memoryTotal": "51527024640",
|
| 80 |
+
"cudaCores": 10752,
|
| 81 |
+
"architecture": "Ampere"
|
| 82 |
+
},
|
| 83 |
+
{
|
| 84 |
+
"name": "NVIDIA RTX A6000",
|
| 85 |
+
"memoryTotal": "51527024640",
|
| 86 |
+
"cudaCores": 10752,
|
| 87 |
+
"architecture": "Ampere"
|
| 88 |
+
},
|
| 89 |
+
{
|
| 90 |
+
"name": "NVIDIA RTX A6000",
|
| 91 |
+
"memoryTotal": "51527024640",
|
| 92 |
+
"cudaCores": 10752,
|
| 93 |
+
"architecture": "Ampere"
|
| 94 |
+
}
|
| 95 |
+
],
|
| 96 |
+
"cudaVersion": "11.8"
|
| 97 |
+
}
|
wandb/run-20241030_225833-frh96rd1/logs/debug-internal.log
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"time":"2024-10-30T22:58:33.403200163-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
|
| 2 |
+
{"time":"2024-10-30T22:58:33.403213043-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_225833-frh96rd1/logs/debug-core.log"}
|
| 3 |
+
{"time":"2024-10-30T22:58:33.510780385-04:00","level":"INFO","msg":"created new stream","id":"frh96rd1"}
|
| 4 |
+
{"time":"2024-10-30T22:58:33.510819765-04:00","level":"INFO","msg":"stream: started","id":"frh96rd1"}
|
| 5 |
+
{"time":"2024-10-30T22:58:33.510841375-04:00","level":"INFO","msg":"sender: started","stream_id":"frh96rd1"}
|
| 6 |
+
{"time":"2024-10-30T22:58:33.510849875-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"frh96rd1"}}
|
| 7 |
+
{"time":"2024-10-30T22:58:33.510823865-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"frh96rd1"}}
|
| 8 |
+
{"time":"2024-10-30T22:58:33.680830326-04:00","level":"INFO","msg":"Starting system monitor"}
|
wandb/run-20241030_225833-frh96rd1/logs/debug.log
ADDED
|
@@ -0,0 +1,26 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
2024-10-30 22:58:33,399 INFO MainThread:451913 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
|
| 2 |
+
2024-10-30 22:58:33,399 INFO MainThread:451913 [wandb_setup.py:_flush():79] Configure stats pid to 451913
|
| 3 |
+
2024-10-30 22:58:33,399 INFO MainThread:451913 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
|
| 4 |
+
2024-10-30 22:58:33,399 INFO MainThread:451913 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
|
| 5 |
+
2024-10-30 22:58:33,399 INFO MainThread:451913 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
|
| 6 |
+
2024-10-30 22:58:33,399 INFO MainThread:451913 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
|
| 7 |
+
2024-10-30 22:58:33,399 INFO MainThread:451913 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
|
| 8 |
+
2024-10-30 22:58:33,399 INFO MainThread:451913 [wandb_setup.py:_flush():79] Applying login settings: {}
|
| 9 |
+
2024-10-30 22:58:33,399 INFO MainThread:451913 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_225833-frh96rd1/logs/debug.log
|
| 10 |
+
2024-10-30 22:58:33,399 INFO MainThread:451913 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_225833-frh96rd1/logs/debug-internal.log
|
| 11 |
+
2024-10-30 22:58:33,399 INFO MainThread:451913 [wandb_init.py:init():621] calling init triggers
|
| 12 |
+
2024-10-30 22:58:33,399 INFO MainThread:451913 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
|
| 13 |
+
config: {}
|
| 14 |
+
2024-10-30 22:58:33,399 INFO MainThread:451913 [wandb_init.py:init():671] starting backend
|
| 15 |
+
2024-10-30 22:58:33,400 INFO MainThread:451913 [wandb_init.py:init():675] sending inform_init request
|
| 16 |
+
2024-10-30 22:58:33,400 INFO MainThread:451913 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
|
| 17 |
+
2024-10-30 22:58:33,401 INFO MainThread:451913 [wandb_init.py:init():688] backend started and connected
|
| 18 |
+
2024-10-30 22:58:33,403 INFO MainThread:451913 [wandb_init.py:init():783] updated telemetry
|
| 19 |
+
2024-10-30 22:58:33,443 INFO MainThread:451913 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
|
| 20 |
+
2024-10-30 22:58:33,677 INFO MainThread:451913 [wandb_init.py:init():867] starting run threads in backend
|
| 21 |
+
2024-10-30 22:58:33,841 INFO MainThread:451913 [wandb_run.py:_console_start():2463] atexit reg
|
| 22 |
+
2024-10-30 22:58:33,841 INFO MainThread:451913 [wandb_run.py:_redirect():2311] redirect: wrap_raw
|
| 23 |
+
2024-10-30 22:58:33,841 INFO MainThread:451913 [wandb_run.py:_redirect():2376] Wrapping output streams.
|
| 24 |
+
2024-10-30 22:58:33,842 INFO MainThread:451913 [wandb_run.py:_redirect():2401] Redirects installed.
|
| 25 |
+
2024-10-30 22:58:33,844 INFO MainThread:451913 [wandb_init.py:init():911] run started, returning control to user process
|
| 26 |
+
2024-10-30 22:58:33,844 INFO MainThread:451913 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'reverse_full', 'train_set': '10M', 'batch_size': 3, 'epoch': 3, 'seed': 0}
|
wandb/run-20241031_000839-cu7972v5/files/output.log
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Loading checkpoint shards: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 2/2 [00:18<00:00, 9.23s/it]
|
| 2 |
+
tokenized_valid: Dataset({
|
| 3 |
+
features: ['input_ids', 'attention_mask'],
|
| 4 |
+
num_rows: 600
|
| 5 |
+
})
|
| 6 |
+
/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/training_args.py:1545: FutureWarning: `evaluation_strategy` is deprecated and will be removed in version 4.46 of 🤗 Transformers. Use `eval_strategy` instead
|
| 7 |
+
warnings.warn(
|
| 8 |
+
[2024-10-31 00:09:00,035] [INFO] [real_accelerator.py:219:get_accelerator] Setting ds_accelerator to cuda (auto detect)
|
| 9 |
+
[2024-10-31 00:09:09,768] [INFO] [comm.py:652:init_distributed] cdb=None
|
| 10 |
+
Installed CUDA version 11.8 does not match the version torch was compiled with 11.7 but since the APIs are compatible, accepting this combination
|
| 11 |
+
Using /home/chunhui/.cache/torch_extensions/py39_cu117 as PyTorch extensions root...
|
| 12 |
+
Loading extension module cpu_adam...
|
| 13 |
+
Time to load cpu_adam op: 5.32434344291687 seconds
|
wandb/run-20241031_000839-cu7972v5/files/requirements.txt
ADDED
|
@@ -0,0 +1,147 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
funcsigs==1.0.2
|
| 2 |
+
sentry-sdk==2.17.0
|
| 3 |
+
multiprocess==0.70.16
|
| 4 |
+
numpy==1.26.2
|
| 5 |
+
pluralizer==1.2.0
|
| 6 |
+
debugpy==1.6.7
|
| 7 |
+
nvidia-cudnn-cu11==8.5.0.96
|
| 8 |
+
deepspeed==0.15.2
|
| 9 |
+
data==0.4
|
| 10 |
+
pandas==2.1.3
|
| 11 |
+
tomli==2.0.1
|
| 12 |
+
charset-normalizer==3.3.2
|
| 13 |
+
attrs==24.2.0
|
| 14 |
+
aiosignal==1.3.1
|
| 15 |
+
fsspec==2023.10.0
|
| 16 |
+
nvidia-cusparse-cu11==11.7.4.91
|
| 17 |
+
zipp==3.12.0
|
| 18 |
+
mypy-extensions==1.0.0
|
| 19 |
+
datasets==3.0.1
|
| 20 |
+
joblib==1.3.2
|
| 21 |
+
hjson==3.1.0
|
| 22 |
+
traitlets==5.7.1
|
| 23 |
+
stack-data==0.6.0
|
| 24 |
+
transformers==4.45.1
|
| 25 |
+
sympy==1.11.1
|
| 26 |
+
Pygments==2.15.0
|
| 27 |
+
docker-pycreds==0.4.0
|
| 28 |
+
dill==0.3.8
|
| 29 |
+
wheel==0.44.0
|
| 30 |
+
prompt-toolkit==3.0.30
|
| 31 |
+
parso==0.8.3
|
| 32 |
+
ipykernel==6.23.1
|
| 33 |
+
pyarrow==17.0.0
|
| 34 |
+
certifi==2023.11.17
|
| 35 |
+
nvidia-cufft-cu11==10.9.0.58
|
| 36 |
+
six==1.16.0
|
| 37 |
+
pydantic==2.9.2
|
| 38 |
+
click==8.1.7
|
| 39 |
+
nest-asyncio==1.5.6
|
| 40 |
+
gmpy2==2.1.0
|
| 41 |
+
matplotlib==3.8.2
|
| 42 |
+
scipy==1.11.4
|
| 43 |
+
typing_extensions==4.12.2
|
| 44 |
+
statsmodels==0.14.0
|
| 45 |
+
huggingface-hub==0.25.0
|
| 46 |
+
frozenlist==1.4.1
|
| 47 |
+
gpustat==1.1.1
|
| 48 |
+
nvidia-nvtx-cu11==11.7.91
|
| 49 |
+
safetensors==0.4.5
|
| 50 |
+
stanza==1.9.2
|
| 51 |
+
decorator==5.1.1
|
| 52 |
+
seaborn==0.13.0
|
| 53 |
+
sentencepiece==0.2.0
|
| 54 |
+
PyYAML==6.0.1
|
| 55 |
+
black==24.8.0
|
| 56 |
+
protobuf==4.25.1
|
| 57 |
+
pickleshare==0.7.5
|
| 58 |
+
peft==0.13.0
|
| 59 |
+
triton==2.0.0
|
| 60 |
+
nvidia-cuda-runtime-cu11==11.7.99
|
| 61 |
+
Jinja2==3.1.2
|
| 62 |
+
nvidia-cusolver-cu11==11.4.0.1
|
| 63 |
+
executing==1.2.0
|
| 64 |
+
jupyter_client==8.1.0
|
| 65 |
+
pluggy==1.3.0
|
| 66 |
+
cmake==3.30.3
|
| 67 |
+
pytz==2023.3.post1
|
| 68 |
+
aiohappyeyeballs==2.4.2
|
| 69 |
+
kiwisolver==1.4.5
|
| 70 |
+
py-cpuinfo==9.0.0
|
| 71 |
+
Pillow==10.1.0
|
| 72 |
+
ptyprocess==0.7.0
|
| 73 |
+
importlib_resources==6.4.5
|
| 74 |
+
GitPython==3.1.43
|
| 75 |
+
importlib-metadata==6.0.0
|
| 76 |
+
iniconfig==2.0.0
|
| 77 |
+
scikit-learn==1.3.2
|
| 78 |
+
exceptiongroup==1.1.0
|
| 79 |
+
networkx==2.8.6
|
| 80 |
+
accelerate==1.0.0
|
| 81 |
+
nltk==3.8.1
|
| 82 |
+
shutilwhich==1.1.0
|
| 83 |
+
fonttools==4.45.1
|
| 84 |
+
future==0.18.3
|
| 85 |
+
aiohttp==3.10.6
|
| 86 |
+
wcwidth==0.2.5
|
| 87 |
+
idna==3.6
|
| 88 |
+
filelock==3.12.2
|
| 89 |
+
pathspec==0.12.1
|
| 90 |
+
jupyter_core==5.1.0
|
| 91 |
+
lit==18.1.8
|
| 92 |
+
nvidia-curand-cu11==10.2.10.91
|
| 93 |
+
nvidia-cublas-cu11==11.10.3.66
|
| 94 |
+
nvidia-ml-py==12.560.30
|
| 95 |
+
msgpack==1.1.0
|
| 96 |
+
python-dateutil==2.8.2
|
| 97 |
+
blessed==1.20.0
|
| 98 |
+
packaging==23.0
|
| 99 |
+
gitdb==4.0.11
|
| 100 |
+
yarl==1.13.0
|
| 101 |
+
emoji==2.8.0
|
| 102 |
+
tzdata==2023.3
|
| 103 |
+
cycler==0.12.1
|
| 104 |
+
tornado==6.2
|
| 105 |
+
backcall==0.2.0
|
| 106 |
+
plotnine==0.12.4
|
| 107 |
+
ninja==1.11.1.1
|
| 108 |
+
latex==0.7.0
|
| 109 |
+
wandb==0.18.5
|
| 110 |
+
setproctitle==1.3.3
|
| 111 |
+
threadpoolctl==3.2.0
|
| 112 |
+
requests==2.32.3
|
| 113 |
+
pyparsing==3.1.1
|
| 114 |
+
smmap==5.0.1
|
| 115 |
+
pyzmq==23.0.0
|
| 116 |
+
async-timeout==4.0.3
|
| 117 |
+
annotated-types==0.7.0
|
| 118 |
+
matplotlib-inline==0.1.6
|
| 119 |
+
latexcodec==1.0.0
|
| 120 |
+
ipython==8.0.0
|
| 121 |
+
patsy==0.5.3
|
| 122 |
+
contourpy==1.2.0
|
| 123 |
+
multidict==6.1.0
|
| 124 |
+
mizani==0.9.3
|
| 125 |
+
urllib3==2.1.0
|
| 126 |
+
tokenizers==0.20.0
|
| 127 |
+
MarkupSafe==2.1.2
|
| 128 |
+
pip==24.2
|
| 129 |
+
pexpect==4.8.0
|
| 130 |
+
tqdm==4.66.5
|
| 131 |
+
jedi==0.18.2
|
| 132 |
+
pydantic_core==2.23.4
|
| 133 |
+
tempdir==0.7.1
|
| 134 |
+
mpmath==1.2.1
|
| 135 |
+
setuptools==72.1.0
|
| 136 |
+
pytest==7.4.3
|
| 137 |
+
pure-eval==0.2.2
|
| 138 |
+
psutil==5.9.1
|
| 139 |
+
comm==0.1.2
|
| 140 |
+
nvidia-cuda-cupti-cu11==11.7.101
|
| 141 |
+
nvidia-cuda-nvrtc-cu11==11.7.99
|
| 142 |
+
regex==2023.10.3
|
| 143 |
+
platformdirs==2.5.2
|
| 144 |
+
asttokens==2.2.1
|
| 145 |
+
torch==2.0.0
|
| 146 |
+
nvidia-nccl-cu11==2.14.3
|
| 147 |
+
xxhash==3.5.0
|
wandb/run-20241031_000839-cu7972v5/files/wandb-metadata.json
ADDED
|
@@ -0,0 +1,97 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
|
| 3 |
+
"python": "3.9.19",
|
| 4 |
+
"startedAt": "2024-10-31T04:08:39.234664Z",
|
| 5 |
+
"args": [
|
| 6 |
+
"--perturbation",
|
| 7 |
+
"reverse_full",
|
| 8 |
+
"--train_set",
|
| 9 |
+
"10M",
|
| 10 |
+
"--batch_size",
|
| 11 |
+
"3",
|
| 12 |
+
"--epoch",
|
| 13 |
+
"6",
|
| 14 |
+
"--seed",
|
| 15 |
+
"0"
|
| 16 |
+
],
|
| 17 |
+
"program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
|
| 18 |
+
"codePath": "train/train_deep_wandb.py",
|
| 19 |
+
"git": {
|
| 20 |
+
"remote": "git@hf.co:Yaning1001/Impossible_llm.git",
|
| 21 |
+
"commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
|
| 22 |
+
},
|
| 23 |
+
"email": "yaning1001@gmail.com",
|
| 24 |
+
"root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
|
| 25 |
+
"host": "mms-large-2",
|
| 26 |
+
"username": "chunhui",
|
| 27 |
+
"executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
|
| 28 |
+
"codePathLocal": "train_deep_wandb.py",
|
| 29 |
+
"cpu_count": 32,
|
| 30 |
+
"cpu_count_logical": 64,
|
| 31 |
+
"gpu": "NVIDIA RTX A6000",
|
| 32 |
+
"gpu_count": 8,
|
| 33 |
+
"disk": {
|
| 34 |
+
"/": {
|
| 35 |
+
"total": "1888559353856",
|
| 36 |
+
"used": "1727270539264"
|
| 37 |
+
}
|
| 38 |
+
},
|
| 39 |
+
"memory": {
|
| 40 |
+
"total": "202617098240"
|
| 41 |
+
},
|
| 42 |
+
"cpu": {
|
| 43 |
+
"count": 32,
|
| 44 |
+
"countLogical": 64
|
| 45 |
+
},
|
| 46 |
+
"gpu_nvidia": [
|
| 47 |
+
{
|
| 48 |
+
"name": "NVIDIA RTX A6000",
|
| 49 |
+
"memoryTotal": "51527024640",
|
| 50 |
+
"cudaCores": 10752,
|
| 51 |
+
"architecture": "Ampere"
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"name": "NVIDIA RTX A6000",
|
| 55 |
+
"memoryTotal": "51527024640",
|
| 56 |
+
"cudaCores": 10752,
|
| 57 |
+
"architecture": "Ampere"
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"name": "NVIDIA RTX A6000",
|
| 61 |
+
"memoryTotal": "51527024640",
|
| 62 |
+
"cudaCores": 10752,
|
| 63 |
+
"architecture": "Ampere"
|
| 64 |
+
},
|
| 65 |
+
{
|
| 66 |
+
"name": "NVIDIA RTX A6000",
|
| 67 |
+
"memoryTotal": "51527024640",
|
| 68 |
+
"cudaCores": 10752,
|
| 69 |
+
"architecture": "Ampere"
|
| 70 |
+
},
|
| 71 |
+
{
|
| 72 |
+
"name": "NVIDIA RTX A6000",
|
| 73 |
+
"memoryTotal": "51527024640",
|
| 74 |
+
"cudaCores": 10752,
|
| 75 |
+
"architecture": "Ampere"
|
| 76 |
+
},
|
| 77 |
+
{
|
| 78 |
+
"name": "NVIDIA RTX A6000",
|
| 79 |
+
"memoryTotal": "51527024640",
|
| 80 |
+
"cudaCores": 10752,
|
| 81 |
+
"architecture": "Ampere"
|
| 82 |
+
},
|
| 83 |
+
{
|
| 84 |
+
"name": "NVIDIA RTX A6000",
|
| 85 |
+
"memoryTotal": "51527024640",
|
| 86 |
+
"cudaCores": 10752,
|
| 87 |
+
"architecture": "Ampere"
|
| 88 |
+
},
|
| 89 |
+
{
|
| 90 |
+
"name": "NVIDIA RTX A6000",
|
| 91 |
+
"memoryTotal": "51527024640",
|
| 92 |
+
"cudaCores": 10752,
|
| 93 |
+
"architecture": "Ampere"
|
| 94 |
+
}
|
| 95 |
+
],
|
| 96 |
+
"cudaVersion": "11.8"
|
| 97 |
+
}
|
wandb/run-20241031_000839-cu7972v5/logs/debug-internal.log
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"time":"2024-10-31T00:08:39.237073048-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
|
| 2 |
+
{"time":"2024-10-31T00:08:39.237089559-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241031_000839-cu7972v5/logs/debug-core.log"}
|
| 3 |
+
{"time":"2024-10-31T00:08:39.341923634-04:00","level":"INFO","msg":"created new stream","id":"cu7972v5"}
|
| 4 |
+
{"time":"2024-10-31T00:08:39.341944484-04:00","level":"INFO","msg":"stream: started","id":"cu7972v5"}
|
| 5 |
+
{"time":"2024-10-31T00:08:39.341969834-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"cu7972v5"}}
|
| 6 |
+
{"time":"2024-10-31T00:08:39.341961834-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"cu7972v5"}}
|
| 7 |
+
{"time":"2024-10-31T00:08:39.341998404-04:00","level":"INFO","msg":"sender: started","stream_id":"cu7972v5"}
|
| 8 |
+
{"time":"2024-10-31T00:08:39.509463672-04:00","level":"INFO","msg":"Starting system monitor"}
|
wandb/run-20241031_000839-cu7972v5/logs/debug.log
ADDED
|
@@ -0,0 +1,26 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
2024-10-31 00:08:39,232 INFO MainThread:477298 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
|
| 2 |
+
2024-10-31 00:08:39,232 INFO MainThread:477298 [wandb_setup.py:_flush():79] Configure stats pid to 477298
|
| 3 |
+
2024-10-31 00:08:39,232 INFO MainThread:477298 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
|
| 4 |
+
2024-10-31 00:08:39,232 INFO MainThread:477298 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
|
| 5 |
+
2024-10-31 00:08:39,232 INFO MainThread:477298 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
|
| 6 |
+
2024-10-31 00:08:39,232 INFO MainThread:477298 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
|
| 7 |
+
2024-10-31 00:08:39,232 INFO MainThread:477298 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
|
| 8 |
+
2024-10-31 00:08:39,232 INFO MainThread:477298 [wandb_setup.py:_flush():79] Applying login settings: {}
|
| 9 |
+
2024-10-31 00:08:39,232 INFO MainThread:477298 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241031_000839-cu7972v5/logs/debug.log
|
| 10 |
+
2024-10-31 00:08:39,232 INFO MainThread:477298 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241031_000839-cu7972v5/logs/debug-internal.log
|
| 11 |
+
2024-10-31 00:08:39,232 INFO MainThread:477298 [wandb_init.py:init():621] calling init triggers
|
| 12 |
+
2024-10-31 00:08:39,232 INFO MainThread:477298 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
|
| 13 |
+
config: {}
|
| 14 |
+
2024-10-31 00:08:39,232 INFO MainThread:477298 [wandb_init.py:init():671] starting backend
|
| 15 |
+
2024-10-31 00:08:39,232 INFO MainThread:477298 [wandb_init.py:init():675] sending inform_init request
|
| 16 |
+
2024-10-31 00:08:39,234 INFO MainThread:477298 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
|
| 17 |
+
2024-10-31 00:08:39,234 INFO MainThread:477298 [wandb_init.py:init():688] backend started and connected
|
| 18 |
+
2024-10-31 00:08:39,239 INFO MainThread:477298 [wandb_init.py:init():783] updated telemetry
|
| 19 |
+
2024-10-31 00:08:39,268 INFO MainThread:477298 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
|
| 20 |
+
2024-10-31 00:08:39,506 INFO MainThread:477298 [wandb_init.py:init():867] starting run threads in backend
|
| 21 |
+
2024-10-31 00:08:39,612 INFO MainThread:477298 [wandb_run.py:_console_start():2463] atexit reg
|
| 22 |
+
2024-10-31 00:08:39,612 INFO MainThread:477298 [wandb_run.py:_redirect():2311] redirect: wrap_raw
|
| 23 |
+
2024-10-31 00:08:39,612 INFO MainThread:477298 [wandb_run.py:_redirect():2376] Wrapping output streams.
|
| 24 |
+
2024-10-31 00:08:39,612 INFO MainThread:477298 [wandb_run.py:_redirect():2401] Redirects installed.
|
| 25 |
+
2024-10-31 00:08:39,613 INFO MainThread:477298 [wandb_init.py:init():911] run started, returning control to user process
|
| 26 |
+
2024-10-31 00:08:39,614 INFO MainThread:477298 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'reverse_full', 'train_set': '10M', 'batch_size': 3, 'epoch': 6, 'seed': 0, 'lr': 1e-05}
|
wandb/run-20241031_000839-cu7972v5/run-cu7972v5.wandb
ADDED
|
Binary file (65.5 kB). View file
|
|
|
wandb/run-20241031_001055-32u9qnul/logs/debug-internal.log
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"time":"2024-10-31T00:10:55.97661836-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
|
| 2 |
+
{"time":"2024-10-31T00:10:55.97666224-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241031_001055-32u9qnul/logs/debug-core.log"}
|
| 3 |
+
{"time":"2024-10-31T00:10:56.085295995-04:00","level":"INFO","msg":"created new stream","id":"32u9qnul"}
|
| 4 |
+
{"time":"2024-10-31T00:10:56.085346146-04:00","level":"INFO","msg":"stream: started","id":"32u9qnul"}
|
| 5 |
+
{"time":"2024-10-31T00:10:56.085429147-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"32u9qnul"}}
|
| 6 |
+
{"time":"2024-10-31T00:10:56.085619888-04:00","level":"INFO","msg":"sender: started","stream_id":"32u9qnul"}
|
| 7 |
+
{"time":"2024-10-31T00:10:56.085632288-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"32u9qnul"}}
|
| 8 |
+
{"time":"2024-10-31T00:10:56.308225083-04:00","level":"INFO","msg":"Starting system monitor"}
|
wandb/run-20241031_001055-32u9qnul/logs/debug.log
ADDED
|
@@ -0,0 +1,26 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
2024-10-31 00:10:55,970 INFO MainThread:479383 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
|
| 2 |
+
2024-10-31 00:10:55,970 INFO MainThread:479383 [wandb_setup.py:_flush():79] Configure stats pid to 479383
|
| 3 |
+
2024-10-31 00:10:55,971 INFO MainThread:479383 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
|
| 4 |
+
2024-10-31 00:10:55,971 INFO MainThread:479383 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
|
| 5 |
+
2024-10-31 00:10:55,971 INFO MainThread:479383 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
|
| 6 |
+
2024-10-31 00:10:55,971 INFO MainThread:479383 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
|
| 7 |
+
2024-10-31 00:10:55,971 INFO MainThread:479383 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
|
| 8 |
+
2024-10-31 00:10:55,971 INFO MainThread:479383 [wandb_setup.py:_flush():79] Applying login settings: {}
|
| 9 |
+
2024-10-31 00:10:55,971 INFO MainThread:479383 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241031_001055-32u9qnul/logs/debug.log
|
| 10 |
+
2024-10-31 00:10:55,971 INFO MainThread:479383 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241031_001055-32u9qnul/logs/debug-internal.log
|
| 11 |
+
2024-10-31 00:10:55,971 INFO MainThread:479383 [wandb_init.py:init():621] calling init triggers
|
| 12 |
+
2024-10-31 00:10:55,971 INFO MainThread:479383 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
|
| 13 |
+
config: {}
|
| 14 |
+
2024-10-31 00:10:55,971 INFO MainThread:479383 [wandb_init.py:init():671] starting backend
|
| 15 |
+
2024-10-31 00:10:55,971 INFO MainThread:479383 [wandb_init.py:init():675] sending inform_init request
|
| 16 |
+
2024-10-31 00:10:55,972 INFO MainThread:479383 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
|
| 17 |
+
2024-10-31 00:10:55,973 INFO MainThread:479383 [wandb_init.py:init():688] backend started and connected
|
| 18 |
+
2024-10-31 00:10:55,976 INFO MainThread:479383 [wandb_init.py:init():783] updated telemetry
|
| 19 |
+
2024-10-31 00:10:56,016 INFO MainThread:479383 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
|
| 20 |
+
2024-10-31 00:10:56,304 INFO MainThread:479383 [wandb_init.py:init():867] starting run threads in backend
|
| 21 |
+
2024-10-31 00:10:56,408 INFO MainThread:479383 [wandb_run.py:_console_start():2463] atexit reg
|
| 22 |
+
2024-10-31 00:10:56,408 INFO MainThread:479383 [wandb_run.py:_redirect():2311] redirect: wrap_raw
|
| 23 |
+
2024-10-31 00:10:56,409 INFO MainThread:479383 [wandb_run.py:_redirect():2376] Wrapping output streams.
|
| 24 |
+
2024-10-31 00:10:56,409 INFO MainThread:479383 [wandb_run.py:_redirect():2401] Redirects installed.
|
| 25 |
+
2024-10-31 00:10:56,410 INFO MainThread:479383 [wandb_init.py:init():911] run started, returning control to user process
|
| 26 |
+
2024-10-31 00:10:56,410 INFO MainThread:479383 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'reverse_full', 'train_set': '10M', 'batch_size': 3, 'epoch': 6, 'seed': 0, 'lr': 1e-05}
|
wandb/run-20241031_002020-u516mysu/files/output.log
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
wandb/run-20241031_114700-jx2hqvx3/files/wandb-metadata.json
ADDED
|
@@ -0,0 +1,97 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
|
| 3 |
+
"python": "3.9.19",
|
| 4 |
+
"startedAt": "2024-10-31T15:47:00.195452Z",
|
| 5 |
+
"args": [
|
| 6 |
+
"--perturbation",
|
| 7 |
+
"reverse_full",
|
| 8 |
+
"--train_set",
|
| 9 |
+
"10M",
|
| 10 |
+
"--batch_size",
|
| 11 |
+
"3",
|
| 12 |
+
"--epoch",
|
| 13 |
+
"6",
|
| 14 |
+
"--seed",
|
| 15 |
+
"0"
|
| 16 |
+
],
|
| 17 |
+
"program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
|
| 18 |
+
"codePath": "train/train_deep_wandb.py",
|
| 19 |
+
"git": {
|
| 20 |
+
"remote": "git@hf.co:Yaning1001/Impossible_llm.git",
|
| 21 |
+
"commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
|
| 22 |
+
},
|
| 23 |
+
"email": "yaning1001@gmail.com",
|
| 24 |
+
"root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
|
| 25 |
+
"host": "mms-large-2",
|
| 26 |
+
"username": "chunhui",
|
| 27 |
+
"executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
|
| 28 |
+
"codePathLocal": "train_deep_wandb.py",
|
| 29 |
+
"cpu_count": 32,
|
| 30 |
+
"cpu_count_logical": 64,
|
| 31 |
+
"gpu": "NVIDIA RTX A6000",
|
| 32 |
+
"gpu_count": 8,
|
| 33 |
+
"disk": {
|
| 34 |
+
"/": {
|
| 35 |
+
"total": "1888559353856",
|
| 36 |
+
"used": "1753158594560"
|
| 37 |
+
}
|
| 38 |
+
},
|
| 39 |
+
"memory": {
|
| 40 |
+
"total": "202617098240"
|
| 41 |
+
},
|
| 42 |
+
"cpu": {
|
| 43 |
+
"count": 32,
|
| 44 |
+
"countLogical": 64
|
| 45 |
+
},
|
| 46 |
+
"gpu_nvidia": [
|
| 47 |
+
{
|
| 48 |
+
"name": "NVIDIA RTX A6000",
|
| 49 |
+
"memoryTotal": "51527024640",
|
| 50 |
+
"cudaCores": 10752,
|
| 51 |
+
"architecture": "Ampere"
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"name": "NVIDIA RTX A6000",
|
| 55 |
+
"memoryTotal": "51527024640",
|
| 56 |
+
"cudaCores": 10752,
|
| 57 |
+
"architecture": "Ampere"
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"name": "NVIDIA RTX A6000",
|
| 61 |
+
"memoryTotal": "51527024640",
|
| 62 |
+
"cudaCores": 10752,
|
| 63 |
+
"architecture": "Ampere"
|
| 64 |
+
},
|
| 65 |
+
{
|
| 66 |
+
"name": "NVIDIA RTX A6000",
|
| 67 |
+
"memoryTotal": "51527024640",
|
| 68 |
+
"cudaCores": 10752,
|
| 69 |
+
"architecture": "Ampere"
|
| 70 |
+
},
|
| 71 |
+
{
|
| 72 |
+
"name": "NVIDIA RTX A6000",
|
| 73 |
+
"memoryTotal": "51527024640",
|
| 74 |
+
"cudaCores": 10752,
|
| 75 |
+
"architecture": "Ampere"
|
| 76 |
+
},
|
| 77 |
+
{
|
| 78 |
+
"name": "NVIDIA RTX A6000",
|
| 79 |
+
"memoryTotal": "51527024640",
|
| 80 |
+
"cudaCores": 10752,
|
| 81 |
+
"architecture": "Ampere"
|
| 82 |
+
},
|
| 83 |
+
{
|
| 84 |
+
"name": "NVIDIA RTX A6000",
|
| 85 |
+
"memoryTotal": "51527024640",
|
| 86 |
+
"cudaCores": 10752,
|
| 87 |
+
"architecture": "Ampere"
|
| 88 |
+
},
|
| 89 |
+
{
|
| 90 |
+
"name": "NVIDIA RTX A6000",
|
| 91 |
+
"memoryTotal": "51527024640",
|
| 92 |
+
"cudaCores": 10752,
|
| 93 |
+
"architecture": "Ampere"
|
| 94 |
+
}
|
| 95 |
+
],
|
| 96 |
+
"cudaVersion": "11.8"
|
| 97 |
+
}
|
wandb/run-20241101_012438-61w48leq/files/config.yaml
ADDED
|
@@ -0,0 +1,49 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
_wandb:
|
| 2 |
+
value:
|
| 3 |
+
cli_version: 0.18.5
|
| 4 |
+
m: []
|
| 5 |
+
python_version: 3.9.19
|
| 6 |
+
t:
|
| 7 |
+
"1":
|
| 8 |
+
- 1
|
| 9 |
+
- 5
|
| 10 |
+
- 11
|
| 11 |
+
- 49
|
| 12 |
+
- 51
|
| 13 |
+
- 53
|
| 14 |
+
- 55
|
| 15 |
+
- 71
|
| 16 |
+
- 98
|
| 17 |
+
"2":
|
| 18 |
+
- 1
|
| 19 |
+
- 5
|
| 20 |
+
- 11
|
| 21 |
+
- 49
|
| 22 |
+
- 51
|
| 23 |
+
- 53
|
| 24 |
+
- 55
|
| 25 |
+
- 71
|
| 26 |
+
- 98
|
| 27 |
+
"3":
|
| 28 |
+
- 13
|
| 29 |
+
- 23
|
| 30 |
+
- 55
|
| 31 |
+
"4": 3.9.19
|
| 32 |
+
"5": 0.18.5
|
| 33 |
+
"6": 4.45.1
|
| 34 |
+
"8":
|
| 35 |
+
- 5
|
| 36 |
+
"12": 0.18.5
|
| 37 |
+
"13": linux-x86_64
|
| 38 |
+
batch_size:
|
| 39 |
+
value: 3
|
| 40 |
+
epoch:
|
| 41 |
+
value: 6
|
| 42 |
+
lr:
|
| 43 |
+
value: 5e-06
|
| 44 |
+
perturbation:
|
| 45 |
+
value: shuffle_nodeterministic
|
| 46 |
+
seed:
|
| 47 |
+
value: 0
|
| 48 |
+
train_set:
|
| 49 |
+
value: 10M
|
wandb/run-20241101_012438-61w48leq/files/output.log
ADDED
|
@@ -0,0 +1,12 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Traceback (most recent call last):
|
| 2 |
+
File "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py", line 164, in <module>
|
| 3 |
+
dataset = load_dataset('babylm_dataset_test.py', name=dataset_name, trust_remote_code=True)
|
| 4 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/datasets/load.py", line 2074, in load_dataset
|
| 5 |
+
builder_instance = load_dataset_builder(
|
| 6 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/datasets/load.py", line 1832, in load_dataset_builder
|
| 7 |
+
builder_instance: DatasetBuilder = builder_cls(
|
| 8 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/datasets/builder.py", line 342, in __init__
|
| 9 |
+
self.config, self.config_id = self._create_builder_config(
|
| 10 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/datasets/builder.py", line 569, in _create_builder_config
|
| 11 |
+
raise ValueError(
|
| 12 |
+
ValueError: BuilderConfig 'babylm_shuffle_nodeterministic_10M_seed0' not found. Available: ['babylm_hop_control_10M_seed0', 'babylm_hop_tokens4_10M_seed0', 'babylm_hop_words4_10M_seed0', 'babylm_reverse_control_10M_seed0', 'babylm_reverse_partial_10M_seed0', 'babylm_reverse_full_10M_seed0', 'babylm_shuffle_control_10M_seed0', 'babylm_shuffle_nondeterministic_10M_seed0', 'babylm_shuffle_deterministic21_10M_seed0', 'babylm_shuffle_deterministic57_10M_seed0', 'babylm_shuffle_deterministic84_10M_seed0', 'babylm_shuffle_local3_10M_seed0', 'babylm_shuffle_local5_10M_seed0', 'babylm_shuffle_local10_10M_seed0', 'babylm_shuffle_even_odd_10M_seed0']
|
wandb/run-20241101_012438-61w48leq/files/wandb-metadata.json
ADDED
|
@@ -0,0 +1,97 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
|
| 3 |
+
"python": "3.9.19",
|
| 4 |
+
"startedAt": "2024-11-01T05:24:38.161201Z",
|
| 5 |
+
"args": [
|
| 6 |
+
"--perturbation",
|
| 7 |
+
"shuffle_nodeterministic",
|
| 8 |
+
"--train_set",
|
| 9 |
+
"10M",
|
| 10 |
+
"--batch_size",
|
| 11 |
+
"3",
|
| 12 |
+
"--epoch",
|
| 13 |
+
"6",
|
| 14 |
+
"--seed",
|
| 15 |
+
"0"
|
| 16 |
+
],
|
| 17 |
+
"program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
|
| 18 |
+
"codePath": "train/train_deep_wandb.py",
|
| 19 |
+
"git": {
|
| 20 |
+
"remote": "git@hf.co:Yaning1001/Impossible_llm.git",
|
| 21 |
+
"commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
|
| 22 |
+
},
|
| 23 |
+
"email": "yaning1001@gmail.com",
|
| 24 |
+
"root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
|
| 25 |
+
"host": "mms-large-2",
|
| 26 |
+
"username": "chunhui",
|
| 27 |
+
"executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
|
| 28 |
+
"codePathLocal": "train_deep_wandb.py",
|
| 29 |
+
"cpu_count": 32,
|
| 30 |
+
"cpu_count_logical": 64,
|
| 31 |
+
"gpu": "NVIDIA RTX A6000",
|
| 32 |
+
"gpu_count": 8,
|
| 33 |
+
"disk": {
|
| 34 |
+
"/": {
|
| 35 |
+
"total": "1888559353856",
|
| 36 |
+
"used": "1753992159232"
|
| 37 |
+
}
|
| 38 |
+
},
|
| 39 |
+
"memory": {
|
| 40 |
+
"total": "202617098240"
|
| 41 |
+
},
|
| 42 |
+
"cpu": {
|
| 43 |
+
"count": 32,
|
| 44 |
+
"countLogical": 64
|
| 45 |
+
},
|
| 46 |
+
"gpu_nvidia": [
|
| 47 |
+
{
|
| 48 |
+
"name": "NVIDIA RTX A6000",
|
| 49 |
+
"memoryTotal": "51527024640",
|
| 50 |
+
"cudaCores": 10752,
|
| 51 |
+
"architecture": "Ampere"
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"name": "NVIDIA RTX A6000",
|
| 55 |
+
"memoryTotal": "51527024640",
|
| 56 |
+
"cudaCores": 10752,
|
| 57 |
+
"architecture": "Ampere"
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"name": "NVIDIA RTX A6000",
|
| 61 |
+
"memoryTotal": "51527024640",
|
| 62 |
+
"cudaCores": 10752,
|
| 63 |
+
"architecture": "Ampere"
|
| 64 |
+
},
|
| 65 |
+
{
|
| 66 |
+
"name": "NVIDIA RTX A6000",
|
| 67 |
+
"memoryTotal": "51527024640",
|
| 68 |
+
"cudaCores": 10752,
|
| 69 |
+
"architecture": "Ampere"
|
| 70 |
+
},
|
| 71 |
+
{
|
| 72 |
+
"name": "NVIDIA RTX A6000",
|
| 73 |
+
"memoryTotal": "51527024640",
|
| 74 |
+
"cudaCores": 10752,
|
| 75 |
+
"architecture": "Ampere"
|
| 76 |
+
},
|
| 77 |
+
{
|
| 78 |
+
"name": "NVIDIA RTX A6000",
|
| 79 |
+
"memoryTotal": "51527024640",
|
| 80 |
+
"cudaCores": 10752,
|
| 81 |
+
"architecture": "Ampere"
|
| 82 |
+
},
|
| 83 |
+
{
|
| 84 |
+
"name": "NVIDIA RTX A6000",
|
| 85 |
+
"memoryTotal": "51527024640",
|
| 86 |
+
"cudaCores": 10752,
|
| 87 |
+
"architecture": "Ampere"
|
| 88 |
+
},
|
| 89 |
+
{
|
| 90 |
+
"name": "NVIDIA RTX A6000",
|
| 91 |
+
"memoryTotal": "51527024640",
|
| 92 |
+
"cudaCores": 10752,
|
| 93 |
+
"architecture": "Ampere"
|
| 94 |
+
}
|
| 95 |
+
],
|
| 96 |
+
"cudaVersion": "11.8"
|
| 97 |
+
}
|
wandb/run-20241101_012438-61w48leq/files/wandb-summary.json
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
{"_wandb":{"runtime":0}}
|
wandb/run-20241101_012438-61w48leq/run-61w48leq.wandb
ADDED
|
Binary file (3.43 kB). View file
|
|
|
wandb/run-20241101_012734-m18lsdzn/files/wandb-metadata.json
ADDED
|
@@ -0,0 +1,97 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
|
| 3 |
+
"python": "3.9.19",
|
| 4 |
+
"startedAt": "2024-11-01T05:27:34.134177Z",
|
| 5 |
+
"args": [
|
| 6 |
+
"--perturbation",
|
| 7 |
+
"shuffle_nondeterministic",
|
| 8 |
+
"--train_set",
|
| 9 |
+
"10M",
|
| 10 |
+
"--batch_size",
|
| 11 |
+
"3",
|
| 12 |
+
"--epoch",
|
| 13 |
+
"6",
|
| 14 |
+
"--seed",
|
| 15 |
+
"0"
|
| 16 |
+
],
|
| 17 |
+
"program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
|
| 18 |
+
"codePath": "train/train_deep_wandb.py",
|
| 19 |
+
"git": {
|
| 20 |
+
"remote": "git@hf.co:Yaning1001/Impossible_llm.git",
|
| 21 |
+
"commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
|
| 22 |
+
},
|
| 23 |
+
"email": "yaning1001@gmail.com",
|
| 24 |
+
"root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
|
| 25 |
+
"host": "mms-large-2",
|
| 26 |
+
"username": "chunhui",
|
| 27 |
+
"executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
|
| 28 |
+
"codePathLocal": "train_deep_wandb.py",
|
| 29 |
+
"cpu_count": 32,
|
| 30 |
+
"cpu_count_logical": 64,
|
| 31 |
+
"gpu": "NVIDIA RTX A6000",
|
| 32 |
+
"gpu_count": 8,
|
| 33 |
+
"disk": {
|
| 34 |
+
"/": {
|
| 35 |
+
"total": "1888559353856",
|
| 36 |
+
"used": "1753992269824"
|
| 37 |
+
}
|
| 38 |
+
},
|
| 39 |
+
"memory": {
|
| 40 |
+
"total": "202617098240"
|
| 41 |
+
},
|
| 42 |
+
"cpu": {
|
| 43 |
+
"count": 32,
|
| 44 |
+
"countLogical": 64
|
| 45 |
+
},
|
| 46 |
+
"gpu_nvidia": [
|
| 47 |
+
{
|
| 48 |
+
"name": "NVIDIA RTX A6000",
|
| 49 |
+
"memoryTotal": "51527024640",
|
| 50 |
+
"cudaCores": 10752,
|
| 51 |
+
"architecture": "Ampere"
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"name": "NVIDIA RTX A6000",
|
| 55 |
+
"memoryTotal": "51527024640",
|
| 56 |
+
"cudaCores": 10752,
|
| 57 |
+
"architecture": "Ampere"
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"name": "NVIDIA RTX A6000",
|
| 61 |
+
"memoryTotal": "51527024640",
|
| 62 |
+
"cudaCores": 10752,
|
| 63 |
+
"architecture": "Ampere"
|
| 64 |
+
},
|
| 65 |
+
{
|
| 66 |
+
"name": "NVIDIA RTX A6000",
|
| 67 |
+
"memoryTotal": "51527024640",
|
| 68 |
+
"cudaCores": 10752,
|
| 69 |
+
"architecture": "Ampere"
|
| 70 |
+
},
|
| 71 |
+
{
|
| 72 |
+
"name": "NVIDIA RTX A6000",
|
| 73 |
+
"memoryTotal": "51527024640",
|
| 74 |
+
"cudaCores": 10752,
|
| 75 |
+
"architecture": "Ampere"
|
| 76 |
+
},
|
| 77 |
+
{
|
| 78 |
+
"name": "NVIDIA RTX A6000",
|
| 79 |
+
"memoryTotal": "51527024640",
|
| 80 |
+
"cudaCores": 10752,
|
| 81 |
+
"architecture": "Ampere"
|
| 82 |
+
},
|
| 83 |
+
{
|
| 84 |
+
"name": "NVIDIA RTX A6000",
|
| 85 |
+
"memoryTotal": "51527024640",
|
| 86 |
+
"cudaCores": 10752,
|
| 87 |
+
"architecture": "Ampere"
|
| 88 |
+
},
|
| 89 |
+
{
|
| 90 |
+
"name": "NVIDIA RTX A6000",
|
| 91 |
+
"memoryTotal": "51527024640",
|
| 92 |
+
"cudaCores": 10752,
|
| 93 |
+
"architecture": "Ampere"
|
| 94 |
+
}
|
| 95 |
+
],
|
| 96 |
+
"cudaVersion": "11.8"
|
| 97 |
+
}
|
wandb/run-20241101_012734-m18lsdzn/logs/debug-internal.log
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"time":"2024-11-01T01:27:34.136793008-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
|
| 2 |
+
{"time":"2024-11-01T01:27:34.136814509-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241101_012734-m18lsdzn/logs/debug-core.log"}
|
| 3 |
+
{"time":"2024-11-01T01:27:34.243830731-04:00","level":"INFO","msg":"created new stream","id":"m18lsdzn"}
|
| 4 |
+
{"time":"2024-11-01T01:27:34.243863612-04:00","level":"INFO","msg":"stream: started","id":"m18lsdzn"}
|
| 5 |
+
{"time":"2024-11-01T01:27:34.244001533-04:00","level":"INFO","msg":"sender: started","stream_id":"m18lsdzn"}
|
| 6 |
+
{"time":"2024-11-01T01:27:34.243939362-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"m18lsdzn"}}
|
| 7 |
+
{"time":"2024-11-01T01:27:34.243890052-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"m18lsdzn"}}
|
| 8 |
+
{"time":"2024-11-01T01:27:34.468881692-04:00","level":"INFO","msg":"Starting system monitor"}
|
wandb/run-20241101_092804-qhsuxbxe/run-qhsuxbxe.wandb
ADDED
|
Binary file (65.5 kB). View file
|
|
|
wandb/run-20241101_200517-77b12390/run-77b12390.wandb
ADDED
|
File without changes
|
wandb/run-20241101_200517-iopieyi0/files/config.yaml
ADDED
|
@@ -0,0 +1,49 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
_wandb:
|
| 2 |
+
value:
|
| 3 |
+
cli_version: 0.18.5
|
| 4 |
+
m: []
|
| 5 |
+
python_version: 3.9.19
|
| 6 |
+
t:
|
| 7 |
+
"1":
|
| 8 |
+
- 1
|
| 9 |
+
- 5
|
| 10 |
+
- 11
|
| 11 |
+
- 49
|
| 12 |
+
- 51
|
| 13 |
+
- 53
|
| 14 |
+
- 55
|
| 15 |
+
- 71
|
| 16 |
+
- 98
|
| 17 |
+
"2":
|
| 18 |
+
- 1
|
| 19 |
+
- 5
|
| 20 |
+
- 11
|
| 21 |
+
- 49
|
| 22 |
+
- 51
|
| 23 |
+
- 53
|
| 24 |
+
- 55
|
| 25 |
+
- 71
|
| 26 |
+
- 98
|
| 27 |
+
"3":
|
| 28 |
+
- 13
|
| 29 |
+
- 23
|
| 30 |
+
- 55
|
| 31 |
+
"4": 3.9.19
|
| 32 |
+
"5": 0.18.5
|
| 33 |
+
"6": 4.45.1
|
| 34 |
+
"8":
|
| 35 |
+
- 5
|
| 36 |
+
"12": 0.18.5
|
| 37 |
+
"13": linux-x86_64
|
| 38 |
+
batch_size:
|
| 39 |
+
value: 3
|
| 40 |
+
epoch:
|
| 41 |
+
value: 3
|
| 42 |
+
lr:
|
| 43 |
+
value: 5e-06
|
| 44 |
+
perturbation:
|
| 45 |
+
value: shuffle_nondeterministic
|
| 46 |
+
seed:
|
| 47 |
+
value: 0
|
| 48 |
+
train_set:
|
| 49 |
+
value: 10M
|
wandb/run-20241101_200517-iopieyi0/files/output.log
ADDED
|
@@ -0,0 +1,42 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Downloading shards: 0%| | 0/2 [00:07<?, ?it/s]
|
| 2 |
+
Error in sys.excepthook:
|
| 3 |
+
Traceback (most recent call last):
|
| 4 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/linecache.py", line 46, in getlines
|
| 5 |
+
return updatecache(filename, module_globals)
|
| 6 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/linecache.py", line 136, in updatecache
|
| 7 |
+
with tokenize.open(fullname) as fp:
|
| 8 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/tokenize.py", line 394, in open
|
| 9 |
+
encoding, lines = detect_encoding(buffer.readline)
|
| 10 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/tokenize.py", line 363, in detect_encoding
|
| 11 |
+
first = read_or_stop()
|
| 12 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/tokenize.py", line 321, in read_or_stop
|
| 13 |
+
return readline()
|
| 14 |
+
KeyboardInterrupt
|
| 15 |
+
|
| 16 |
+
Original exception was:
|
| 17 |
+
Traceback (most recent call last):
|
| 18 |
+
File "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py", line 173, in <module>
|
| 19 |
+
model = AutoModelForCausalLM.from_pretrained(model_name,
|
| 20 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/models/auto/auto_factory.py", line 564, in from_pretrained
|
| 21 |
+
return model_class.from_pretrained(
|
| 22 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/modeling_utils.py", line 3769, in from_pretrained
|
| 23 |
+
resolved_archive_file, sharded_metadata = get_checkpoint_shard_files(
|
| 24 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/utils/hub.py", line 1098, in get_checkpoint_shard_files
|
| 25 |
+
cached_filename = cached_file(
|
| 26 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/utils/hub.py", line 403, in cached_file
|
| 27 |
+
resolved_file = hf_hub_download(
|
| 28 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/utils/_deprecation.py", line 101, in inner_f
|
| 29 |
+
return f(*args, **kwargs)
|
| 30 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/utils/_validators.py", line 114, in _inner_fn
|
| 31 |
+
return fn(*args, **kwargs)
|
| 32 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/file_download.py", line 1232, in hf_hub_download
|
| 33 |
+
return _hf_hub_download_to_cache_dir(
|
| 34 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/file_download.py", line 1380, in _hf_hub_download_to_cache_dir
|
| 35 |
+
with WeakFileLock(lock_path):
|
| 36 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/contextlib.py", line 119, in __enter__
|
| 37 |
+
return next(self.gen)
|
| 38 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/utils/_fixes.py", line 98, in WeakFileLock
|
| 39 |
+
lock.acquire()
|
| 40 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/filelock/_api.py", line 225, in acquire
|
| 41 |
+
time.sleep(poll_interval)
|
| 42 |
+
KeyboardInterrupt
|
wandb/run-20241101_200517-iopieyi0/files/requirements.txt
ADDED
|
@@ -0,0 +1,147 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
funcsigs==1.0.2
|
| 2 |
+
sentry-sdk==2.17.0
|
| 3 |
+
multiprocess==0.70.16
|
| 4 |
+
numpy==1.26.2
|
| 5 |
+
pluralizer==1.2.0
|
| 6 |
+
debugpy==1.6.7
|
| 7 |
+
nvidia-cudnn-cu11==8.5.0.96
|
| 8 |
+
deepspeed==0.15.2
|
| 9 |
+
data==0.4
|
| 10 |
+
pandas==2.1.3
|
| 11 |
+
tomli==2.0.1
|
| 12 |
+
charset-normalizer==3.3.2
|
| 13 |
+
attrs==24.2.0
|
| 14 |
+
aiosignal==1.3.1
|
| 15 |
+
fsspec==2023.10.0
|
| 16 |
+
nvidia-cusparse-cu11==11.7.4.91
|
| 17 |
+
zipp==3.12.0
|
| 18 |
+
mypy-extensions==1.0.0
|
| 19 |
+
datasets==3.0.1
|
| 20 |
+
joblib==1.3.2
|
| 21 |
+
hjson==3.1.0
|
| 22 |
+
traitlets==5.7.1
|
| 23 |
+
stack-data==0.6.0
|
| 24 |
+
transformers==4.45.1
|
| 25 |
+
sympy==1.11.1
|
| 26 |
+
Pygments==2.15.0
|
| 27 |
+
docker-pycreds==0.4.0
|
| 28 |
+
dill==0.3.8
|
| 29 |
+
wheel==0.44.0
|
| 30 |
+
prompt-toolkit==3.0.30
|
| 31 |
+
parso==0.8.3
|
| 32 |
+
ipykernel==6.23.1
|
| 33 |
+
pyarrow==17.0.0
|
| 34 |
+
certifi==2023.11.17
|
| 35 |
+
nvidia-cufft-cu11==10.9.0.58
|
| 36 |
+
six==1.16.0
|
| 37 |
+
pydantic==2.9.2
|
| 38 |
+
click==8.1.7
|
| 39 |
+
nest-asyncio==1.5.6
|
| 40 |
+
gmpy2==2.1.0
|
| 41 |
+
matplotlib==3.8.2
|
| 42 |
+
scipy==1.11.4
|
| 43 |
+
typing_extensions==4.12.2
|
| 44 |
+
statsmodels==0.14.0
|
| 45 |
+
huggingface-hub==0.25.0
|
| 46 |
+
frozenlist==1.4.1
|
| 47 |
+
gpustat==1.1.1
|
| 48 |
+
nvidia-nvtx-cu11==11.7.91
|
| 49 |
+
safetensors==0.4.5
|
| 50 |
+
stanza==1.9.2
|
| 51 |
+
decorator==5.1.1
|
| 52 |
+
seaborn==0.13.0
|
| 53 |
+
sentencepiece==0.2.0
|
| 54 |
+
PyYAML==6.0.1
|
| 55 |
+
black==24.8.0
|
| 56 |
+
protobuf==4.25.1
|
| 57 |
+
pickleshare==0.7.5
|
| 58 |
+
peft==0.13.0
|
| 59 |
+
triton==2.0.0
|
| 60 |
+
nvidia-cuda-runtime-cu11==11.7.99
|
| 61 |
+
Jinja2==3.1.2
|
| 62 |
+
nvidia-cusolver-cu11==11.4.0.1
|
| 63 |
+
executing==1.2.0
|
| 64 |
+
jupyter_client==8.1.0
|
| 65 |
+
pluggy==1.3.0
|
| 66 |
+
cmake==3.30.3
|
| 67 |
+
pytz==2023.3.post1
|
| 68 |
+
aiohappyeyeballs==2.4.2
|
| 69 |
+
kiwisolver==1.4.5
|
| 70 |
+
py-cpuinfo==9.0.0
|
| 71 |
+
Pillow==10.1.0
|
| 72 |
+
ptyprocess==0.7.0
|
| 73 |
+
importlib_resources==6.4.5
|
| 74 |
+
GitPython==3.1.43
|
| 75 |
+
importlib-metadata==6.0.0
|
| 76 |
+
iniconfig==2.0.0
|
| 77 |
+
scikit-learn==1.3.2
|
| 78 |
+
exceptiongroup==1.1.0
|
| 79 |
+
networkx==2.8.6
|
| 80 |
+
accelerate==1.0.0
|
| 81 |
+
nltk==3.8.1
|
| 82 |
+
shutilwhich==1.1.0
|
| 83 |
+
fonttools==4.45.1
|
| 84 |
+
future==0.18.3
|
| 85 |
+
aiohttp==3.10.6
|
| 86 |
+
wcwidth==0.2.5
|
| 87 |
+
idna==3.6
|
| 88 |
+
filelock==3.12.2
|
| 89 |
+
pathspec==0.12.1
|
| 90 |
+
jupyter_core==5.1.0
|
| 91 |
+
lit==18.1.8
|
| 92 |
+
nvidia-curand-cu11==10.2.10.91
|
| 93 |
+
nvidia-cublas-cu11==11.10.3.66
|
| 94 |
+
nvidia-ml-py==12.560.30
|
| 95 |
+
msgpack==1.1.0
|
| 96 |
+
python-dateutil==2.8.2
|
| 97 |
+
blessed==1.20.0
|
| 98 |
+
packaging==23.0
|
| 99 |
+
gitdb==4.0.11
|
| 100 |
+
yarl==1.13.0
|
| 101 |
+
emoji==2.8.0
|
| 102 |
+
tzdata==2023.3
|
| 103 |
+
cycler==0.12.1
|
| 104 |
+
tornado==6.2
|
| 105 |
+
backcall==0.2.0
|
| 106 |
+
plotnine==0.12.4
|
| 107 |
+
ninja==1.11.1.1
|
| 108 |
+
latex==0.7.0
|
| 109 |
+
wandb==0.18.5
|
| 110 |
+
setproctitle==1.3.3
|
| 111 |
+
threadpoolctl==3.2.0
|
| 112 |
+
requests==2.32.3
|
| 113 |
+
pyparsing==3.1.1
|
| 114 |
+
smmap==5.0.1
|
| 115 |
+
pyzmq==23.0.0
|
| 116 |
+
async-timeout==4.0.3
|
| 117 |
+
annotated-types==0.7.0
|
| 118 |
+
matplotlib-inline==0.1.6
|
| 119 |
+
latexcodec==1.0.0
|
| 120 |
+
ipython==8.0.0
|
| 121 |
+
patsy==0.5.3
|
| 122 |
+
contourpy==1.2.0
|
| 123 |
+
multidict==6.1.0
|
| 124 |
+
mizani==0.9.3
|
| 125 |
+
urllib3==2.1.0
|
| 126 |
+
tokenizers==0.20.0
|
| 127 |
+
MarkupSafe==2.1.2
|
| 128 |
+
pip==24.2
|
| 129 |
+
pexpect==4.8.0
|
| 130 |
+
tqdm==4.66.5
|
| 131 |
+
jedi==0.18.2
|
| 132 |
+
pydantic_core==2.23.4
|
| 133 |
+
tempdir==0.7.1
|
| 134 |
+
mpmath==1.2.1
|
| 135 |
+
setuptools==72.1.0
|
| 136 |
+
pytest==7.4.3
|
| 137 |
+
pure-eval==0.2.2
|
| 138 |
+
psutil==5.9.1
|
| 139 |
+
comm==0.1.2
|
| 140 |
+
nvidia-cuda-cupti-cu11==11.7.101
|
| 141 |
+
nvidia-cuda-nvrtc-cu11==11.7.99
|
| 142 |
+
regex==2023.10.3
|
| 143 |
+
platformdirs==2.5.2
|
| 144 |
+
asttokens==2.2.1
|
| 145 |
+
torch==2.0.0
|
| 146 |
+
nvidia-nccl-cu11==2.14.3
|
| 147 |
+
xxhash==3.5.0
|
wandb/run-20241101_200517-iopieyi0/files/wandb-metadata.json
ADDED
|
@@ -0,0 +1,97 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
|
| 3 |
+
"python": "3.9.19",
|
| 4 |
+
"startedAt": "2024-11-02T00:05:17.140953Z",
|
| 5 |
+
"args": [
|
| 6 |
+
"--perturbation",
|
| 7 |
+
"shuffle_nondeterministic",
|
| 8 |
+
"--train_set",
|
| 9 |
+
"10M",
|
| 10 |
+
"--batch_size",
|
| 11 |
+
"3",
|
| 12 |
+
"--epoch",
|
| 13 |
+
"3",
|
| 14 |
+
"--seed",
|
| 15 |
+
"0"
|
| 16 |
+
],
|
| 17 |
+
"program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
|
| 18 |
+
"codePath": "train/train_deep_wandb.py",
|
| 19 |
+
"git": {
|
| 20 |
+
"remote": "git@hf.co:Yaning1001/Impossible_llm.git",
|
| 21 |
+
"commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
|
| 22 |
+
},
|
| 23 |
+
"email": "yaning1001@gmail.com",
|
| 24 |
+
"root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
|
| 25 |
+
"host": "mms-large-2",
|
| 26 |
+
"username": "chunhui",
|
| 27 |
+
"executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
|
| 28 |
+
"codePathLocal": "train_deep_wandb.py",
|
| 29 |
+
"cpu_count": 32,
|
| 30 |
+
"cpu_count_logical": 64,
|
| 31 |
+
"gpu": "NVIDIA RTX A6000",
|
| 32 |
+
"gpu_count": 8,
|
| 33 |
+
"disk": {
|
| 34 |
+
"/": {
|
| 35 |
+
"total": "1888559353856",
|
| 36 |
+
"used": "1754801557504"
|
| 37 |
+
}
|
| 38 |
+
},
|
| 39 |
+
"memory": {
|
| 40 |
+
"total": "202617098240"
|
| 41 |
+
},
|
| 42 |
+
"cpu": {
|
| 43 |
+
"count": 32,
|
| 44 |
+
"countLogical": 64
|
| 45 |
+
},
|
| 46 |
+
"gpu_nvidia": [
|
| 47 |
+
{
|
| 48 |
+
"name": "NVIDIA RTX A6000",
|
| 49 |
+
"memoryTotal": "51527024640",
|
| 50 |
+
"cudaCores": 10752,
|
| 51 |
+
"architecture": "Ampere"
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"name": "NVIDIA RTX A6000",
|
| 55 |
+
"memoryTotal": "51527024640",
|
| 56 |
+
"cudaCores": 10752,
|
| 57 |
+
"architecture": "Ampere"
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"name": "NVIDIA RTX A6000",
|
| 61 |
+
"memoryTotal": "51527024640",
|
| 62 |
+
"cudaCores": 10752,
|
| 63 |
+
"architecture": "Ampere"
|
| 64 |
+
},
|
| 65 |
+
{
|
| 66 |
+
"name": "NVIDIA RTX A6000",
|
| 67 |
+
"memoryTotal": "51527024640",
|
| 68 |
+
"cudaCores": 10752,
|
| 69 |
+
"architecture": "Ampere"
|
| 70 |
+
},
|
| 71 |
+
{
|
| 72 |
+
"name": "NVIDIA RTX A6000",
|
| 73 |
+
"memoryTotal": "51527024640",
|
| 74 |
+
"cudaCores": 10752,
|
| 75 |
+
"architecture": "Ampere"
|
| 76 |
+
},
|
| 77 |
+
{
|
| 78 |
+
"name": "NVIDIA RTX A6000",
|
| 79 |
+
"memoryTotal": "51527024640",
|
| 80 |
+
"cudaCores": 10752,
|
| 81 |
+
"architecture": "Ampere"
|
| 82 |
+
},
|
| 83 |
+
{
|
| 84 |
+
"name": "NVIDIA RTX A6000",
|
| 85 |
+
"memoryTotal": "51527024640",
|
| 86 |
+
"cudaCores": 10752,
|
| 87 |
+
"architecture": "Ampere"
|
| 88 |
+
},
|
| 89 |
+
{
|
| 90 |
+
"name": "NVIDIA RTX A6000",
|
| 91 |
+
"memoryTotal": "51527024640",
|
| 92 |
+
"cudaCores": 10752,
|
| 93 |
+
"architecture": "Ampere"
|
| 94 |
+
}
|
| 95 |
+
],
|
| 96 |
+
"cudaVersion": "11.8"
|
| 97 |
+
}
|
wandb/run-20241101_200517-iopieyi0/files/wandb-summary.json
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
{"_wandb":{"runtime":8}}
|
wandb/run-20241101_200517-iopieyi0/logs/debug-internal.log
ADDED
|
@@ -0,0 +1,11 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"time":"2024-11-01T20:05:17.143962046-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
|
| 2 |
+
{"time":"2024-11-01T20:05:17.143982966-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241101_200517-iopieyi0/logs/debug-core.log"}
|
| 3 |
+
{"time":"2024-11-01T20:05:17.249463211-04:00","level":"INFO","msg":"created new stream","id":"iopieyi0"}
|
| 4 |
+
{"time":"2024-11-01T20:05:17.249485021-04:00","level":"INFO","msg":"stream: started","id":"iopieyi0"}
|
| 5 |
+
{"time":"2024-11-01T20:05:17.249553251-04:00","level":"INFO","msg":"sender: started","stream_id":"iopieyi0"}
|
| 6 |
+
{"time":"2024-11-01T20:05:17.249556141-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"iopieyi0"}}
|
| 7 |
+
{"time":"2024-11-01T20:05:17.249534441-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"iopieyi0"}}
|
| 8 |
+
{"time":"2024-11-01T20:05:17.488885363-04:00","level":"INFO","msg":"Starting system monitor"}
|
| 9 |
+
{"time":"2024-11-01T20:05:25.263236061-04:00","level":"INFO","msg":"stream: closing","id":"iopieyi0"}
|
| 10 |
+
{"time":"2024-11-01T20:05:25.263331811-04:00","level":"INFO","msg":"Stopping system monitor"}
|
| 11 |
+
{"time":"2024-11-01T20:05:25.264008616-04:00","level":"INFO","msg":"Stopped system monitor"}
|
wandb/run-20241101_201927-8tmqrwpx/logs/debug-internal.log
ADDED
|
@@ -0,0 +1,16 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"time":"2024-11-01T20:19:27.015302603-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
|
| 2 |
+
{"time":"2024-11-01T20:19:27.015321923-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241101_201927-8tmqrwpx/logs/debug-core.log"}
|
| 3 |
+
{"time":"2024-11-01T20:19:27.121571158-04:00","level":"INFO","msg":"created new stream","id":"8tmqrwpx"}
|
| 4 |
+
{"time":"2024-11-01T20:19:27.121598178-04:00","level":"INFO","msg":"stream: started","id":"8tmqrwpx"}
|
| 5 |
+
{"time":"2024-11-01T20:19:27.121628798-04:00","level":"INFO","msg":"sender: started","stream_id":"8tmqrwpx"}
|
| 6 |
+
{"time":"2024-11-01T20:19:27.121672758-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"8tmqrwpx"}}
|
| 7 |
+
{"time":"2024-11-01T20:19:27.121613458-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"8tmqrwpx"}}
|
| 8 |
+
{"time":"2024-11-01T20:19:27.355884937-04:00","level":"INFO","msg":"Starting system monitor"}
|
| 9 |
+
{"time":"2024-11-01T20:20:25.212793119-04:00","level":"INFO","msg":"stream: closing","id":"8tmqrwpx"}
|
| 10 |
+
{"time":"2024-11-01T20:20:25.212864349-04:00","level":"INFO","msg":"Stopping system monitor"}
|
| 11 |
+
{"time":"2024-11-01T20:20:25.214203517-04:00","level":"INFO","msg":"Stopped system monitor"}
|
| 12 |
+
{"time":"2024-11-01T20:20:25.558844402-04:00","level":"INFO","msg":"fileTransfer: Close: file transfer manager closed"}
|
| 13 |
+
{"time":"2024-11-01T20:20:25.688669501-04:00","level":"INFO","msg":"handler: closed","stream_id":{"value":"8tmqrwpx"}}
|
| 14 |
+
{"time":"2024-11-01T20:20:25.688704032-04:00","level":"INFO","msg":"writer: Close: closed","stream_id":{"value":"8tmqrwpx"}}
|
| 15 |
+
{"time":"2024-11-01T20:20:25.688738252-04:00","level":"INFO","msg":"sender: closed","stream_id":"8tmqrwpx"}
|
| 16 |
+
{"time":"2024-11-01T20:20:25.688758612-04:00","level":"INFO","msg":"stream: closed","id":"8tmqrwpx"}
|
wandb/run-20241105_160059-czoj7ear/files/wandb-summary.json
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
{"_wandb":{"runtime":5}}
|
wandb/run-20241105_160059-czoj7ear/logs/debug-internal.log
ADDED
|
@@ -0,0 +1,17 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"time":"2024-11-05T16:00:59.421616162-05:00","level":"INFO","msg":"using version","core version":"0.18.5"}
|
| 2 |
+
{"time":"2024-11-05T16:00:59.421628512-05:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241105_160059-czoj7ear/logs/debug-core.log"}
|
| 3 |
+
{"time":"2024-11-05T16:01:04.448895853-05:00","level":"INFO","msg":"created new stream","id":"czoj7ear"}
|
| 4 |
+
{"time":"2024-11-05T16:01:04.448941013-05:00","level":"INFO","msg":"stream: started","id":"czoj7ear"}
|
| 5 |
+
{"time":"2024-11-05T16:01:04.449026033-05:00","level":"INFO","msg":"sender: started","stream_id":"czoj7ear"}
|
| 6 |
+
{"time":"2024-11-05T16:01:04.448975523-05:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"czoj7ear"}}
|
| 7 |
+
{"time":"2024-11-05T16:01:04.449027943-05:00","level":"INFO","msg":"handler: started","stream_id":{"value":"czoj7ear"}}
|
| 8 |
+
{"time":"2024-11-05T16:01:04.684712499-05:00","level":"INFO","msg":"Starting system monitor"}
|
| 9 |
+
{"time":"2024-11-05T16:01:04.794485285-05:00","level":"INFO","msg":"stream: closing","id":"czoj7ear"}
|
| 10 |
+
{"time":"2024-11-05T16:01:04.794512205-05:00","level":"INFO","msg":"Stopping system monitor"}
|
| 11 |
+
{"time":"2024-11-05T16:01:04.794592695-05:00","level":"INFO","msg":"Stopped system monitor"}
|
| 12 |
+
{"time":"2024-11-05T16:01:05.093812959-05:00","level":"ERROR","msg":"sender: sendDefer: failed to build job artifact","error":"failed to write data to file: write /tmp/tmpfile-3701223180: no space left on device"}
|
| 13 |
+
{"time":"2024-11-05T16:01:05.340839422-05:00","level":"INFO","msg":"fileTransfer: Close: file transfer manager closed"}
|
| 14 |
+
{"time":"2024-11-05T16:01:05.518421834-05:00","level":"INFO","msg":"handler: closed","stream_id":{"value":"czoj7ear"}}
|
| 15 |
+
{"time":"2024-11-05T16:01:05.518467794-05:00","level":"INFO","msg":"writer: Close: closed","stream_id":{"value":"czoj7ear"}}
|
| 16 |
+
{"time":"2024-11-05T16:01:05.518480744-05:00","level":"INFO","msg":"sender: closed","stream_id":"czoj7ear"}
|
| 17 |
+
{"time":"2024-11-05T16:01:05.518572734-05:00","level":"INFO","msg":"stream: closed","id":"czoj7ear"}
|
wandb/run-20241105_160059-czoj7ear/run-czoj7ear.wandb
ADDED
|
Binary file (3.78 kB). View file
|
|
|
wandb/run-20241105_161832-c18fx9uc/files/requirements.txt
ADDED
|
@@ -0,0 +1,147 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
funcsigs==1.0.2
|
| 2 |
+
sentry-sdk==2.17.0
|
| 3 |
+
multiprocess==0.70.16
|
| 4 |
+
numpy==1.26.2
|
| 5 |
+
pluralizer==1.2.0
|
| 6 |
+
debugpy==1.6.7
|
| 7 |
+
nvidia-cudnn-cu11==8.5.0.96
|
| 8 |
+
deepspeed==0.15.2
|
| 9 |
+
data==0.4
|
| 10 |
+
pandas==2.1.3
|
| 11 |
+
tomli==2.0.1
|
| 12 |
+
charset-normalizer==3.3.2
|
| 13 |
+
attrs==24.2.0
|
| 14 |
+
aiosignal==1.3.1
|
| 15 |
+
fsspec==2023.10.0
|
| 16 |
+
nvidia-cusparse-cu11==11.7.4.91
|
| 17 |
+
zipp==3.12.0
|
| 18 |
+
mypy-extensions==1.0.0
|
| 19 |
+
datasets==3.0.1
|
| 20 |
+
joblib==1.3.2
|
| 21 |
+
hjson==3.1.0
|
| 22 |
+
traitlets==5.7.1
|
| 23 |
+
stack-data==0.6.0
|
| 24 |
+
transformers==4.45.1
|
| 25 |
+
sympy==1.11.1
|
| 26 |
+
Pygments==2.15.0
|
| 27 |
+
docker-pycreds==0.4.0
|
| 28 |
+
dill==0.3.8
|
| 29 |
+
wheel==0.44.0
|
| 30 |
+
prompt-toolkit==3.0.30
|
| 31 |
+
parso==0.8.3
|
| 32 |
+
ipykernel==6.23.1
|
| 33 |
+
pyarrow==17.0.0
|
| 34 |
+
certifi==2023.11.17
|
| 35 |
+
nvidia-cufft-cu11==10.9.0.58
|
| 36 |
+
six==1.16.0
|
| 37 |
+
pydantic==2.9.2
|
| 38 |
+
click==8.1.7
|
| 39 |
+
nest-asyncio==1.5.6
|
| 40 |
+
gmpy2==2.1.0
|
| 41 |
+
matplotlib==3.8.2
|
| 42 |
+
scipy==1.11.4
|
| 43 |
+
typing_extensions==4.12.2
|
| 44 |
+
statsmodels==0.14.0
|
| 45 |
+
huggingface-hub==0.25.0
|
| 46 |
+
frozenlist==1.4.1
|
| 47 |
+
gpustat==1.1.1
|
| 48 |
+
nvidia-nvtx-cu11==11.7.91
|
| 49 |
+
safetensors==0.4.5
|
| 50 |
+
stanza==1.9.2
|
| 51 |
+
decorator==5.1.1
|
| 52 |
+
seaborn==0.13.0
|
| 53 |
+
sentencepiece==0.2.0
|
| 54 |
+
PyYAML==6.0.1
|
| 55 |
+
black==24.8.0
|
| 56 |
+
protobuf==4.25.1
|
| 57 |
+
pickleshare==0.7.5
|
| 58 |
+
peft==0.13.0
|
| 59 |
+
triton==2.0.0
|
| 60 |
+
nvidia-cuda-runtime-cu11==11.7.99
|
| 61 |
+
Jinja2==3.1.2
|
| 62 |
+
nvidia-cusolver-cu11==11.4.0.1
|
| 63 |
+
executing==1.2.0
|
| 64 |
+
jupyter_client==8.1.0
|
| 65 |
+
pluggy==1.3.0
|
| 66 |
+
cmake==3.30.3
|
| 67 |
+
pytz==2023.3.post1
|
| 68 |
+
aiohappyeyeballs==2.4.2
|
| 69 |
+
kiwisolver==1.4.5
|
| 70 |
+
py-cpuinfo==9.0.0
|
| 71 |
+
Pillow==10.1.0
|
| 72 |
+
ptyprocess==0.7.0
|
| 73 |
+
importlib_resources==6.4.5
|
| 74 |
+
GitPython==3.1.43
|
| 75 |
+
importlib-metadata==6.0.0
|
| 76 |
+
iniconfig==2.0.0
|
| 77 |
+
scikit-learn==1.3.2
|
| 78 |
+
exceptiongroup==1.1.0
|
| 79 |
+
networkx==2.8.6
|
| 80 |
+
accelerate==1.0.0
|
| 81 |
+
nltk==3.8.1
|
| 82 |
+
shutilwhich==1.1.0
|
| 83 |
+
fonttools==4.45.1
|
| 84 |
+
future==0.18.3
|
| 85 |
+
aiohttp==3.10.6
|
| 86 |
+
wcwidth==0.2.5
|
| 87 |
+
idna==3.6
|
| 88 |
+
filelock==3.12.2
|
| 89 |
+
pathspec==0.12.1
|
| 90 |
+
jupyter_core==5.1.0
|
| 91 |
+
lit==18.1.8
|
| 92 |
+
nvidia-curand-cu11==10.2.10.91
|
| 93 |
+
nvidia-cublas-cu11==11.10.3.66
|
| 94 |
+
nvidia-ml-py==12.560.30
|
| 95 |
+
msgpack==1.1.0
|
| 96 |
+
python-dateutil==2.8.2
|
| 97 |
+
blessed==1.20.0
|
| 98 |
+
packaging==23.0
|
| 99 |
+
gitdb==4.0.11
|
| 100 |
+
yarl==1.13.0
|
| 101 |
+
emoji==2.8.0
|
| 102 |
+
tzdata==2023.3
|
| 103 |
+
cycler==0.12.1
|
| 104 |
+
tornado==6.2
|
| 105 |
+
backcall==0.2.0
|
| 106 |
+
plotnine==0.12.4
|
| 107 |
+
ninja==1.11.1.1
|
| 108 |
+
latex==0.7.0
|
| 109 |
+
wandb==0.18.5
|
| 110 |
+
setproctitle==1.3.3
|
| 111 |
+
threadpoolctl==3.2.0
|
| 112 |
+
requests==2.32.3
|
| 113 |
+
pyparsing==3.1.1
|
| 114 |
+
smmap==5.0.1
|
| 115 |
+
pyzmq==23.0.0
|
| 116 |
+
async-timeout==4.0.3
|
| 117 |
+
annotated-types==0.7.0
|
| 118 |
+
matplotlib-inline==0.1.6
|
| 119 |
+
latexcodec==1.0.0
|
| 120 |
+
ipython==8.0.0
|
| 121 |
+
patsy==0.5.3
|
| 122 |
+
contourpy==1.2.0
|
| 123 |
+
multidict==6.1.0
|
| 124 |
+
mizani==0.9.3
|
| 125 |
+
urllib3==2.1.0
|
| 126 |
+
tokenizers==0.20.0
|
| 127 |
+
MarkupSafe==2.1.2
|
| 128 |
+
pip==24.2
|
| 129 |
+
pexpect==4.8.0
|
| 130 |
+
tqdm==4.66.5
|
| 131 |
+
jedi==0.18.2
|
| 132 |
+
pydantic_core==2.23.4
|
| 133 |
+
tempdir==0.7.1
|
| 134 |
+
mpmath==1.2.1
|
| 135 |
+
setuptools==72.1.0
|
| 136 |
+
pytest==7.4.3
|
| 137 |
+
pure-eval==0.2.2
|
| 138 |
+
psutil==5.9.1
|
| 139 |
+
comm==0.1.2
|
| 140 |
+
nvidia-cuda-cupti-cu11==11.7.101
|
| 141 |
+
nvidia-cuda-nvrtc-cu11==11.7.99
|
| 142 |
+
regex==2023.10.3
|
| 143 |
+
platformdirs==2.5.2
|
| 144 |
+
asttokens==2.2.1
|
| 145 |
+
torch==2.0.0
|
| 146 |
+
nvidia-nccl-cu11==2.14.3
|
| 147 |
+
xxhash==3.5.0
|
wandb/run-20241105_162858-6py0unak/files/config.yaml
ADDED
|
@@ -0,0 +1,49 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
_wandb:
|
| 2 |
+
value:
|
| 3 |
+
cli_version: 0.18.5
|
| 4 |
+
m: []
|
| 5 |
+
python_version: 3.9.19
|
| 6 |
+
t:
|
| 7 |
+
"1":
|
| 8 |
+
- 1
|
| 9 |
+
- 5
|
| 10 |
+
- 11
|
| 11 |
+
- 49
|
| 12 |
+
- 51
|
| 13 |
+
- 53
|
| 14 |
+
- 55
|
| 15 |
+
- 71
|
| 16 |
+
- 98
|
| 17 |
+
"2":
|
| 18 |
+
- 1
|
| 19 |
+
- 5
|
| 20 |
+
- 11
|
| 21 |
+
- 49
|
| 22 |
+
- 51
|
| 23 |
+
- 53
|
| 24 |
+
- 55
|
| 25 |
+
- 71
|
| 26 |
+
- 98
|
| 27 |
+
"3":
|
| 28 |
+
- 13
|
| 29 |
+
- 23
|
| 30 |
+
- 55
|
| 31 |
+
"4": 3.9.19
|
| 32 |
+
"5": 0.18.5
|
| 33 |
+
"6": 4.45.1
|
| 34 |
+
"8":
|
| 35 |
+
- 5
|
| 36 |
+
"12": 0.18.5
|
| 37 |
+
"13": linux-x86_64
|
| 38 |
+
batch_size:
|
| 39 |
+
value: 3
|
| 40 |
+
epoch:
|
| 41 |
+
value: 3
|
| 42 |
+
lr:
|
| 43 |
+
value: 5e-06
|
| 44 |
+
perturbation:
|
| 45 |
+
value: shuffle_deterministic57
|
| 46 |
+
seed:
|
| 47 |
+
value: 0
|
| 48 |
+
train_set:
|
| 49 |
+
value: 10M
|
wandb/run-20241105_162858-6py0unak/files/output.log
ADDED
|
@@ -0,0 +1,34 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Downloading shards: 0%| | 0/2 [00:00<?, ?it/s]
|
| 2 |
+
Error in sys.excepthook:
|
| 3 |
+
Traceback (most recent call last):
|
| 4 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/wandb/sdk/lib/exit_hooks.py", line 41, in exc_handler
|
| 5 |
+
def exc_handler(
|
| 6 |
+
KeyboardInterrupt
|
| 7 |
+
|
| 8 |
+
Original exception was:
|
| 9 |
+
Traceback (most recent call last):
|
| 10 |
+
File "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py", line 174, in <module>
|
| 11 |
+
model = AutoModelForCausalLM.from_pretrained(model_name,
|
| 12 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/models/auto/auto_factory.py", line 564, in from_pretrained
|
| 13 |
+
return model_class.from_pretrained(
|
| 14 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/modeling_utils.py", line 3769, in from_pretrained
|
| 15 |
+
resolved_archive_file, sharded_metadata = get_checkpoint_shard_files(
|
| 16 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/utils/hub.py", line 1098, in get_checkpoint_shard_files
|
| 17 |
+
cached_filename = cached_file(
|
| 18 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/utils/hub.py", line 403, in cached_file
|
| 19 |
+
resolved_file = hf_hub_download(
|
| 20 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/utils/_deprecation.py", line 101, in inner_f
|
| 21 |
+
return f(*args, **kwargs)
|
| 22 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/utils/_validators.py", line 114, in _inner_fn
|
| 23 |
+
return fn(*args, **kwargs)
|
| 24 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/file_download.py", line 1232, in hf_hub_download
|
| 25 |
+
return _hf_hub_download_to_cache_dir(
|
| 26 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/file_download.py", line 1380, in _hf_hub_download_to_cache_dir
|
| 27 |
+
with WeakFileLock(lock_path):
|
| 28 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/contextlib.py", line 119, in __enter__
|
| 29 |
+
return next(self.gen)
|
| 30 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/utils/_fixes.py", line 98, in WeakFileLock
|
| 31 |
+
lock.acquire()
|
| 32 |
+
File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/filelock/_api.py", line 225, in acquire
|
| 33 |
+
time.sleep(poll_interval)
|
| 34 |
+
KeyboardInterrupt
|
wandb/run-20241105_162858-6py0unak/files/requirements.txt
ADDED
|
@@ -0,0 +1,147 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
funcsigs==1.0.2
|
| 2 |
+
sentry-sdk==2.17.0
|
| 3 |
+
multiprocess==0.70.16
|
| 4 |
+
numpy==1.26.2
|
| 5 |
+
pluralizer==1.2.0
|
| 6 |
+
debugpy==1.6.7
|
| 7 |
+
nvidia-cudnn-cu11==8.5.0.96
|
| 8 |
+
deepspeed==0.15.2
|
| 9 |
+
data==0.4
|
| 10 |
+
pandas==2.1.3
|
| 11 |
+
tomli==2.0.1
|
| 12 |
+
charset-normalizer==3.3.2
|
| 13 |
+
attrs==24.2.0
|
| 14 |
+
aiosignal==1.3.1
|
| 15 |
+
fsspec==2023.10.0
|
| 16 |
+
nvidia-cusparse-cu11==11.7.4.91
|
| 17 |
+
zipp==3.12.0
|
| 18 |
+
mypy-extensions==1.0.0
|
| 19 |
+
datasets==3.0.1
|
| 20 |
+
joblib==1.3.2
|
| 21 |
+
hjson==3.1.0
|
| 22 |
+
traitlets==5.7.1
|
| 23 |
+
stack-data==0.6.0
|
| 24 |
+
transformers==4.45.1
|
| 25 |
+
sympy==1.11.1
|
| 26 |
+
Pygments==2.15.0
|
| 27 |
+
docker-pycreds==0.4.0
|
| 28 |
+
dill==0.3.8
|
| 29 |
+
wheel==0.44.0
|
| 30 |
+
prompt-toolkit==3.0.30
|
| 31 |
+
parso==0.8.3
|
| 32 |
+
ipykernel==6.23.1
|
| 33 |
+
pyarrow==17.0.0
|
| 34 |
+
certifi==2023.11.17
|
| 35 |
+
nvidia-cufft-cu11==10.9.0.58
|
| 36 |
+
six==1.16.0
|
| 37 |
+
pydantic==2.9.2
|
| 38 |
+
click==8.1.7
|
| 39 |
+
nest-asyncio==1.5.6
|
| 40 |
+
gmpy2==2.1.0
|
| 41 |
+
matplotlib==3.8.2
|
| 42 |
+
scipy==1.11.4
|
| 43 |
+
typing_extensions==4.12.2
|
| 44 |
+
statsmodels==0.14.0
|
| 45 |
+
huggingface-hub==0.25.0
|
| 46 |
+
frozenlist==1.4.1
|
| 47 |
+
gpustat==1.1.1
|
| 48 |
+
nvidia-nvtx-cu11==11.7.91
|
| 49 |
+
safetensors==0.4.5
|
| 50 |
+
stanza==1.9.2
|
| 51 |
+
decorator==5.1.1
|
| 52 |
+
seaborn==0.13.0
|
| 53 |
+
sentencepiece==0.2.0
|
| 54 |
+
PyYAML==6.0.1
|
| 55 |
+
black==24.8.0
|
| 56 |
+
protobuf==4.25.1
|
| 57 |
+
pickleshare==0.7.5
|
| 58 |
+
peft==0.13.0
|
| 59 |
+
triton==2.0.0
|
| 60 |
+
nvidia-cuda-runtime-cu11==11.7.99
|
| 61 |
+
Jinja2==3.1.2
|
| 62 |
+
nvidia-cusolver-cu11==11.4.0.1
|
| 63 |
+
executing==1.2.0
|
| 64 |
+
jupyter_client==8.1.0
|
| 65 |
+
pluggy==1.3.0
|
| 66 |
+
cmake==3.30.3
|
| 67 |
+
pytz==2023.3.post1
|
| 68 |
+
aiohappyeyeballs==2.4.2
|
| 69 |
+
kiwisolver==1.4.5
|
| 70 |
+
py-cpuinfo==9.0.0
|
| 71 |
+
Pillow==10.1.0
|
| 72 |
+
ptyprocess==0.7.0
|
| 73 |
+
importlib_resources==6.4.5
|
| 74 |
+
GitPython==3.1.43
|
| 75 |
+
importlib-metadata==6.0.0
|
| 76 |
+
iniconfig==2.0.0
|
| 77 |
+
scikit-learn==1.3.2
|
| 78 |
+
exceptiongroup==1.1.0
|
| 79 |
+
networkx==2.8.6
|
| 80 |
+
accelerate==1.0.0
|
| 81 |
+
nltk==3.8.1
|
| 82 |
+
shutilwhich==1.1.0
|
| 83 |
+
fonttools==4.45.1
|
| 84 |
+
future==0.18.3
|
| 85 |
+
aiohttp==3.10.6
|
| 86 |
+
wcwidth==0.2.5
|
| 87 |
+
idna==3.6
|
| 88 |
+
filelock==3.12.2
|
| 89 |
+
pathspec==0.12.1
|
| 90 |
+
jupyter_core==5.1.0
|
| 91 |
+
lit==18.1.8
|
| 92 |
+
nvidia-curand-cu11==10.2.10.91
|
| 93 |
+
nvidia-cublas-cu11==11.10.3.66
|
| 94 |
+
nvidia-ml-py==12.560.30
|
| 95 |
+
msgpack==1.1.0
|
| 96 |
+
python-dateutil==2.8.2
|
| 97 |
+
blessed==1.20.0
|
| 98 |
+
packaging==23.0
|
| 99 |
+
gitdb==4.0.11
|
| 100 |
+
yarl==1.13.0
|
| 101 |
+
emoji==2.8.0
|
| 102 |
+
tzdata==2023.3
|
| 103 |
+
cycler==0.12.1
|
| 104 |
+
tornado==6.2
|
| 105 |
+
backcall==0.2.0
|
| 106 |
+
plotnine==0.12.4
|
| 107 |
+
ninja==1.11.1.1
|
| 108 |
+
latex==0.7.0
|
| 109 |
+
wandb==0.18.5
|
| 110 |
+
setproctitle==1.3.3
|
| 111 |
+
threadpoolctl==3.2.0
|
| 112 |
+
requests==2.32.3
|
| 113 |
+
pyparsing==3.1.1
|
| 114 |
+
smmap==5.0.1
|
| 115 |
+
pyzmq==23.0.0
|
| 116 |
+
async-timeout==4.0.3
|
| 117 |
+
annotated-types==0.7.0
|
| 118 |
+
matplotlib-inline==0.1.6
|
| 119 |
+
latexcodec==1.0.0
|
| 120 |
+
ipython==8.0.0
|
| 121 |
+
patsy==0.5.3
|
| 122 |
+
contourpy==1.2.0
|
| 123 |
+
multidict==6.1.0
|
| 124 |
+
mizani==0.9.3
|
| 125 |
+
urllib3==2.1.0
|
| 126 |
+
tokenizers==0.20.0
|
| 127 |
+
MarkupSafe==2.1.2
|
| 128 |
+
pip==24.2
|
| 129 |
+
pexpect==4.8.0
|
| 130 |
+
tqdm==4.66.5
|
| 131 |
+
jedi==0.18.2
|
| 132 |
+
pydantic_core==2.23.4
|
| 133 |
+
tempdir==0.7.1
|
| 134 |
+
mpmath==1.2.1
|
| 135 |
+
setuptools==72.1.0
|
| 136 |
+
pytest==7.4.3
|
| 137 |
+
pure-eval==0.2.2
|
| 138 |
+
psutil==5.9.1
|
| 139 |
+
comm==0.1.2
|
| 140 |
+
nvidia-cuda-cupti-cu11==11.7.101
|
| 141 |
+
nvidia-cuda-nvrtc-cu11==11.7.99
|
| 142 |
+
regex==2023.10.3
|
| 143 |
+
platformdirs==2.5.2
|
| 144 |
+
asttokens==2.2.1
|
| 145 |
+
torch==2.0.0
|
| 146 |
+
nvidia-nccl-cu11==2.14.3
|
| 147 |
+
xxhash==3.5.0
|
wandb/run-20241105_162858-6py0unak/files/wandb-metadata.json
ADDED
|
@@ -0,0 +1,97 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
|
| 3 |
+
"python": "3.9.19",
|
| 4 |
+
"startedAt": "2024-11-05T21:28:58.840904Z",
|
| 5 |
+
"args": [
|
| 6 |
+
"--perturbation",
|
| 7 |
+
"shuffle_deterministic57",
|
| 8 |
+
"--train_set",
|
| 9 |
+
"10M",
|
| 10 |
+
"--batch_size",
|
| 11 |
+
"3",
|
| 12 |
+
"--epoch",
|
| 13 |
+
"3",
|
| 14 |
+
"--seed",
|
| 15 |
+
"0"
|
| 16 |
+
],
|
| 17 |
+
"program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
|
| 18 |
+
"codePath": "train/train_deep_wandb.py",
|
| 19 |
+
"git": {
|
| 20 |
+
"remote": "git@hf.co:Yaning1001/Impossible_llm.git",
|
| 21 |
+
"commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
|
| 22 |
+
},
|
| 23 |
+
"email": "yaning1001@gmail.com",
|
| 24 |
+
"root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
|
| 25 |
+
"host": "mms-large-2",
|
| 26 |
+
"username": "chunhui",
|
| 27 |
+
"executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
|
| 28 |
+
"codePathLocal": "train_deep_wandb.py",
|
| 29 |
+
"cpu_count": 32,
|
| 30 |
+
"cpu_count_logical": 64,
|
| 31 |
+
"gpu": "NVIDIA RTX A6000",
|
| 32 |
+
"gpu_count": 8,
|
| 33 |
+
"disk": {
|
| 34 |
+
"/": {
|
| 35 |
+
"total": "1888559353856",
|
| 36 |
+
"used": "1785811787776"
|
| 37 |
+
}
|
| 38 |
+
},
|
| 39 |
+
"memory": {
|
| 40 |
+
"total": "202617098240"
|
| 41 |
+
},
|
| 42 |
+
"cpu": {
|
| 43 |
+
"count": 32,
|
| 44 |
+
"countLogical": 64
|
| 45 |
+
},
|
| 46 |
+
"gpu_nvidia": [
|
| 47 |
+
{
|
| 48 |
+
"name": "NVIDIA RTX A6000",
|
| 49 |
+
"memoryTotal": "51527024640",
|
| 50 |
+
"cudaCores": 10752,
|
| 51 |
+
"architecture": "Ampere"
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"name": "NVIDIA RTX A6000",
|
| 55 |
+
"memoryTotal": "51527024640",
|
| 56 |
+
"cudaCores": 10752,
|
| 57 |
+
"architecture": "Ampere"
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"name": "NVIDIA RTX A6000",
|
| 61 |
+
"memoryTotal": "51527024640",
|
| 62 |
+
"cudaCores": 10752,
|
| 63 |
+
"architecture": "Ampere"
|
| 64 |
+
},
|
| 65 |
+
{
|
| 66 |
+
"name": "NVIDIA RTX A6000",
|
| 67 |
+
"memoryTotal": "51527024640",
|
| 68 |
+
"cudaCores": 10752,
|
| 69 |
+
"architecture": "Ampere"
|
| 70 |
+
},
|
| 71 |
+
{
|
| 72 |
+
"name": "NVIDIA RTX A6000",
|
| 73 |
+
"memoryTotal": "51527024640",
|
| 74 |
+
"cudaCores": 10752,
|
| 75 |
+
"architecture": "Ampere"
|
| 76 |
+
},
|
| 77 |
+
{
|
| 78 |
+
"name": "NVIDIA RTX A6000",
|
| 79 |
+
"memoryTotal": "51527024640",
|
| 80 |
+
"cudaCores": 10752,
|
| 81 |
+
"architecture": "Ampere"
|
| 82 |
+
},
|
| 83 |
+
{
|
| 84 |
+
"name": "NVIDIA RTX A6000",
|
| 85 |
+
"memoryTotal": "51527024640",
|
| 86 |
+
"cudaCores": 10752,
|
| 87 |
+
"architecture": "Ampere"
|
| 88 |
+
},
|
| 89 |
+
{
|
| 90 |
+
"name": "NVIDIA RTX A6000",
|
| 91 |
+
"memoryTotal": "51527024640",
|
| 92 |
+
"cudaCores": 10752,
|
| 93 |
+
"architecture": "Ampere"
|
| 94 |
+
}
|
| 95 |
+
],
|
| 96 |
+
"cudaVersion": "11.8"
|
| 97 |
+
}
|
wandb/run-20241105_162858-6py0unak/files/wandb-summary.json
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
{"_wandb":{"runtime":23}}
|
wandb/run-20241105_162858-6py0unak/logs/debug-internal.log
ADDED
|
@@ -0,0 +1,12 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"time":"2024-11-05T16:28:58.843383096-05:00","level":"INFO","msg":"using version","core version":"0.18.5"}
|
| 2 |
+
{"time":"2024-11-05T16:28:58.843400236-05:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241105_162858-6py0unak/logs/debug-core.log"}
|
| 3 |
+
{"time":"2024-11-05T16:28:58.951956143-05:00","level":"INFO","msg":"created new stream","id":"6py0unak"}
|
| 4 |
+
{"time":"2024-11-05T16:28:58.952002873-05:00","level":"INFO","msg":"stream: started","id":"6py0unak"}
|
| 5 |
+
{"time":"2024-11-05T16:28:58.952045833-05:00","level":"INFO","msg":"sender: started","stream_id":"6py0unak"}
|
| 6 |
+
{"time":"2024-11-05T16:28:58.952018493-05:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"6py0unak"}}
|
| 7 |
+
{"time":"2024-11-05T16:28:58.952045693-05:00","level":"INFO","msg":"handler: started","stream_id":{"value":"6py0unak"}}
|
| 8 |
+
{"time":"2024-11-05T16:28:59.159891406-05:00","level":"INFO","msg":"Starting system monitor"}
|
| 9 |
+
{"time":"2024-11-05T16:29:22.458107091-05:00","level":"INFO","msg":"stream: closing","id":"6py0unak"}
|
| 10 |
+
{"time":"2024-11-05T16:29:22.458143501-05:00","level":"INFO","msg":"Stopping system monitor"}
|
| 11 |
+
{"time":"2024-11-05T16:29:22.458735014-05:00","level":"INFO","msg":"Stopped system monitor"}
|
| 12 |
+
{"time":"2024-11-05T16:29:22.68105964-05:00","level":"INFO","msg":"api: retrying HTTP error","status":503,"url":"https://api.wandb.ai/graphql"}
|
wandb/run-20241105_162858-6py0unak/logs/debug.log
ADDED
|
@@ -0,0 +1,27 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
2024-11-05 16:28:58,838 INFO MainThread:1778375 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
|
| 2 |
+
2024-11-05 16:28:58,838 INFO MainThread:1778375 [wandb_setup.py:_flush():79] Configure stats pid to 1778375
|
| 3 |
+
2024-11-05 16:28:58,838 INFO MainThread:1778375 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
|
| 4 |
+
2024-11-05 16:28:58,838 INFO MainThread:1778375 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
|
| 5 |
+
2024-11-05 16:28:58,838 INFO MainThread:1778375 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
|
| 6 |
+
2024-11-05 16:28:58,838 INFO MainThread:1778375 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
|
| 7 |
+
2024-11-05 16:28:58,838 INFO MainThread:1778375 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
|
| 8 |
+
2024-11-05 16:28:58,838 INFO MainThread:1778375 [wandb_setup.py:_flush():79] Applying login settings: {}
|
| 9 |
+
2024-11-05 16:28:58,839 INFO MainThread:1778375 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241105_162858-6py0unak/logs/debug.log
|
| 10 |
+
2024-11-05 16:28:58,839 INFO MainThread:1778375 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241105_162858-6py0unak/logs/debug-internal.log
|
| 11 |
+
2024-11-05 16:28:58,839 INFO MainThread:1778375 [wandb_init.py:init():621] calling init triggers
|
| 12 |
+
2024-11-05 16:28:58,839 INFO MainThread:1778375 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
|
| 13 |
+
config: {}
|
| 14 |
+
2024-11-05 16:28:58,839 INFO MainThread:1778375 [wandb_init.py:init():671] starting backend
|
| 15 |
+
2024-11-05 16:28:58,839 INFO MainThread:1778375 [wandb_init.py:init():675] sending inform_init request
|
| 16 |
+
2024-11-05 16:28:58,840 INFO MainThread:1778375 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
|
| 17 |
+
2024-11-05 16:28:58,840 INFO MainThread:1778375 [wandb_init.py:init():688] backend started and connected
|
| 18 |
+
2024-11-05 16:28:58,844 INFO MainThread:1778375 [wandb_init.py:init():783] updated telemetry
|
| 19 |
+
2024-11-05 16:28:58,865 INFO MainThread:1778375 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
|
| 20 |
+
2024-11-05 16:28:59,157 INFO MainThread:1778375 [wandb_init.py:init():867] starting run threads in backend
|
| 21 |
+
2024-11-05 16:28:59,247 INFO MainThread:1778375 [wandb_run.py:_console_start():2463] atexit reg
|
| 22 |
+
2024-11-05 16:28:59,247 INFO MainThread:1778375 [wandb_run.py:_redirect():2311] redirect: wrap_raw
|
| 23 |
+
2024-11-05 16:28:59,247 INFO MainThread:1778375 [wandb_run.py:_redirect():2376] Wrapping output streams.
|
| 24 |
+
2024-11-05 16:28:59,247 INFO MainThread:1778375 [wandb_run.py:_redirect():2401] Redirects installed.
|
| 25 |
+
2024-11-05 16:28:59,249 INFO MainThread:1778375 [wandb_init.py:init():911] run started, returning control to user process
|
| 26 |
+
2024-11-05 16:28:59,249 INFO MainThread:1778375 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'shuffle_deterministic57', 'train_set': '10M', 'batch_size': 3, 'epoch': 3, 'seed': 0, 'lr': 5e-06}
|
| 27 |
+
2024-11-05 16:29:22,458 WARNING MsgRouterThr:1778375 [router.py:message_loop():77] message_loop has been closed
|
wandb/run-20241105_163001-vvohahtj/run-vvohahtj.wandb
ADDED
|
File without changes
|