Yaning1001 commited on
Commit
907dcdb
·
verified ·
1 Parent(s): b87502c

Add files using upload-large-folder tool

Browse files
This view is limited to 50 files because it contains too many changes.   See raw diff
Files changed (50) hide show
  1. .gitattributes +5 -0
  2. wandb/run-20241030_013339-pahr4hk1/files/output.log +21 -0
  3. wandb/run-20241030_013339-pahr4hk1/files/requirements.txt +147 -0
  4. wandb/run-20241030_013339-pahr4hk1/files/wandb-metadata.json +97 -0
  5. wandb/run-20241030_013339-pahr4hk1/logs/debug-internal.log +8 -0
  6. wandb/run-20241030_013339-pahr4hk1/logs/debug.log +26 -0
  7. wandb/run-20241030_112852-av3r7rx8/run-av3r7rx8.wandb +3 -0
  8. wandb/run-20241030_112852-mfvd6tgw/files/config.yaml +48 -0
  9. wandb/run-20241030_112852-mfvd6tgw/files/output.log +17 -0
  10. wandb/run-20241030_112852-mfvd6tgw/files/wandb-metadata.json +97 -0
  11. wandb/run-20241030_112852-mfvd6tgw/files/wandb-summary.json +1 -0
  12. wandb/run-20241030_112852-mfvd6tgw/logs/debug-internal.log +108 -0
  13. wandb/run-20241030_112852-mfvd6tgw/logs/debug.log +33 -0
  14. wandb/run-20241030_233740-np98q8en/run-np98q8en.wandb +3 -0
  15. wandb/run-20241031_000839-acpkxm8c/files/output.log +13 -0
  16. wandb/run-20241031_000839-acpkxm8c/files/requirements.txt +147 -0
  17. wandb/run-20241031_000839-acpkxm8c/files/wandb-metadata.json +97 -0
  18. wandb/run-20241031_000839-acpkxm8c/logs/debug-internal.log +8 -0
  19. wandb/run-20241031_000839-acpkxm8c/logs/debug.log +26 -0
  20. wandb/run-20241031_000839-acpkxm8c/run-acpkxm8c.wandb +0 -0
  21. wandb/run-20241101_012733-3tsgnm2p/run-3tsgnm2p.wandb +3 -0
  22. wandb/run-20241101_200502-7hem25r3/files/output.log +1 -0
  23. wandb/run-20241101_200502-7hem25r3/files/requirements.txt +147 -0
  24. wandb/run-20241101_200502-7hem25r3/files/wandb-metadata.json +97 -0
  25. wandb/run-20241101_200502-7hem25r3/logs/debug-internal.log +8 -0
  26. wandb/run-20241101_200502-7hem25r3/logs/debug.log +26 -0
  27. wandb/run-20241101_200502-7hem25r3/run-7hem25r3.wandb +0 -0
  28. wandb/run-20241101_200517-vzv5zg2q/files/config.yaml +49 -0
  29. wandb/run-20241101_200517-vzv5zg2q/files/output.log +34 -0
  30. wandb/run-20241101_200517-vzv5zg2q/files/requirements.txt +147 -0
  31. wandb/run-20241101_200517-vzv5zg2q/files/wandb-metadata.json +97 -0
  32. wandb/run-20241101_200517-vzv5zg2q/files/wandb-summary.json +1 -0
  33. wandb/run-20241101_200517-vzv5zg2q/logs/debug-internal.log +11 -0
  34. wandb/run-20241101_200517-vzv5zg2q/logs/debug.log +27 -0
  35. wandb/run-20241101_200517-vzv5zg2q/run-vzv5zg2q.wandb +0 -0
  36. wandb/run-20241101_201630-e5gt2fir/files/config.yaml +49 -0
  37. wandb/run-20241101_201630-e5gt2fir/files/output.log +12 -0
  38. wandb/run-20241101_201630-e5gt2fir/files/wandb-summary.json +1 -0
  39. wandb/run-20241101_201630-e5gt2fir/logs/debug-internal.log +16 -0
  40. wandb/run-20241101_201630-e5gt2fir/logs/debug.log +27 -0
  41. wandb/run-20241101_202058-hjyig8so/run-hjyig8so.wandb +3 -0
  42. wandb/run-20241105_160652-v9udw9ab/files/config.yaml +49 -0
  43. wandb/run-20241105_160652-v9udw9ab/files/output.log +8 -0
  44. wandb/run-20241105_160652-v9udw9ab/files/requirements.txt +147 -0
  45. wandb/run-20241105_160652-v9udw9ab/files/wandb-metadata.json +97 -0
  46. wandb/run-20241105_160652-v9udw9ab/files/wandb-summary.json +1 -0
  47. wandb/run-20241105_160652-v9udw9ab/logs/debug-internal.log +17 -0
  48. wandb/run-20241105_160652-v9udw9ab/logs/debug.log +27 -0
  49. wandb/run-20241105_160652-v9udw9ab/run-v9udw9ab.wandb +0 -0
  50. wandb/run-20241105_162858-hqnfirxi/files/config.yaml +49 -0
.gitattributes CHANGED
@@ -92,3 +92,8 @@ wandb/run-20241106_232725-f16bcfrx/run-f16bcfrx.wandb filter=lfs diff=lfs merge=
92
  wandb/run-20241105_163244-59l4qxgx/run-59l4qxgx.wandb filter=lfs diff=lfs merge=lfs -text
93
  wandb/run-20241101_202058-ptl7coag/run-ptl7coag.wandb filter=lfs diff=lfs merge=lfs -text
94
  wandb/run-20241129_235322-bxqdruiw/run-bxqdruiw.wandb filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
92
  wandb/run-20241105_163244-59l4qxgx/run-59l4qxgx.wandb filter=lfs diff=lfs merge=lfs -text
93
  wandb/run-20241101_202058-ptl7coag/run-ptl7coag.wandb filter=lfs diff=lfs merge=lfs -text
94
  wandb/run-20241129_235322-bxqdruiw/run-bxqdruiw.wandb filter=lfs diff=lfs merge=lfs -text
95
+ wandb/run-20241030_233740-np98q8en/run-np98q8en.wandb filter=lfs diff=lfs merge=lfs -text
96
+ wandb/run-20241101_202058-hjyig8so/run-hjyig8so.wandb filter=lfs diff=lfs merge=lfs -text
97
+ wandb/run-20241101_012733-3tsgnm2p/run-3tsgnm2p.wandb filter=lfs diff=lfs merge=lfs -text
98
+ wandb/run-20241030_112852-av3r7rx8/run-av3r7rx8.wandb filter=lfs diff=lfs merge=lfs -text
99
+ wandb/run-20241115_125218-rrve0rbk/run-rrve0rbk.wandb filter=lfs diff=lfs merge=lfs -text
wandb/run-20241030_013339-pahr4hk1/files/output.log ADDED
@@ -0,0 +1,21 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ config.json: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 844/844 [00:00<00:00, 279kB/s]
2
+ model.safetensors.index.json: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████| 20.9k/20.9k [00:00<00:00, 18.9MB/s]
3
+ model-00001-of-00002.safetensors: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 4.97G/4.97G [01:57<00:00, 42.1MB/s]
4
+ model-00002-of-00002.safetensors: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 1.46G/1.46G [00:34<00:00, 42.5MB/s]
5
+ Downloading shards: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 2/2 [02:32<00:00, 76.19s/it]
6
+ Loading checkpoint shards: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 2/2 [00:04<00:00, 2.10s/it]
7
+ generation_config.json: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 185/185 [00:00<00:00, 74.7kB/s]
8
+ Map: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 17519/17519 [00:57<00:00, 302.93 examples/s]
9
+ Map: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 18140/18140 [00:57<00:00, 316.26 examples/s]
10
+ tokenized_valid: Dataset({
11
+ features: ['input_ids', 'attention_mask'],
12
+ num_rows: 600
13
+ })
14
+ /mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/training_args.py:1545: FutureWarning: `evaluation_strategy` is deprecated and will be removed in version 4.46 of 🤗 Transformers. Use `eval_strategy` instead
15
+ warnings.warn(
16
+ [2024-10-30 01:38:13,749] [INFO] [real_accelerator.py:219:get_accelerator] Setting ds_accelerator to cuda (auto detect)
17
+ [2024-10-30 01:38:21,729] [INFO] [comm.py:652:init_distributed] cdb=None
18
+ Installed CUDA version 11.8 does not match the version torch was compiled with 11.7 but since the APIs are compatible, accepting this combination
19
+ Using /home/chunhui/.cache/torch_extensions/py39_cu117 as PyTorch extensions root...
20
+ Loading extension module cpu_adam...
21
+ Time to load cpu_adam op: 4.509259939193726 seconds
wandb/run-20241030_013339-pahr4hk1/files/requirements.txt ADDED
@@ -0,0 +1,147 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ funcsigs==1.0.2
2
+ sentry-sdk==2.17.0
3
+ multiprocess==0.70.16
4
+ numpy==1.26.2
5
+ pluralizer==1.2.0
6
+ debugpy==1.6.7
7
+ nvidia-cudnn-cu11==8.5.0.96
8
+ deepspeed==0.15.2
9
+ data==0.4
10
+ pandas==2.1.3
11
+ tomli==2.0.1
12
+ charset-normalizer==3.3.2
13
+ attrs==24.2.0
14
+ aiosignal==1.3.1
15
+ fsspec==2023.10.0
16
+ nvidia-cusparse-cu11==11.7.4.91
17
+ zipp==3.12.0
18
+ mypy-extensions==1.0.0
19
+ datasets==3.0.1
20
+ joblib==1.3.2
21
+ hjson==3.1.0
22
+ traitlets==5.7.1
23
+ stack-data==0.6.0
24
+ transformers==4.45.1
25
+ sympy==1.11.1
26
+ Pygments==2.15.0
27
+ docker-pycreds==0.4.0
28
+ dill==0.3.8
29
+ wheel==0.44.0
30
+ prompt-toolkit==3.0.30
31
+ parso==0.8.3
32
+ ipykernel==6.23.1
33
+ pyarrow==17.0.0
34
+ certifi==2023.11.17
35
+ nvidia-cufft-cu11==10.9.0.58
36
+ six==1.16.0
37
+ pydantic==2.9.2
38
+ click==8.1.7
39
+ nest-asyncio==1.5.6
40
+ gmpy2==2.1.0
41
+ matplotlib==3.8.2
42
+ scipy==1.11.4
43
+ typing_extensions==4.12.2
44
+ statsmodels==0.14.0
45
+ huggingface-hub==0.25.0
46
+ frozenlist==1.4.1
47
+ gpustat==1.1.1
48
+ nvidia-nvtx-cu11==11.7.91
49
+ safetensors==0.4.5
50
+ stanza==1.9.2
51
+ decorator==5.1.1
52
+ seaborn==0.13.0
53
+ sentencepiece==0.2.0
54
+ PyYAML==6.0.1
55
+ black==24.8.0
56
+ protobuf==4.25.1
57
+ pickleshare==0.7.5
58
+ peft==0.13.0
59
+ triton==2.0.0
60
+ nvidia-cuda-runtime-cu11==11.7.99
61
+ Jinja2==3.1.2
62
+ nvidia-cusolver-cu11==11.4.0.1
63
+ executing==1.2.0
64
+ jupyter_client==8.1.0
65
+ pluggy==1.3.0
66
+ cmake==3.30.3
67
+ pytz==2023.3.post1
68
+ aiohappyeyeballs==2.4.2
69
+ kiwisolver==1.4.5
70
+ py-cpuinfo==9.0.0
71
+ Pillow==10.1.0
72
+ ptyprocess==0.7.0
73
+ importlib_resources==6.4.5
74
+ GitPython==3.1.43
75
+ importlib-metadata==6.0.0
76
+ iniconfig==2.0.0
77
+ scikit-learn==1.3.2
78
+ exceptiongroup==1.1.0
79
+ networkx==2.8.6
80
+ accelerate==1.0.0
81
+ nltk==3.8.1
82
+ shutilwhich==1.1.0
83
+ fonttools==4.45.1
84
+ future==0.18.3
85
+ aiohttp==3.10.6
86
+ wcwidth==0.2.5
87
+ idna==3.6
88
+ filelock==3.12.2
89
+ pathspec==0.12.1
90
+ jupyter_core==5.1.0
91
+ lit==18.1.8
92
+ nvidia-curand-cu11==10.2.10.91
93
+ nvidia-cublas-cu11==11.10.3.66
94
+ nvidia-ml-py==12.560.30
95
+ msgpack==1.1.0
96
+ python-dateutil==2.8.2
97
+ blessed==1.20.0
98
+ packaging==23.0
99
+ gitdb==4.0.11
100
+ yarl==1.13.0
101
+ emoji==2.8.0
102
+ tzdata==2023.3
103
+ cycler==0.12.1
104
+ tornado==6.2
105
+ backcall==0.2.0
106
+ plotnine==0.12.4
107
+ ninja==1.11.1.1
108
+ latex==0.7.0
109
+ wandb==0.18.5
110
+ setproctitle==1.3.3
111
+ threadpoolctl==3.2.0
112
+ requests==2.32.3
113
+ pyparsing==3.1.1
114
+ smmap==5.0.1
115
+ pyzmq==23.0.0
116
+ async-timeout==4.0.3
117
+ annotated-types==0.7.0
118
+ matplotlib-inline==0.1.6
119
+ latexcodec==1.0.0
120
+ ipython==8.0.0
121
+ patsy==0.5.3
122
+ contourpy==1.2.0
123
+ multidict==6.1.0
124
+ mizani==0.9.3
125
+ urllib3==2.1.0
126
+ tokenizers==0.20.0
127
+ MarkupSafe==2.1.2
128
+ pip==24.2
129
+ pexpect==4.8.0
130
+ tqdm==4.66.5
131
+ jedi==0.18.2
132
+ pydantic_core==2.23.4
133
+ tempdir==0.7.1
134
+ mpmath==1.2.1
135
+ setuptools==72.1.0
136
+ pytest==7.4.3
137
+ pure-eval==0.2.2
138
+ psutil==5.9.1
139
+ comm==0.1.2
140
+ nvidia-cuda-cupti-cu11==11.7.101
141
+ nvidia-cuda-nvrtc-cu11==11.7.99
142
+ regex==2023.10.3
143
+ platformdirs==2.5.2
144
+ asttokens==2.2.1
145
+ torch==2.0.0
146
+ nvidia-nccl-cu11==2.14.3
147
+ xxhash==3.5.0
wandb/run-20241030_013339-pahr4hk1/files/wandb-metadata.json ADDED
@@ -0,0 +1,97 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
3
+ "python": "3.9.19",
4
+ "startedAt": "2024-10-30T05:33:39.715818Z",
5
+ "args": [
6
+ "--perturbation",
7
+ "reverse_full",
8
+ "--train_set",
9
+ "10M",
10
+ "--batch_size",
11
+ "3",
12
+ "--epoch",
13
+ "7",
14
+ "--seed",
15
+ "0"
16
+ ],
17
+ "program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
18
+ "codePath": "train/train_deep_wandb.py",
19
+ "git": {
20
+ "remote": "git@hf.co:Yaning1001/Impossible_llm.git",
21
+ "commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
22
+ },
23
+ "email": "yaning1001@gmail.com",
24
+ "root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
25
+ "host": "mms-large-2",
26
+ "username": "chunhui",
27
+ "executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
28
+ "codePathLocal": "train_deep_wandb.py",
29
+ "cpu_count": 32,
30
+ "cpu_count_logical": 64,
31
+ "gpu": "NVIDIA RTX A6000",
32
+ "gpu_count": 8,
33
+ "disk": {
34
+ "/": {
35
+ "total": "1888559353856",
36
+ "used": "1710081847296"
37
+ }
38
+ },
39
+ "memory": {
40
+ "total": "202617098240"
41
+ },
42
+ "cpu": {
43
+ "count": 32,
44
+ "countLogical": 64
45
+ },
46
+ "gpu_nvidia": [
47
+ {
48
+ "name": "NVIDIA RTX A6000",
49
+ "memoryTotal": "51527024640",
50
+ "cudaCores": 10752,
51
+ "architecture": "Ampere"
52
+ },
53
+ {
54
+ "name": "NVIDIA RTX A6000",
55
+ "memoryTotal": "51527024640",
56
+ "cudaCores": 10752,
57
+ "architecture": "Ampere"
58
+ },
59
+ {
60
+ "name": "NVIDIA RTX A6000",
61
+ "memoryTotal": "51527024640",
62
+ "cudaCores": 10752,
63
+ "architecture": "Ampere"
64
+ },
65
+ {
66
+ "name": "NVIDIA RTX A6000",
67
+ "memoryTotal": "51527024640",
68
+ "cudaCores": 10752,
69
+ "architecture": "Ampere"
70
+ },
71
+ {
72
+ "name": "NVIDIA RTX A6000",
73
+ "memoryTotal": "51527024640",
74
+ "cudaCores": 10752,
75
+ "architecture": "Ampere"
76
+ },
77
+ {
78
+ "name": "NVIDIA RTX A6000",
79
+ "memoryTotal": "51527024640",
80
+ "cudaCores": 10752,
81
+ "architecture": "Ampere"
82
+ },
83
+ {
84
+ "name": "NVIDIA RTX A6000",
85
+ "memoryTotal": "51527024640",
86
+ "cudaCores": 10752,
87
+ "architecture": "Ampere"
88
+ },
89
+ {
90
+ "name": "NVIDIA RTX A6000",
91
+ "memoryTotal": "51527024640",
92
+ "cudaCores": 10752,
93
+ "architecture": "Ampere"
94
+ }
95
+ ],
96
+ "cudaVersion": "11.8"
97
+ }
wandb/run-20241030_013339-pahr4hk1/logs/debug-internal.log ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {"time":"2024-10-30T01:33:39.718219018-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
2
+ {"time":"2024-10-30T01:33:39.718234308-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_013339-pahr4hk1/logs/debug-core.log"}
3
+ {"time":"2024-10-30T01:33:39.825323399-04:00","level":"INFO","msg":"created new stream","id":"pahr4hk1"}
4
+ {"time":"2024-10-30T01:33:39.825359099-04:00","level":"INFO","msg":"stream: started","id":"pahr4hk1"}
5
+ {"time":"2024-10-30T01:33:39.825400679-04:00","level":"INFO","msg":"sender: started","stream_id":"pahr4hk1"}
6
+ {"time":"2024-10-30T01:33:39.825385829-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"pahr4hk1"}}
7
+ {"time":"2024-10-30T01:33:39.825383059-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"pahr4hk1"}}
8
+ {"time":"2024-10-30T01:33:39.990823292-04:00","level":"INFO","msg":"Starting system monitor"}
wandb/run-20241030_013339-pahr4hk1/logs/debug.log ADDED
@@ -0,0 +1,26 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 2024-10-30 01:33:39,713 INFO MainThread:337258 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
2
+ 2024-10-30 01:33:39,714 INFO MainThread:337258 [wandb_setup.py:_flush():79] Configure stats pid to 337258
3
+ 2024-10-30 01:33:39,714 INFO MainThread:337258 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
4
+ 2024-10-30 01:33:39,714 INFO MainThread:337258 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
5
+ 2024-10-30 01:33:39,714 INFO MainThread:337258 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
6
+ 2024-10-30 01:33:39,714 INFO MainThread:337258 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
7
+ 2024-10-30 01:33:39,714 INFO MainThread:337258 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
8
+ 2024-10-30 01:33:39,714 INFO MainThread:337258 [wandb_setup.py:_flush():79] Applying login settings: {}
9
+ 2024-10-30 01:33:39,714 INFO MainThread:337258 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_013339-pahr4hk1/logs/debug.log
10
+ 2024-10-30 01:33:39,714 INFO MainThread:337258 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_013339-pahr4hk1/logs/debug-internal.log
11
+ 2024-10-30 01:33:39,714 INFO MainThread:337258 [wandb_init.py:init():621] calling init triggers
12
+ 2024-10-30 01:33:39,714 INFO MainThread:337258 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
13
+ config: {}
14
+ 2024-10-30 01:33:39,714 INFO MainThread:337258 [wandb_init.py:init():671] starting backend
15
+ 2024-10-30 01:33:39,714 INFO MainThread:337258 [wandb_init.py:init():675] sending inform_init request
16
+ 2024-10-30 01:33:39,715 INFO MainThread:337258 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
17
+ 2024-10-30 01:33:39,715 INFO MainThread:337258 [wandb_init.py:init():688] backend started and connected
18
+ 2024-10-30 01:33:39,718 INFO MainThread:337258 [wandb_init.py:init():783] updated telemetry
19
+ 2024-10-30 01:33:39,748 INFO MainThread:337258 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
20
+ 2024-10-30 01:33:39,986 INFO MainThread:337258 [wandb_init.py:init():867] starting run threads in backend
21
+ 2024-10-30 01:33:40,076 INFO MainThread:337258 [wandb_run.py:_console_start():2463] atexit reg
22
+ 2024-10-30 01:33:40,076 INFO MainThread:337258 [wandb_run.py:_redirect():2311] redirect: wrap_raw
23
+ 2024-10-30 01:33:40,076 INFO MainThread:337258 [wandb_run.py:_redirect():2376] Wrapping output streams.
24
+ 2024-10-30 01:33:40,076 INFO MainThread:337258 [wandb_run.py:_redirect():2401] Redirects installed.
25
+ 2024-10-30 01:33:40,077 INFO MainThread:337258 [wandb_init.py:init():911] run started, returning control to user process
26
+ 2024-10-30 01:33:40,078 INFO MainThread:337258 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'reverse_full', 'train_set': '10M', 'batch_size': 3, 'epoch': 7, 'seed': 0}
wandb/run-20241030_112852-av3r7rx8/run-av3r7rx8.wandb ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:82826fe5130b74ba5b11ab4c02aeb2ce26500de1060a76c495fae27c1f928f5f
3
+ size 14185176
wandb/run-20241030_112852-mfvd6tgw/files/config.yaml ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ _wandb:
2
+ value:
3
+ cli_version: 0.18.5
4
+ m: []
5
+ python_version: 3.9.19
6
+ t:
7
+ "1":
8
+ - 1
9
+ - 5
10
+ - 11
11
+ - 49
12
+ - 51
13
+ - 53
14
+ - 55
15
+ - 71
16
+ - 98
17
+ "2":
18
+ - 1
19
+ - 5
20
+ - 11
21
+ - 49
22
+ - 51
23
+ - 53
24
+ - 55
25
+ - 71
26
+ - 98
27
+ "3":
28
+ - 2
29
+ - 13
30
+ - 23
31
+ - 55
32
+ "4": 3.9.19
33
+ "5": 0.18.5
34
+ "6": 4.45.1
35
+ "8":
36
+ - 5
37
+ "12": 0.18.5
38
+ "13": linux-x86_64
39
+ batch_size:
40
+ value: 3
41
+ epoch:
42
+ value: 3
43
+ perturbation:
44
+ value: reverse_control
45
+ seed:
46
+ value: 0
47
+ train_set:
48
+ value: 10M
wandb/run-20241030_112852-mfvd6tgw/files/output.log ADDED
@@ -0,0 +1,17 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Downloading shards: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 2/2 [02:08<00:00, 64.23s/it]
2
+ Loading checkpoint shards: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 2/2 [00:04<00:00, 2.16s/it]
3
+ generation_config.json: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 185/185 [00:00<00:00, 57.7kB/s]
4
+ Map: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 18140/18140 [00:49<00:00, 367.87 examples/s]
5
+ tokenized_valid: Dataset({
6
+ features: ['input_ids', 'attention_mask'],
7
+ num_rows: 600
8
+ })
9
+ /mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/training_args.py:1545: FutureWarning: `evaluation_strategy` is deprecated and will be removed in version 4.46 of 🤗 Transformers. Use `eval_strategy` instead
10
+ warnings.warn(
11
+ [2024-10-30 11:31:57,082] [INFO] [real_accelerator.py:219:get_accelerator] Setting ds_accelerator to cuda (auto detect)
12
+ [2024-10-30 11:32:05,240] [INFO] [comm.py:652:init_distributed] cdb=None
13
+ Installed CUDA version 11.8 does not match the version torch was compiled with 11.7 but since the APIs are compatible, accepting this combination
14
+ Using /home/chunhui/.cache/torch_extensions/py39_cu117 as PyTorch extensions root...
15
+ Loading extension module cpu_adam...
16
+ Time to load cpu_adam op: 4.680240869522095 seconds
17
+ wandb: WARNING Fatal error while uploading data. Some run data will not be synced, but it will still be written to disk. Use `wandb sync` at the end of the run to try uploading.
wandb/run-20241030_112852-mfvd6tgw/files/wandb-metadata.json ADDED
@@ -0,0 +1,97 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
3
+ "python": "3.9.19",
4
+ "startedAt": "2024-10-30T15:28:52.868095Z",
5
+ "args": [
6
+ "--perturbation",
7
+ "reverse_control",
8
+ "--train_set",
9
+ "10M",
10
+ "--batch_size",
11
+ "3",
12
+ "--epoch",
13
+ "3",
14
+ "--seed",
15
+ "0"
16
+ ],
17
+ "program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
18
+ "codePath": "train/train_deep_wandb.py",
19
+ "git": {
20
+ "remote": "git@hf.co:Yaning1001/Impossible_llm.git",
21
+ "commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
22
+ },
23
+ "email": "yaning1001@gmail.com",
24
+ "root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
25
+ "host": "mms-large-2",
26
+ "username": "chunhui",
27
+ "executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
28
+ "codePathLocal": "train_deep_wandb.py",
29
+ "cpu_count": 32,
30
+ "cpu_count_logical": 64,
31
+ "gpu": "NVIDIA RTX A6000",
32
+ "gpu_count": 8,
33
+ "disk": {
34
+ "/": {
35
+ "total": "1888559353856",
36
+ "used": "1710831611904"
37
+ }
38
+ },
39
+ "memory": {
40
+ "total": "202617098240"
41
+ },
42
+ "cpu": {
43
+ "count": 32,
44
+ "countLogical": 64
45
+ },
46
+ "gpu_nvidia": [
47
+ {
48
+ "name": "NVIDIA RTX A6000",
49
+ "memoryTotal": "51527024640",
50
+ "cudaCores": 10752,
51
+ "architecture": "Ampere"
52
+ },
53
+ {
54
+ "name": "NVIDIA RTX A6000",
55
+ "memoryTotal": "51527024640",
56
+ "cudaCores": 10752,
57
+ "architecture": "Ampere"
58
+ },
59
+ {
60
+ "name": "NVIDIA RTX A6000",
61
+ "memoryTotal": "51527024640",
62
+ "cudaCores": 10752,
63
+ "architecture": "Ampere"
64
+ },
65
+ {
66
+ "name": "NVIDIA RTX A6000",
67
+ "memoryTotal": "51527024640",
68
+ "cudaCores": 10752,
69
+ "architecture": "Ampere"
70
+ },
71
+ {
72
+ "name": "NVIDIA RTX A6000",
73
+ "memoryTotal": "51527024640",
74
+ "cudaCores": 10752,
75
+ "architecture": "Ampere"
76
+ },
77
+ {
78
+ "name": "NVIDIA RTX A6000",
79
+ "memoryTotal": "51527024640",
80
+ "cudaCores": 10752,
81
+ "architecture": "Ampere"
82
+ },
83
+ {
84
+ "name": "NVIDIA RTX A6000",
85
+ "memoryTotal": "51527024640",
86
+ "cudaCores": 10752,
87
+ "architecture": "Ampere"
88
+ },
89
+ {
90
+ "name": "NVIDIA RTX A6000",
91
+ "memoryTotal": "51527024640",
92
+ "cudaCores": 10752,
93
+ "architecture": "Ampere"
94
+ }
95
+ ],
96
+ "cudaVersion": "11.8"
97
+ }
wandb/run-20241030_112852-mfvd6tgw/files/wandb-summary.json ADDED
@@ -0,0 +1 @@
 
 
1
+ {"_wandb":{"runtime":23503}}
wandb/run-20241030_112852-mfvd6tgw/logs/debug-internal.log ADDED
@@ -0,0 +1,108 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {"time":"2024-10-30T11:28:52.871579904-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
2
+ {"time":"2024-10-30T11:28:52.871600584-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_112852-mfvd6tgw/logs/debug-core.log"}
3
+ {"time":"2024-10-30T11:28:52.978841939-04:00","level":"INFO","msg":"created new stream","id":"mfvd6tgw"}
4
+ {"time":"2024-10-30T11:28:52.978902739-04:00","level":"INFO","msg":"stream: started","id":"mfvd6tgw"}
5
+ {"time":"2024-10-30T11:28:52.978926819-04:00","level":"INFO","msg":"sender: started","stream_id":"mfvd6tgw"}
6
+ {"time":"2024-10-30T11:28:52.97893628-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"mfvd6tgw"}}
7
+ {"time":"2024-10-30T11:28:52.978909109-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"mfvd6tgw"}}
8
+ {"time":"2024-10-30T11:28:53.155627531-04:00","level":"INFO","msg":"Starting system monitor"}
9
+ {"time":"2024-10-30T14:03:08.474581851-04:00","level":"ERROR","msg":"HTTP error","status":404,"method":"POST","url":"https://api.wandb.ai/files/yaning1001-dartmouth-college/impossible_llm_reverse/mfvd6tgw/file_stream"}
10
+ {"time":"2024-10-30T14:03:08.487108589-04:00","level":"ERROR+4","msg":"filestream: fatal error: filestream: failed to upload: 404 Not Found path=files/yaning1001-dartmouth-college/impossible_llm_reverse/mfvd6tgw/file_stream: {\"error\":\"run impossible_llm_reverse/mfvd6tgw not found while streaming file\"}"}
11
+ {"time":"2024-10-30T15:29:42.382518231-04:00","level":"INFO","msg":"api: retrying error","error":"Post \"https://api.wandb.ai/graphql\": net/http: request canceled (Client.Timeout exceeded while awaiting headers)"}
12
+ {"time":"2024-10-30T18:00:35.96830223-04:00","level":"INFO","msg":"Stopping system monitor"}
13
+ {"time":"2024-10-30T18:00:35.984486769-04:00","level":"INFO","msg":"Stopped system monitor"}
14
+ {"time":"2024-10-30T18:00:36.012170229-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
15
+ {"time":"2024-10-30T18:00:36.969090118-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":1.01417858,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
16
+ {"time":"2024-10-30T18:00:38.26382741-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
17
+ {"time":"2024-10-30T18:00:42.839562556-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
18
+ {"time":"2024-10-30T18:00:52.373625385-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
19
+ {"time":"2024-10-30T18:01:08.798689764-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
20
+ {"time":"2024-10-30T18:01:37.002153025-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":61.047236517,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
21
+ {"time":"2024-10-30T18:01:48.282746193-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
22
+ {"time":"2024-10-30T18:02:37.021865642-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":121.066947764,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
23
+ {"time":"2024-10-30T18:02:48.335092366-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
24
+ {"time":"2024-10-30T18:03:37.04344863-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":181.088531722,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
25
+ {"time":"2024-10-30T18:03:48.393221692-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
26
+ {"time":"2024-10-30T18:04:37.067486978-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":241.11257044,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
27
+ {"time":"2024-10-30T18:04:48.448890729-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
28
+ {"time":"2024-10-30T18:05:37.08852491-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":301.133606772,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
29
+ {"time":"2024-10-30T18:05:48.502838457-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
30
+ {"time":"2024-10-30T18:06:37.117047582-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":361.162133104,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
31
+ {"time":"2024-10-30T18:06:48.554665436-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
32
+ {"time":"2024-10-30T18:07:37.139331089-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":421.184416131,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
33
+ {"time":"2024-10-30T18:07:48.61386156-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
34
+ {"time":"2024-10-30T18:08:37.166847552-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":481.211928004,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
35
+ {"time":"2024-10-30T18:08:48.666190809-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
36
+ {"time":"2024-10-30T18:09:37.194996856-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":541.240081368,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
37
+ {"time":"2024-10-30T18:09:48.721077294-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
38
+ {"time":"2024-10-30T18:10:35.955456557-04:00","level":"WARN","msg":"sender: taking a long time","seconds":600.000428857,"work":"WorkRecord(*service_go_proto.Record_Telemetry); Control(connection_id:\"127.0.0.1:35326\")"}
39
+ {"time":"2024-10-30T18:10:37.216651416-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":601.261725058,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
40
+ {"time":"2024-10-30T18:10:48.773966229-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
41
+ {"time":"2024-10-30T18:11:37.244172578-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":661.28925569,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
42
+ {"time":"2024-10-30T18:11:48.824895668-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
43
+ {"time":"2024-10-30T18:12:37.298658553-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":721.343740985,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
44
+ {"time":"2024-10-30T18:12:48.922946274-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
45
+ {"time":"2024-10-30T18:13:37.319508977-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":781.364593549,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
46
+ {"time":"2024-10-30T18:13:48.976226497-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
47
+ {"time":"2024-10-30T18:14:37.34947864-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":841.394560982,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
48
+ {"time":"2024-10-30T18:14:49.026538556-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
49
+ {"time":"2024-10-30T18:15:37.371409433-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":901.416490594,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
50
+ {"time":"2024-10-30T18:15:49.077219977-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
51
+ {"time":"2024-10-30T18:16:37.394261418-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":961.43934326,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
52
+ {"time":"2024-10-30T18:16:49.129011563-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
53
+ {"time":"2024-10-30T18:16:49.129142475-04:00","level":"ERROR","msg":"sender: sendConfig:","error":"api: failed sending: POST https://api.wandb.ai/graphql giving up after 21 attempt(s)"}
54
+ {"time":"2024-10-30T18:16:49.129310967-04:00","level":"INFO","msg":"sender: succeeded after taking longer than expected","seconds":973.174329898,"work":"WorkRecord(*service_go_proto.Record_Telemetry); Control(connection_id:\"127.0.0.1:35326\")"}
55
+ {"time":"2024-10-30T18:16:49.185435119-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
56
+ {"time":"2024-10-30T18:16:49.231279214-04:00","level":"ERROR","msg":"HTTP error","status":404,"method":"POST","url":"https://api.wandb.ai/graphql"}
57
+ {"time":"2024-10-30T18:16:49.231344485-04:00","level":"ERROR","msg":"runfiles: CreateRunFiles returned error: returned error 404 Not Found: {\"errors\":[{\"message\":\"run impossible_llm_reverse/mfvd6tgw not found during createRunFiles\",\"path\":[\"createRunFiles\"]}],\"data\":{\"createRunFiles\":null}}"}
58
+ {"time":"2024-10-30T18:16:51.551546918-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
59
+ {"time":"2024-10-30T18:16:56.224687098-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
60
+ {"time":"2024-10-30T18:17:04.811694165-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
61
+ {"time":"2024-10-30T18:17:23.782790961-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
62
+ {"time":"2024-10-30T18:17:37.423013569-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":48.293350338,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
63
+ {"time":"2024-10-30T18:18:01.969943103-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
64
+ {"time":"2024-10-30T18:18:37.444499646-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":108.314835984,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
65
+ {"time":"2024-10-30T18:19:02.031162287-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
66
+ {"time":"2024-10-30T18:19:37.46831367-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":168.338651879,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
67
+ {"time":"2024-10-30T18:20:02.086803086-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
68
+ {"time":"2024-10-30T18:20:37.49484466-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":228.365184589,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
69
+ {"time":"2024-10-30T18:21:02.14808835-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
70
+ {"time":"2024-10-30T18:21:37.518777326-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":288.389114975,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
71
+ {"time":"2024-10-30T18:22:02.208668814-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
72
+ {"time":"2024-10-30T18:22:37.542083025-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":348.412418624,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
73
+ {"time":"2024-10-30T18:23:02.265589282-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
74
+ {"time":"2024-10-30T18:23:37.568249482-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":408.438587131,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
75
+ {"time":"2024-10-30T18:24:02.332185182-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
76
+ {"time":"2024-10-30T18:24:37.59016655-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":468.460504129,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
77
+ {"time":"2024-10-30T18:25:02.383191439-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
78
+ {"time":"2024-10-30T18:25:37.611919348-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":528.482255227,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
79
+ {"time":"2024-10-30T18:26:02.439712218-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
80
+ {"time":"2024-10-30T18:26:37.63751757-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":588.507844589,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
81
+ {"time":"2024-10-30T18:26:49.130687726-04:00","level":"WARN","msg":"sender: taking a long time","seconds":600.00065645,"work":"WorkRecord(*service_go_proto.Request_Defer); Control(local:true always_send:true)"}
82
+ {"time":"2024-10-30T18:27:02.49576742-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
83
+ {"time":"2024-10-30T18:27:37.658073199-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":648.528411148,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
84
+ {"time":"2024-10-30T18:28:02.556603286-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
85
+ {"time":"2024-10-30T18:28:37.678804315-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":708.549129014,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
86
+ {"time":"2024-10-30T18:29:02.616108339-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
87
+ {"time":"2024-10-30T18:29:37.705580303-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":768.575917741,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
88
+ {"time":"2024-10-30T18:30:02.669248598-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
89
+ {"time":"2024-10-30T18:30:37.73008217-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":828.600419649,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
90
+ {"time":"2024-10-30T18:31:02.724511612-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
91
+ {"time":"2024-10-30T18:31:37.753990664-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":888.624324932,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
92
+ {"time":"2024-10-30T18:32:02.784891698-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
93
+ {"time":"2024-10-30T18:32:37.778640619-04:00","level":"INFO","msg":"handler: operation stats","stats":{"operations":[{"desc":"updating run config","runtime_seconds":948.648979298,"error_status":"retrying HTTP 409 Conflict"}],"total_operations":1}}
94
+ {"time":"2024-10-30T18:33:02.873762515-04:00","level":"INFO","msg":"api: retrying HTTP error","status":409,"url":"https://api.wandb.ai/graphql"}
95
+ {"time":"2024-10-30T18:33:02.873850586-04:00","level":"ERROR","msg":"sender: sendConfig:","error":"api: failed sending: POST https://api.wandb.ai/graphql giving up after 21 attempt(s)"}
96
+ {"time":"2024-10-30T18:33:02.874229689-04:00","level":"INFO","msg":"sender: succeeded after taking longer than expected","seconds":973.744352185,"work":"WorkRecord(*service_go_proto.Request_Defer); Control(local:true always_send:true)"}
97
+ {"time":"2024-10-30T18:33:02.972769976-04:00","level":"ERROR","msg":"HTTP error","status":404,"method":"POST","url":"https://api.wandb.ai/graphql"}
98
+ {"time":"2024-10-30T18:33:02.972826436-04:00","level":"ERROR","msg":"runfiles: CreateRunFiles returned error: returned error 404 Not Found: {\"errors\":[{\"message\":\"run impossible_llm_reverse/mfvd6tgw not found during createRunFiles\",\"path\":[\"createRunFiles\"]}],\"data\":{\"createRunFiles\":null}}"}
99
+ {"time":"2024-10-30T18:33:03.110980965-04:00","level":"ERROR","msg":"HTTP error","status":404,"method":"POST","url":"https://api.wandb.ai/graphql"}
100
+ {"time":"2024-10-30T18:33:03.111192917-04:00","level":"ERROR","msg":"sender: failed to save job artifact: ArtifactSaver.createManifest: returned error 404 Not Found: {\"errors\":[{\"message\":\"failed to find run impossible_llm_reverse/mfvd6tgw\",\"path\":[\"createArtifactManifest\"]}],\"data\":{\"createArtifactManifest\":null}}"}
101
+ {"time":"2024-10-30T18:33:03.162250522-04:00","level":"INFO","msg":"fileTransfer: Close: file transfer manager closed"}
102
+ {"time":"2024-10-30T18:33:03.212467488-04:00","level":"ERROR","msg":"HTTP error","status":404,"method":"POST","url":"https://api.wandb.ai/graphql"}
103
+ {"time":"2024-10-30T18:33:03.212514949-04:00","level":"ERROR","msg":"runfiles: CreateRunFiles returned error: returned error 404 Not Found: {\"errors\":[{\"message\":\"run impossible_llm_reverse/mfvd6tgw not found during createRunFiles\",\"path\":[\"createRunFiles\"]}],\"data\":{\"createRunFiles\":null}}"}
104
+ {"time":"2024-10-30T18:33:04.172896443-04:00","level":"INFO","msg":"stream: closing","id":"mfvd6tgw"}
105
+ {"time":"2024-10-30T18:33:04.172936993-04:00","level":"INFO","msg":"handler: closed","stream_id":{"value":"mfvd6tgw"}}
106
+ {"time":"2024-10-30T18:33:04.172965223-04:00","level":"INFO","msg":"sender: closed","stream_id":"mfvd6tgw"}
107
+ {"time":"2024-10-30T18:33:04.172958023-04:00","level":"INFO","msg":"writer: Close: closed","stream_id":{"value":"mfvd6tgw"}}
108
+ {"time":"2024-10-30T18:33:04.173060174-04:00","level":"INFO","msg":"stream: closed","id":"mfvd6tgw"}
wandb/run-20241030_112852-mfvd6tgw/logs/debug.log ADDED
@@ -0,0 +1,33 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 2024-10-30 11:28:52,866 INFO MainThread:367766 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
2
+ 2024-10-30 11:28:52,866 INFO MainThread:367766 [wandb_setup.py:_flush():79] Configure stats pid to 367766
3
+ 2024-10-30 11:28:52,866 INFO MainThread:367766 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
4
+ 2024-10-30 11:28:52,866 INFO MainThread:367766 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
5
+ 2024-10-30 11:28:52,866 INFO MainThread:367766 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
6
+ 2024-10-30 11:28:52,866 INFO MainThread:367766 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
7
+ 2024-10-30 11:28:52,866 INFO MainThread:367766 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
8
+ 2024-10-30 11:28:52,866 INFO MainThread:367766 [wandb_setup.py:_flush():79] Applying login settings: {}
9
+ 2024-10-30 11:28:52,866 INFO MainThread:367766 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_112852-mfvd6tgw/logs/debug.log
10
+ 2024-10-30 11:28:52,866 INFO MainThread:367766 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241030_112852-mfvd6tgw/logs/debug-internal.log
11
+ 2024-10-30 11:28:52,866 INFO MainThread:367766 [wandb_init.py:init():621] calling init triggers
12
+ 2024-10-30 11:28:52,866 INFO MainThread:367766 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
13
+ config: {}
14
+ 2024-10-30 11:28:52,866 INFO MainThread:367766 [wandb_init.py:init():671] starting backend
15
+ 2024-10-30 11:28:52,866 INFO MainThread:367766 [wandb_init.py:init():675] sending inform_init request
16
+ 2024-10-30 11:28:52,867 INFO MainThread:367766 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
17
+ 2024-10-30 11:28:52,867 INFO MainThread:367766 [wandb_init.py:init():688] backend started and connected
18
+ 2024-10-30 11:28:52,871 INFO MainThread:367766 [wandb_init.py:init():783] updated telemetry
19
+ 2024-10-30 11:28:52,896 INFO MainThread:367766 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
20
+ 2024-10-30 11:28:53,152 INFO MainThread:367766 [wandb_init.py:init():867] starting run threads in backend
21
+ 2024-10-30 11:28:53,239 INFO MainThread:367766 [wandb_run.py:_console_start():2463] atexit reg
22
+ 2024-10-30 11:28:53,239 INFO MainThread:367766 [wandb_run.py:_redirect():2311] redirect: wrap_raw
23
+ 2024-10-30 11:28:53,239 INFO MainThread:367766 [wandb_run.py:_redirect():2376] Wrapping output streams.
24
+ 2024-10-30 11:28:53,239 INFO MainThread:367766 [wandb_run.py:_redirect():2401] Redirects installed.
25
+ 2024-10-30 11:28:53,240 INFO MainThread:367766 [wandb_init.py:init():911] run started, returning control to user process
26
+ 2024-10-30 11:28:53,240 INFO MainThread:367766 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'reverse_control', 'train_set': '10M', 'batch_size': 3, 'epoch': 3, 'seed': 0}
27
+ 2024-10-30 18:00:35,945 INFO MainThread:367766 [wandb_run.py:_finish():2158] finishing run yaning1001-dartmouth-college/impossible_llm_reverse/mfvd6tgw
28
+ 2024-10-30 18:00:35,954 INFO MainThread:367766 [wandb_run.py:_atexit_cleanup():2426] got exitcode: 0
29
+ 2024-10-30 18:00:35,955 INFO MainThread:367766 [wandb_run.py:_restore():2408] restore
30
+ 2024-10-30 18:00:35,955 INFO MainThread:367766 [wandb_run.py:_restore():2414] restore done
31
+ 2024-10-30 18:33:04,165 INFO MainThread:367766 [wandb_run.py:_footer_history_summary_info():3975] rendering history
32
+ 2024-10-30 18:33:04,166 INFO MainThread:367766 [wandb_run.py:_footer_history_summary_info():4007] rendering summary
33
+ 2024-10-30 18:33:04,172 INFO MainThread:367766 [wandb_run.py:_footer_sync_info():3934] logging synced files
wandb/run-20241030_233740-np98q8en/run-np98q8en.wandb ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a85d468fd545c171fcd433a279c3cfc6210c775149911b0f456f66e43c4b7c64
3
+ size 851968
wandb/run-20241031_000839-acpkxm8c/files/output.log ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Loading checkpoint shards: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 2/2 [00:18<00:00, 9.31s/it]
2
+ tokenized_valid: Dataset({
3
+ features: ['input_ids', 'attention_mask'],
4
+ num_rows: 600
5
+ })
6
+ /mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/training_args.py:1545: FutureWarning: `evaluation_strategy` is deprecated and will be removed in version 4.46 of 🤗 Transformers. Use `eval_strategy` instead
7
+ warnings.warn(
8
+ [2024-10-31 00:09:00,175] [INFO] [real_accelerator.py:219:get_accelerator] Setting ds_accelerator to cuda (auto detect)
9
+ [2024-10-31 00:09:10,031] [INFO] [comm.py:652:init_distributed] cdb=None
10
+ Installed CUDA version 11.8 does not match the version torch was compiled with 11.7 but since the APIs are compatible, accepting this combination
11
+ Using /home/chunhui/.cache/torch_extensions/py39_cu117 as PyTorch extensions root...
12
+ Loading extension module cpu_adam...
13
+ Time to load cpu_adam op: 5.321483373641968 seconds
wandb/run-20241031_000839-acpkxm8c/files/requirements.txt ADDED
@@ -0,0 +1,147 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ funcsigs==1.0.2
2
+ sentry-sdk==2.17.0
3
+ multiprocess==0.70.16
4
+ numpy==1.26.2
5
+ pluralizer==1.2.0
6
+ debugpy==1.6.7
7
+ nvidia-cudnn-cu11==8.5.0.96
8
+ deepspeed==0.15.2
9
+ data==0.4
10
+ pandas==2.1.3
11
+ tomli==2.0.1
12
+ charset-normalizer==3.3.2
13
+ attrs==24.2.0
14
+ aiosignal==1.3.1
15
+ fsspec==2023.10.0
16
+ nvidia-cusparse-cu11==11.7.4.91
17
+ zipp==3.12.0
18
+ mypy-extensions==1.0.0
19
+ datasets==3.0.1
20
+ joblib==1.3.2
21
+ hjson==3.1.0
22
+ traitlets==5.7.1
23
+ stack-data==0.6.0
24
+ transformers==4.45.1
25
+ sympy==1.11.1
26
+ Pygments==2.15.0
27
+ docker-pycreds==0.4.0
28
+ dill==0.3.8
29
+ wheel==0.44.0
30
+ prompt-toolkit==3.0.30
31
+ parso==0.8.3
32
+ ipykernel==6.23.1
33
+ pyarrow==17.0.0
34
+ certifi==2023.11.17
35
+ nvidia-cufft-cu11==10.9.0.58
36
+ six==1.16.0
37
+ pydantic==2.9.2
38
+ click==8.1.7
39
+ nest-asyncio==1.5.6
40
+ gmpy2==2.1.0
41
+ matplotlib==3.8.2
42
+ scipy==1.11.4
43
+ typing_extensions==4.12.2
44
+ statsmodels==0.14.0
45
+ huggingface-hub==0.25.0
46
+ frozenlist==1.4.1
47
+ gpustat==1.1.1
48
+ nvidia-nvtx-cu11==11.7.91
49
+ safetensors==0.4.5
50
+ stanza==1.9.2
51
+ decorator==5.1.1
52
+ seaborn==0.13.0
53
+ sentencepiece==0.2.0
54
+ PyYAML==6.0.1
55
+ black==24.8.0
56
+ protobuf==4.25.1
57
+ pickleshare==0.7.5
58
+ peft==0.13.0
59
+ triton==2.0.0
60
+ nvidia-cuda-runtime-cu11==11.7.99
61
+ Jinja2==3.1.2
62
+ nvidia-cusolver-cu11==11.4.0.1
63
+ executing==1.2.0
64
+ jupyter_client==8.1.0
65
+ pluggy==1.3.0
66
+ cmake==3.30.3
67
+ pytz==2023.3.post1
68
+ aiohappyeyeballs==2.4.2
69
+ kiwisolver==1.4.5
70
+ py-cpuinfo==9.0.0
71
+ Pillow==10.1.0
72
+ ptyprocess==0.7.0
73
+ importlib_resources==6.4.5
74
+ GitPython==3.1.43
75
+ importlib-metadata==6.0.0
76
+ iniconfig==2.0.0
77
+ scikit-learn==1.3.2
78
+ exceptiongroup==1.1.0
79
+ networkx==2.8.6
80
+ accelerate==1.0.0
81
+ nltk==3.8.1
82
+ shutilwhich==1.1.0
83
+ fonttools==4.45.1
84
+ future==0.18.3
85
+ aiohttp==3.10.6
86
+ wcwidth==0.2.5
87
+ idna==3.6
88
+ filelock==3.12.2
89
+ pathspec==0.12.1
90
+ jupyter_core==5.1.0
91
+ lit==18.1.8
92
+ nvidia-curand-cu11==10.2.10.91
93
+ nvidia-cublas-cu11==11.10.3.66
94
+ nvidia-ml-py==12.560.30
95
+ msgpack==1.1.0
96
+ python-dateutil==2.8.2
97
+ blessed==1.20.0
98
+ packaging==23.0
99
+ gitdb==4.0.11
100
+ yarl==1.13.0
101
+ emoji==2.8.0
102
+ tzdata==2023.3
103
+ cycler==0.12.1
104
+ tornado==6.2
105
+ backcall==0.2.0
106
+ plotnine==0.12.4
107
+ ninja==1.11.1.1
108
+ latex==0.7.0
109
+ wandb==0.18.5
110
+ setproctitle==1.3.3
111
+ threadpoolctl==3.2.0
112
+ requests==2.32.3
113
+ pyparsing==3.1.1
114
+ smmap==5.0.1
115
+ pyzmq==23.0.0
116
+ async-timeout==4.0.3
117
+ annotated-types==0.7.0
118
+ matplotlib-inline==0.1.6
119
+ latexcodec==1.0.0
120
+ ipython==8.0.0
121
+ patsy==0.5.3
122
+ contourpy==1.2.0
123
+ multidict==6.1.0
124
+ mizani==0.9.3
125
+ urllib3==2.1.0
126
+ tokenizers==0.20.0
127
+ MarkupSafe==2.1.2
128
+ pip==24.2
129
+ pexpect==4.8.0
130
+ tqdm==4.66.5
131
+ jedi==0.18.2
132
+ pydantic_core==2.23.4
133
+ tempdir==0.7.1
134
+ mpmath==1.2.1
135
+ setuptools==72.1.0
136
+ pytest==7.4.3
137
+ pure-eval==0.2.2
138
+ psutil==5.9.1
139
+ comm==0.1.2
140
+ nvidia-cuda-cupti-cu11==11.7.101
141
+ nvidia-cuda-nvrtc-cu11==11.7.99
142
+ regex==2023.10.3
143
+ platformdirs==2.5.2
144
+ asttokens==2.2.1
145
+ torch==2.0.0
146
+ nvidia-nccl-cu11==2.14.3
147
+ xxhash==3.5.0
wandb/run-20241031_000839-acpkxm8c/files/wandb-metadata.json ADDED
@@ -0,0 +1,97 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
3
+ "python": "3.9.19",
4
+ "startedAt": "2024-10-31T04:08:39.164895Z",
5
+ "args": [
6
+ "--perturbation",
7
+ "reverse_full",
8
+ "--train_set",
9
+ "10M",
10
+ "--batch_size",
11
+ "3",
12
+ "--epoch",
13
+ "6",
14
+ "--seed",
15
+ "0"
16
+ ],
17
+ "program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
18
+ "codePath": "train/train_deep_wandb.py",
19
+ "git": {
20
+ "remote": "git@hf.co:Yaning1001/Impossible_llm.git",
21
+ "commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
22
+ },
23
+ "email": "yaning1001@gmail.com",
24
+ "root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
25
+ "host": "mms-large-2",
26
+ "username": "chunhui",
27
+ "executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
28
+ "codePathLocal": "train_deep_wandb.py",
29
+ "cpu_count": 32,
30
+ "cpu_count_logical": 64,
31
+ "gpu": "NVIDIA RTX A6000",
32
+ "gpu_count": 8,
33
+ "disk": {
34
+ "/": {
35
+ "total": "1888559353856",
36
+ "used": "1727270539264"
37
+ }
38
+ },
39
+ "memory": {
40
+ "total": "202617098240"
41
+ },
42
+ "cpu": {
43
+ "count": 32,
44
+ "countLogical": 64
45
+ },
46
+ "gpu_nvidia": [
47
+ {
48
+ "name": "NVIDIA RTX A6000",
49
+ "memoryTotal": "51527024640",
50
+ "cudaCores": 10752,
51
+ "architecture": "Ampere"
52
+ },
53
+ {
54
+ "name": "NVIDIA RTX A6000",
55
+ "memoryTotal": "51527024640",
56
+ "cudaCores": 10752,
57
+ "architecture": "Ampere"
58
+ },
59
+ {
60
+ "name": "NVIDIA RTX A6000",
61
+ "memoryTotal": "51527024640",
62
+ "cudaCores": 10752,
63
+ "architecture": "Ampere"
64
+ },
65
+ {
66
+ "name": "NVIDIA RTX A6000",
67
+ "memoryTotal": "51527024640",
68
+ "cudaCores": 10752,
69
+ "architecture": "Ampere"
70
+ },
71
+ {
72
+ "name": "NVIDIA RTX A6000",
73
+ "memoryTotal": "51527024640",
74
+ "cudaCores": 10752,
75
+ "architecture": "Ampere"
76
+ },
77
+ {
78
+ "name": "NVIDIA RTX A6000",
79
+ "memoryTotal": "51527024640",
80
+ "cudaCores": 10752,
81
+ "architecture": "Ampere"
82
+ },
83
+ {
84
+ "name": "NVIDIA RTX A6000",
85
+ "memoryTotal": "51527024640",
86
+ "cudaCores": 10752,
87
+ "architecture": "Ampere"
88
+ },
89
+ {
90
+ "name": "NVIDIA RTX A6000",
91
+ "memoryTotal": "51527024640",
92
+ "cudaCores": 10752,
93
+ "architecture": "Ampere"
94
+ }
95
+ ],
96
+ "cudaVersion": "11.8"
97
+ }
wandb/run-20241031_000839-acpkxm8c/logs/debug-internal.log ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {"time":"2024-10-31T00:08:39.166868509-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
2
+ {"time":"2024-10-31T00:08:39.166882039-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241031_000839-acpkxm8c/logs/debug-core.log"}
3
+ {"time":"2024-10-31T00:08:39.272997415-04:00","level":"INFO","msg":"created new stream","id":"acpkxm8c"}
4
+ {"time":"2024-10-31T00:08:39.273045656-04:00","level":"INFO","msg":"stream: started","id":"acpkxm8c"}
5
+ {"time":"2024-10-31T00:08:39.273081016-04:00","level":"INFO","msg":"sender: started","stream_id":"acpkxm8c"}
6
+ {"time":"2024-10-31T00:08:39.273075736-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"acpkxm8c"}}
7
+ {"time":"2024-10-31T00:08:39.273091156-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"acpkxm8c"}}
8
+ {"time":"2024-10-31T00:08:39.475372241-04:00","level":"INFO","msg":"Starting system monitor"}
wandb/run-20241031_000839-acpkxm8c/logs/debug.log ADDED
@@ -0,0 +1,26 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 2024-10-31 00:08:39,162 INFO MainThread:477299 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
2
+ 2024-10-31 00:08:39,163 INFO MainThread:477299 [wandb_setup.py:_flush():79] Configure stats pid to 477299
3
+ 2024-10-31 00:08:39,163 INFO MainThread:477299 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
4
+ 2024-10-31 00:08:39,163 INFO MainThread:477299 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
5
+ 2024-10-31 00:08:39,163 INFO MainThread:477299 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
6
+ 2024-10-31 00:08:39,163 INFO MainThread:477299 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
7
+ 2024-10-31 00:08:39,163 INFO MainThread:477299 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
8
+ 2024-10-31 00:08:39,163 INFO MainThread:477299 [wandb_setup.py:_flush():79] Applying login settings: {}
9
+ 2024-10-31 00:08:39,163 INFO MainThread:477299 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241031_000839-acpkxm8c/logs/debug.log
10
+ 2024-10-31 00:08:39,163 INFO MainThread:477299 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241031_000839-acpkxm8c/logs/debug-internal.log
11
+ 2024-10-31 00:08:39,163 INFO MainThread:477299 [wandb_init.py:init():621] calling init triggers
12
+ 2024-10-31 00:08:39,163 INFO MainThread:477299 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
13
+ config: {}
14
+ 2024-10-31 00:08:39,163 INFO MainThread:477299 [wandb_init.py:init():671] starting backend
15
+ 2024-10-31 00:08:39,163 INFO MainThread:477299 [wandb_init.py:init():675] sending inform_init request
16
+ 2024-10-31 00:08:39,164 INFO MainThread:477299 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
17
+ 2024-10-31 00:08:39,164 INFO MainThread:477299 [wandb_init.py:init():688] backend started and connected
18
+ 2024-10-31 00:08:39,167 INFO MainThread:477299 [wandb_init.py:init():783] updated telemetry
19
+ 2024-10-31 00:08:39,194 INFO MainThread:477299 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
20
+ 2024-10-31 00:08:39,472 INFO MainThread:477299 [wandb_init.py:init():867] starting run threads in backend
21
+ 2024-10-31 00:08:39,571 INFO MainThread:477299 [wandb_run.py:_console_start():2463] atexit reg
22
+ 2024-10-31 00:08:39,571 INFO MainThread:477299 [wandb_run.py:_redirect():2311] redirect: wrap_raw
23
+ 2024-10-31 00:08:39,571 INFO MainThread:477299 [wandb_run.py:_redirect():2376] Wrapping output streams.
24
+ 2024-10-31 00:08:39,571 INFO MainThread:477299 [wandb_run.py:_redirect():2401] Redirects installed.
25
+ 2024-10-31 00:08:39,572 INFO MainThread:477299 [wandb_init.py:init():911] run started, returning control to user process
26
+ 2024-10-31 00:08:39,572 INFO MainThread:477299 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'reverse_full', 'train_set': '10M', 'batch_size': 3, 'epoch': 6, 'seed': 0, 'lr': 1e-05}
wandb/run-20241031_000839-acpkxm8c/run-acpkxm8c.wandb ADDED
Binary file (65.5 kB). View file
 
wandb/run-20241101_012733-3tsgnm2p/run-3tsgnm2p.wandb ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:157ac6fcdd37aa5de6b664b2b512542501c187eda74f9151bcca4bd636d73df0
3
+ size 1015808
wandb/run-20241101_200502-7hem25r3/files/output.log ADDED
@@ -0,0 +1 @@
 
 
1
+ Loading checkpoint shards: 0%| | 0/2 [00:00<?, ?it/s]
wandb/run-20241101_200502-7hem25r3/files/requirements.txt ADDED
@@ -0,0 +1,147 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ funcsigs==1.0.2
2
+ sentry-sdk==2.17.0
3
+ multiprocess==0.70.16
4
+ numpy==1.26.2
5
+ pluralizer==1.2.0
6
+ debugpy==1.6.7
7
+ nvidia-cudnn-cu11==8.5.0.96
8
+ deepspeed==0.15.2
9
+ data==0.4
10
+ pandas==2.1.3
11
+ tomli==2.0.1
12
+ charset-normalizer==3.3.2
13
+ attrs==24.2.0
14
+ aiosignal==1.3.1
15
+ fsspec==2023.10.0
16
+ nvidia-cusparse-cu11==11.7.4.91
17
+ zipp==3.12.0
18
+ mypy-extensions==1.0.0
19
+ datasets==3.0.1
20
+ joblib==1.3.2
21
+ hjson==3.1.0
22
+ traitlets==5.7.1
23
+ stack-data==0.6.0
24
+ transformers==4.45.1
25
+ sympy==1.11.1
26
+ Pygments==2.15.0
27
+ docker-pycreds==0.4.0
28
+ dill==0.3.8
29
+ wheel==0.44.0
30
+ prompt-toolkit==3.0.30
31
+ parso==0.8.3
32
+ ipykernel==6.23.1
33
+ pyarrow==17.0.0
34
+ certifi==2023.11.17
35
+ nvidia-cufft-cu11==10.9.0.58
36
+ six==1.16.0
37
+ pydantic==2.9.2
38
+ click==8.1.7
39
+ nest-asyncio==1.5.6
40
+ gmpy2==2.1.0
41
+ matplotlib==3.8.2
42
+ scipy==1.11.4
43
+ typing_extensions==4.12.2
44
+ statsmodels==0.14.0
45
+ huggingface-hub==0.25.0
46
+ frozenlist==1.4.1
47
+ gpustat==1.1.1
48
+ nvidia-nvtx-cu11==11.7.91
49
+ safetensors==0.4.5
50
+ stanza==1.9.2
51
+ decorator==5.1.1
52
+ seaborn==0.13.0
53
+ sentencepiece==0.2.0
54
+ PyYAML==6.0.1
55
+ black==24.8.0
56
+ protobuf==4.25.1
57
+ pickleshare==0.7.5
58
+ peft==0.13.0
59
+ triton==2.0.0
60
+ nvidia-cuda-runtime-cu11==11.7.99
61
+ Jinja2==3.1.2
62
+ nvidia-cusolver-cu11==11.4.0.1
63
+ executing==1.2.0
64
+ jupyter_client==8.1.0
65
+ pluggy==1.3.0
66
+ cmake==3.30.3
67
+ pytz==2023.3.post1
68
+ aiohappyeyeballs==2.4.2
69
+ kiwisolver==1.4.5
70
+ py-cpuinfo==9.0.0
71
+ Pillow==10.1.0
72
+ ptyprocess==0.7.0
73
+ importlib_resources==6.4.5
74
+ GitPython==3.1.43
75
+ importlib-metadata==6.0.0
76
+ iniconfig==2.0.0
77
+ scikit-learn==1.3.2
78
+ exceptiongroup==1.1.0
79
+ networkx==2.8.6
80
+ accelerate==1.0.0
81
+ nltk==3.8.1
82
+ shutilwhich==1.1.0
83
+ fonttools==4.45.1
84
+ future==0.18.3
85
+ aiohttp==3.10.6
86
+ wcwidth==0.2.5
87
+ idna==3.6
88
+ filelock==3.12.2
89
+ pathspec==0.12.1
90
+ jupyter_core==5.1.0
91
+ lit==18.1.8
92
+ nvidia-curand-cu11==10.2.10.91
93
+ nvidia-cublas-cu11==11.10.3.66
94
+ nvidia-ml-py==12.560.30
95
+ msgpack==1.1.0
96
+ python-dateutil==2.8.2
97
+ blessed==1.20.0
98
+ packaging==23.0
99
+ gitdb==4.0.11
100
+ yarl==1.13.0
101
+ emoji==2.8.0
102
+ tzdata==2023.3
103
+ cycler==0.12.1
104
+ tornado==6.2
105
+ backcall==0.2.0
106
+ plotnine==0.12.4
107
+ ninja==1.11.1.1
108
+ latex==0.7.0
109
+ wandb==0.18.5
110
+ setproctitle==1.3.3
111
+ threadpoolctl==3.2.0
112
+ requests==2.32.3
113
+ pyparsing==3.1.1
114
+ smmap==5.0.1
115
+ pyzmq==23.0.0
116
+ async-timeout==4.0.3
117
+ annotated-types==0.7.0
118
+ matplotlib-inline==0.1.6
119
+ latexcodec==1.0.0
120
+ ipython==8.0.0
121
+ patsy==0.5.3
122
+ contourpy==1.2.0
123
+ multidict==6.1.0
124
+ mizani==0.9.3
125
+ urllib3==2.1.0
126
+ tokenizers==0.20.0
127
+ MarkupSafe==2.1.2
128
+ pip==24.2
129
+ pexpect==4.8.0
130
+ tqdm==4.66.5
131
+ jedi==0.18.2
132
+ pydantic_core==2.23.4
133
+ tempdir==0.7.1
134
+ mpmath==1.2.1
135
+ setuptools==72.1.0
136
+ pytest==7.4.3
137
+ pure-eval==0.2.2
138
+ psutil==5.9.1
139
+ comm==0.1.2
140
+ nvidia-cuda-cupti-cu11==11.7.101
141
+ nvidia-cuda-nvrtc-cu11==11.7.99
142
+ regex==2023.10.3
143
+ platformdirs==2.5.2
144
+ asttokens==2.2.1
145
+ torch==2.0.0
146
+ nvidia-nccl-cu11==2.14.3
147
+ xxhash==3.5.0
wandb/run-20241101_200502-7hem25r3/files/wandb-metadata.json ADDED
@@ -0,0 +1,97 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
3
+ "python": "3.9.19",
4
+ "startedAt": "2024-11-02T00:05:02.628844Z",
5
+ "args": [
6
+ "--perturbation",
7
+ "shuffle_nondeterministic",
8
+ "--train_set",
9
+ "10M",
10
+ "--batch_size",
11
+ "3",
12
+ "--epoch",
13
+ "3",
14
+ "--seed",
15
+ "0"
16
+ ],
17
+ "program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
18
+ "codePath": "train/train_deep_wandb.py",
19
+ "git": {
20
+ "remote": "git@hf.co:Yaning1001/Impossible_llm.git",
21
+ "commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
22
+ },
23
+ "email": "yaning1001@gmail.com",
24
+ "root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
25
+ "host": "mms-large-2",
26
+ "username": "chunhui",
27
+ "executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
28
+ "codePathLocal": "train_deep_wandb.py",
29
+ "cpu_count": 32,
30
+ "cpu_count_logical": 64,
31
+ "gpu": "NVIDIA RTX A6000",
32
+ "gpu_count": 8,
33
+ "disk": {
34
+ "/": {
35
+ "total": "1888559353856",
36
+ "used": "1754801463296"
37
+ }
38
+ },
39
+ "memory": {
40
+ "total": "202617098240"
41
+ },
42
+ "cpu": {
43
+ "count": 32,
44
+ "countLogical": 64
45
+ },
46
+ "gpu_nvidia": [
47
+ {
48
+ "name": "NVIDIA RTX A6000",
49
+ "memoryTotal": "51527024640",
50
+ "cudaCores": 10752,
51
+ "architecture": "Ampere"
52
+ },
53
+ {
54
+ "name": "NVIDIA RTX A6000",
55
+ "memoryTotal": "51527024640",
56
+ "cudaCores": 10752,
57
+ "architecture": "Ampere"
58
+ },
59
+ {
60
+ "name": "NVIDIA RTX A6000",
61
+ "memoryTotal": "51527024640",
62
+ "cudaCores": 10752,
63
+ "architecture": "Ampere"
64
+ },
65
+ {
66
+ "name": "NVIDIA RTX A6000",
67
+ "memoryTotal": "51527024640",
68
+ "cudaCores": 10752,
69
+ "architecture": "Ampere"
70
+ },
71
+ {
72
+ "name": "NVIDIA RTX A6000",
73
+ "memoryTotal": "51527024640",
74
+ "cudaCores": 10752,
75
+ "architecture": "Ampere"
76
+ },
77
+ {
78
+ "name": "NVIDIA RTX A6000",
79
+ "memoryTotal": "51527024640",
80
+ "cudaCores": 10752,
81
+ "architecture": "Ampere"
82
+ },
83
+ {
84
+ "name": "NVIDIA RTX A6000",
85
+ "memoryTotal": "51527024640",
86
+ "cudaCores": 10752,
87
+ "architecture": "Ampere"
88
+ },
89
+ {
90
+ "name": "NVIDIA RTX A6000",
91
+ "memoryTotal": "51527024640",
92
+ "cudaCores": 10752,
93
+ "architecture": "Ampere"
94
+ }
95
+ ],
96
+ "cudaVersion": "11.8"
97
+ }
wandb/run-20241101_200502-7hem25r3/logs/debug-internal.log ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {"time":"2024-11-01T20:05:02.633735347-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
2
+ {"time":"2024-11-01T20:05:02.633752607-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241101_200502-7hem25r3/logs/debug-core.log"}
3
+ {"time":"2024-11-01T20:05:02.844674605-04:00","level":"INFO","msg":"created new stream","id":"7hem25r3"}
4
+ {"time":"2024-11-01T20:05:02.844749605-04:00","level":"INFO","msg":"stream: started","id":"7hem25r3"}
5
+ {"time":"2024-11-01T20:05:02.844804386-04:00","level":"INFO","msg":"sender: started","stream_id":"7hem25r3"}
6
+ {"time":"2024-11-01T20:05:02.844777586-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"7hem25r3"}}
7
+ {"time":"2024-11-01T20:05:02.844794496-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"7hem25r3"}}
8
+ {"time":"2024-11-01T20:05:03.055749493-04:00","level":"INFO","msg":"Starting system monitor"}
wandb/run-20241101_200502-7hem25r3/logs/debug.log ADDED
@@ -0,0 +1,26 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 2024-11-01 20:05:02,626 INFO MainThread:869512 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
2
+ 2024-11-01 20:05:02,626 INFO MainThread:869512 [wandb_setup.py:_flush():79] Configure stats pid to 869512
3
+ 2024-11-01 20:05:02,626 INFO MainThread:869512 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
4
+ 2024-11-01 20:05:02,626 INFO MainThread:869512 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
5
+ 2024-11-01 20:05:02,626 INFO MainThread:869512 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
6
+ 2024-11-01 20:05:02,626 INFO MainThread:869512 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
7
+ 2024-11-01 20:05:02,626 INFO MainThread:869512 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
8
+ 2024-11-01 20:05:02,626 INFO MainThread:869512 [wandb_setup.py:_flush():79] Applying login settings: {}
9
+ 2024-11-01 20:05:02,626 INFO MainThread:869512 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241101_200502-7hem25r3/logs/debug.log
10
+ 2024-11-01 20:05:02,626 INFO MainThread:869512 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241101_200502-7hem25r3/logs/debug-internal.log
11
+ 2024-11-01 20:05:02,626 INFO MainThread:869512 [wandb_init.py:init():621] calling init triggers
12
+ 2024-11-01 20:05:02,627 INFO MainThread:869512 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
13
+ config: {}
14
+ 2024-11-01 20:05:02,627 INFO MainThread:869512 [wandb_init.py:init():671] starting backend
15
+ 2024-11-01 20:05:02,627 INFO MainThread:869512 [wandb_init.py:init():675] sending inform_init request
16
+ 2024-11-01 20:05:02,628 INFO MainThread:869512 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
17
+ 2024-11-01 20:05:02,628 INFO MainThread:869512 [wandb_init.py:init():688] backend started and connected
18
+ 2024-11-01 20:05:02,631 INFO MainThread:869512 [wandb_init.py:init():783] updated telemetry
19
+ 2024-11-01 20:05:02,651 INFO MainThread:869512 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
20
+ 2024-11-01 20:05:03,053 INFO MainThread:869512 [wandb_init.py:init():867] starting run threads in backend
21
+ 2024-11-01 20:05:03,140 INFO MainThread:869512 [wandb_run.py:_console_start():2463] atexit reg
22
+ 2024-11-01 20:05:03,140 INFO MainThread:869512 [wandb_run.py:_redirect():2311] redirect: wrap_raw
23
+ 2024-11-01 20:05:03,140 INFO MainThread:869512 [wandb_run.py:_redirect():2376] Wrapping output streams.
24
+ 2024-11-01 20:05:03,140 INFO MainThread:869512 [wandb_run.py:_redirect():2401] Redirects installed.
25
+ 2024-11-01 20:05:03,142 INFO MainThread:869512 [wandb_init.py:init():911] run started, returning control to user process
26
+ 2024-11-01 20:05:03,142 INFO MainThread:869512 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'shuffle_nondeterministic', 'train_set': '10M', 'batch_size': 3, 'epoch': 3, 'seed': 0, 'lr': 5e-06}
wandb/run-20241101_200502-7hem25r3/run-7hem25r3.wandb ADDED
File without changes
wandb/run-20241101_200517-vzv5zg2q/files/config.yaml ADDED
@@ -0,0 +1,49 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ _wandb:
2
+ value:
3
+ cli_version: 0.18.5
4
+ m: []
5
+ python_version: 3.9.19
6
+ t:
7
+ "1":
8
+ - 1
9
+ - 5
10
+ - 11
11
+ - 49
12
+ - 51
13
+ - 53
14
+ - 55
15
+ - 71
16
+ - 98
17
+ "2":
18
+ - 1
19
+ - 5
20
+ - 11
21
+ - 49
22
+ - 51
23
+ - 53
24
+ - 55
25
+ - 71
26
+ - 98
27
+ "3":
28
+ - 13
29
+ - 23
30
+ - 55
31
+ "4": 3.9.19
32
+ "5": 0.18.5
33
+ "6": 4.45.1
34
+ "8":
35
+ - 5
36
+ "12": 0.18.5
37
+ "13": linux-x86_64
38
+ batch_size:
39
+ value: 3
40
+ epoch:
41
+ value: 3
42
+ lr:
43
+ value: 5e-06
44
+ perturbation:
45
+ value: shuffle_nondeterministic
46
+ seed:
47
+ value: 0
48
+ train_set:
49
+ value: 10M
wandb/run-20241101_200517-vzv5zg2q/files/output.log ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Downloading shards: 0%| | 0/2 [00:07<?, ?it/s]
2
+ Error in sys.excepthook:
3
+ Traceback (most recent call last):
4
+ File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/wandb/sdk/lib/exit_hooks.py", line 52, in exc_handler
5
+ traceback.print_exception(exc_type, exc, tb)
6
+ KeyboardInterrupt
7
+
8
+ Original exception was:
9
+ Traceback (most recent call last):
10
+ File "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py", line 173, in <module>
11
+ model = AutoModelForCausalLM.from_pretrained(model_name,
12
+ File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/models/auto/auto_factory.py", line 564, in from_pretrained
13
+ return model_class.from_pretrained(
14
+ File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/modeling_utils.py", line 3769, in from_pretrained
15
+ resolved_archive_file, sharded_metadata = get_checkpoint_shard_files(
16
+ File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/utils/hub.py", line 1098, in get_checkpoint_shard_files
17
+ cached_filename = cached_file(
18
+ File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/transformers/utils/hub.py", line 403, in cached_file
19
+ resolved_file = hf_hub_download(
20
+ File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/utils/_deprecation.py", line 101, in inner_f
21
+ return f(*args, **kwargs)
22
+ File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/utils/_validators.py", line 114, in _inner_fn
23
+ return fn(*args, **kwargs)
24
+ File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/file_download.py", line 1232, in hf_hub_download
25
+ return _hf_hub_download_to_cache_dir(
26
+ File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/file_download.py", line 1380, in _hf_hub_download_to_cache_dir
27
+ with WeakFileLock(lock_path):
28
+ File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/contextlib.py", line 119, in __enter__
29
+ return next(self.gen)
30
+ File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/huggingface_hub/utils/_fixes.py", line 98, in WeakFileLock
31
+ lock.acquire()
32
+ File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/filelock/_api.py", line 225, in acquire
33
+ time.sleep(poll_interval)
34
+ KeyboardInterrupt
wandb/run-20241101_200517-vzv5zg2q/files/requirements.txt ADDED
@@ -0,0 +1,147 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ funcsigs==1.0.2
2
+ sentry-sdk==2.17.0
3
+ multiprocess==0.70.16
4
+ numpy==1.26.2
5
+ pluralizer==1.2.0
6
+ debugpy==1.6.7
7
+ nvidia-cudnn-cu11==8.5.0.96
8
+ deepspeed==0.15.2
9
+ data==0.4
10
+ pandas==2.1.3
11
+ tomli==2.0.1
12
+ charset-normalizer==3.3.2
13
+ attrs==24.2.0
14
+ aiosignal==1.3.1
15
+ fsspec==2023.10.0
16
+ nvidia-cusparse-cu11==11.7.4.91
17
+ zipp==3.12.0
18
+ mypy-extensions==1.0.0
19
+ datasets==3.0.1
20
+ joblib==1.3.2
21
+ hjson==3.1.0
22
+ traitlets==5.7.1
23
+ stack-data==0.6.0
24
+ transformers==4.45.1
25
+ sympy==1.11.1
26
+ Pygments==2.15.0
27
+ docker-pycreds==0.4.0
28
+ dill==0.3.8
29
+ wheel==0.44.0
30
+ prompt-toolkit==3.0.30
31
+ parso==0.8.3
32
+ ipykernel==6.23.1
33
+ pyarrow==17.0.0
34
+ certifi==2023.11.17
35
+ nvidia-cufft-cu11==10.9.0.58
36
+ six==1.16.0
37
+ pydantic==2.9.2
38
+ click==8.1.7
39
+ nest-asyncio==1.5.6
40
+ gmpy2==2.1.0
41
+ matplotlib==3.8.2
42
+ scipy==1.11.4
43
+ typing_extensions==4.12.2
44
+ statsmodels==0.14.0
45
+ huggingface-hub==0.25.0
46
+ frozenlist==1.4.1
47
+ gpustat==1.1.1
48
+ nvidia-nvtx-cu11==11.7.91
49
+ safetensors==0.4.5
50
+ stanza==1.9.2
51
+ decorator==5.1.1
52
+ seaborn==0.13.0
53
+ sentencepiece==0.2.0
54
+ PyYAML==6.0.1
55
+ black==24.8.0
56
+ protobuf==4.25.1
57
+ pickleshare==0.7.5
58
+ peft==0.13.0
59
+ triton==2.0.0
60
+ nvidia-cuda-runtime-cu11==11.7.99
61
+ Jinja2==3.1.2
62
+ nvidia-cusolver-cu11==11.4.0.1
63
+ executing==1.2.0
64
+ jupyter_client==8.1.0
65
+ pluggy==1.3.0
66
+ cmake==3.30.3
67
+ pytz==2023.3.post1
68
+ aiohappyeyeballs==2.4.2
69
+ kiwisolver==1.4.5
70
+ py-cpuinfo==9.0.0
71
+ Pillow==10.1.0
72
+ ptyprocess==0.7.0
73
+ importlib_resources==6.4.5
74
+ GitPython==3.1.43
75
+ importlib-metadata==6.0.0
76
+ iniconfig==2.0.0
77
+ scikit-learn==1.3.2
78
+ exceptiongroup==1.1.0
79
+ networkx==2.8.6
80
+ accelerate==1.0.0
81
+ nltk==3.8.1
82
+ shutilwhich==1.1.0
83
+ fonttools==4.45.1
84
+ future==0.18.3
85
+ aiohttp==3.10.6
86
+ wcwidth==0.2.5
87
+ idna==3.6
88
+ filelock==3.12.2
89
+ pathspec==0.12.1
90
+ jupyter_core==5.1.0
91
+ lit==18.1.8
92
+ nvidia-curand-cu11==10.2.10.91
93
+ nvidia-cublas-cu11==11.10.3.66
94
+ nvidia-ml-py==12.560.30
95
+ msgpack==1.1.0
96
+ python-dateutil==2.8.2
97
+ blessed==1.20.0
98
+ packaging==23.0
99
+ gitdb==4.0.11
100
+ yarl==1.13.0
101
+ emoji==2.8.0
102
+ tzdata==2023.3
103
+ cycler==0.12.1
104
+ tornado==6.2
105
+ backcall==0.2.0
106
+ plotnine==0.12.4
107
+ ninja==1.11.1.1
108
+ latex==0.7.0
109
+ wandb==0.18.5
110
+ setproctitle==1.3.3
111
+ threadpoolctl==3.2.0
112
+ requests==2.32.3
113
+ pyparsing==3.1.1
114
+ smmap==5.0.1
115
+ pyzmq==23.0.0
116
+ async-timeout==4.0.3
117
+ annotated-types==0.7.0
118
+ matplotlib-inline==0.1.6
119
+ latexcodec==1.0.0
120
+ ipython==8.0.0
121
+ patsy==0.5.3
122
+ contourpy==1.2.0
123
+ multidict==6.1.0
124
+ mizani==0.9.3
125
+ urllib3==2.1.0
126
+ tokenizers==0.20.0
127
+ MarkupSafe==2.1.2
128
+ pip==24.2
129
+ pexpect==4.8.0
130
+ tqdm==4.66.5
131
+ jedi==0.18.2
132
+ pydantic_core==2.23.4
133
+ tempdir==0.7.1
134
+ mpmath==1.2.1
135
+ setuptools==72.1.0
136
+ pytest==7.4.3
137
+ pure-eval==0.2.2
138
+ psutil==5.9.1
139
+ comm==0.1.2
140
+ nvidia-cuda-cupti-cu11==11.7.101
141
+ nvidia-cuda-nvrtc-cu11==11.7.99
142
+ regex==2023.10.3
143
+ platformdirs==2.5.2
144
+ asttokens==2.2.1
145
+ torch==2.0.0
146
+ nvidia-nccl-cu11==2.14.3
147
+ xxhash==3.5.0
wandb/run-20241101_200517-vzv5zg2q/files/wandb-metadata.json ADDED
@@ -0,0 +1,97 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
3
+ "python": "3.9.19",
4
+ "startedAt": "2024-11-02T00:05:17.600434Z",
5
+ "args": [
6
+ "--perturbation",
7
+ "shuffle_nondeterministic",
8
+ "--train_set",
9
+ "10M",
10
+ "--batch_size",
11
+ "3",
12
+ "--epoch",
13
+ "3",
14
+ "--seed",
15
+ "0"
16
+ ],
17
+ "program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
18
+ "codePath": "train/train_deep_wandb.py",
19
+ "git": {
20
+ "remote": "git@hf.co:Yaning1001/Impossible_llm.git",
21
+ "commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
22
+ },
23
+ "email": "yaning1001@gmail.com",
24
+ "root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
25
+ "host": "mms-large-2",
26
+ "username": "chunhui",
27
+ "executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
28
+ "codePathLocal": "train_deep_wandb.py",
29
+ "cpu_count": 32,
30
+ "cpu_count_logical": 64,
31
+ "gpu": "NVIDIA RTX A6000",
32
+ "gpu_count": 8,
33
+ "disk": {
34
+ "/": {
35
+ "total": "1888559353856",
36
+ "used": "1754801557504"
37
+ }
38
+ },
39
+ "memory": {
40
+ "total": "202617098240"
41
+ },
42
+ "cpu": {
43
+ "count": 32,
44
+ "countLogical": 64
45
+ },
46
+ "gpu_nvidia": [
47
+ {
48
+ "name": "NVIDIA RTX A6000",
49
+ "memoryTotal": "51527024640",
50
+ "cudaCores": 10752,
51
+ "architecture": "Ampere"
52
+ },
53
+ {
54
+ "name": "NVIDIA RTX A6000",
55
+ "memoryTotal": "51527024640",
56
+ "cudaCores": 10752,
57
+ "architecture": "Ampere"
58
+ },
59
+ {
60
+ "name": "NVIDIA RTX A6000",
61
+ "memoryTotal": "51527024640",
62
+ "cudaCores": 10752,
63
+ "architecture": "Ampere"
64
+ },
65
+ {
66
+ "name": "NVIDIA RTX A6000",
67
+ "memoryTotal": "51527024640",
68
+ "cudaCores": 10752,
69
+ "architecture": "Ampere"
70
+ },
71
+ {
72
+ "name": "NVIDIA RTX A6000",
73
+ "memoryTotal": "51527024640",
74
+ "cudaCores": 10752,
75
+ "architecture": "Ampere"
76
+ },
77
+ {
78
+ "name": "NVIDIA RTX A6000",
79
+ "memoryTotal": "51527024640",
80
+ "cudaCores": 10752,
81
+ "architecture": "Ampere"
82
+ },
83
+ {
84
+ "name": "NVIDIA RTX A6000",
85
+ "memoryTotal": "51527024640",
86
+ "cudaCores": 10752,
87
+ "architecture": "Ampere"
88
+ },
89
+ {
90
+ "name": "NVIDIA RTX A6000",
91
+ "memoryTotal": "51527024640",
92
+ "cudaCores": 10752,
93
+ "architecture": "Ampere"
94
+ }
95
+ ],
96
+ "cudaVersion": "11.8"
97
+ }
wandb/run-20241101_200517-vzv5zg2q/files/wandb-summary.json ADDED
@@ -0,0 +1 @@
 
 
1
+ {"_wandb":{"runtime":7}}
wandb/run-20241101_200517-vzv5zg2q/logs/debug-internal.log ADDED
@@ -0,0 +1,11 @@
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {"time":"2024-11-01T20:05:17.602749928-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
2
+ {"time":"2024-11-01T20:05:17.602763728-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241101_200517-vzv5zg2q/logs/debug-core.log"}
3
+ {"time":"2024-11-01T20:05:17.709188149-04:00","level":"INFO","msg":"created new stream","id":"vzv5zg2q"}
4
+ {"time":"2024-11-01T20:05:17.70922197-04:00","level":"INFO","msg":"stream: started","id":"vzv5zg2q"}
5
+ {"time":"2024-11-01T20:05:17.70926242-04:00","level":"INFO","msg":"sender: started","stream_id":"vzv5zg2q"}
6
+ {"time":"2024-11-01T20:05:17.70924561-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"vzv5zg2q"}}
7
+ {"time":"2024-11-01T20:05:17.709359941-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"vzv5zg2q"}}
8
+ {"time":"2024-11-01T20:05:17.906758123-04:00","level":"INFO","msg":"Starting system monitor"}
9
+ {"time":"2024-11-01T20:05:25.260555442-04:00","level":"INFO","msg":"stream: closing","id":"vzv5zg2q"}
10
+ {"time":"2024-11-01T20:05:25.260678473-04:00","level":"INFO","msg":"Stopping system monitor"}
11
+ {"time":"2024-11-01T20:05:25.261811351-04:00","level":"INFO","msg":"Stopped system monitor"}
wandb/run-20241101_200517-vzv5zg2q/logs/debug.log ADDED
@@ -0,0 +1,27 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 2024-11-01 20:05:17,597 INFO MainThread:870384 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
2
+ 2024-11-01 20:05:17,597 INFO MainThread:870384 [wandb_setup.py:_flush():79] Configure stats pid to 870384
3
+ 2024-11-01 20:05:17,597 INFO MainThread:870384 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
4
+ 2024-11-01 20:05:17,597 INFO MainThread:870384 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
5
+ 2024-11-01 20:05:17,597 INFO MainThread:870384 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
6
+ 2024-11-01 20:05:17,597 INFO MainThread:870384 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
7
+ 2024-11-01 20:05:17,597 INFO MainThread:870384 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
8
+ 2024-11-01 20:05:17,597 INFO MainThread:870384 [wandb_setup.py:_flush():79] Applying login settings: {}
9
+ 2024-11-01 20:05:17,597 INFO MainThread:870384 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241101_200517-vzv5zg2q/logs/debug.log
10
+ 2024-11-01 20:05:17,597 INFO MainThread:870384 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241101_200517-vzv5zg2q/logs/debug-internal.log
11
+ 2024-11-01 20:05:17,598 INFO MainThread:870384 [wandb_init.py:init():621] calling init triggers
12
+ 2024-11-01 20:05:17,598 INFO MainThread:870384 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
13
+ config: {}
14
+ 2024-11-01 20:05:17,598 INFO MainThread:870384 [wandb_init.py:init():671] starting backend
15
+ 2024-11-01 20:05:17,598 INFO MainThread:870384 [wandb_init.py:init():675] sending inform_init request
16
+ 2024-11-01 20:05:17,599 INFO MainThread:870384 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
17
+ 2024-11-01 20:05:17,600 INFO MainThread:870384 [wandb_init.py:init():688] backend started and connected
18
+ 2024-11-01 20:05:17,603 INFO MainThread:870384 [wandb_init.py:init():783] updated telemetry
19
+ 2024-11-01 20:05:17,631 INFO MainThread:870384 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
20
+ 2024-11-01 20:05:17,903 INFO MainThread:870384 [wandb_init.py:init():867] starting run threads in backend
21
+ 2024-11-01 20:05:17,993 INFO MainThread:870384 [wandb_run.py:_console_start():2463] atexit reg
22
+ 2024-11-01 20:05:17,993 INFO MainThread:870384 [wandb_run.py:_redirect():2311] redirect: wrap_raw
23
+ 2024-11-01 20:05:17,993 INFO MainThread:870384 [wandb_run.py:_redirect():2376] Wrapping output streams.
24
+ 2024-11-01 20:05:17,993 INFO MainThread:870384 [wandb_run.py:_redirect():2401] Redirects installed.
25
+ 2024-11-01 20:05:17,994 INFO MainThread:870384 [wandb_init.py:init():911] run started, returning control to user process
26
+ 2024-11-01 20:05:17,994 INFO MainThread:870384 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'shuffle_nondeterministic', 'train_set': '10M', 'batch_size': 3, 'epoch': 3, 'seed': 0, 'lr': 5e-06}
27
+ 2024-11-01 20:05:25,260 WARNING MsgRouterThr:870384 [router.py:message_loop():77] message_loop has been closed
wandb/run-20241101_200517-vzv5zg2q/run-vzv5zg2q.wandb ADDED
File without changes
wandb/run-20241101_201630-e5gt2fir/files/config.yaml ADDED
@@ -0,0 +1,49 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ _wandb:
2
+ value:
3
+ cli_version: 0.18.5
4
+ m: []
5
+ python_version: 3.9.19
6
+ t:
7
+ "1":
8
+ - 1
9
+ - 5
10
+ - 11
11
+ - 49
12
+ - 51
13
+ - 53
14
+ - 55
15
+ - 71
16
+ - 98
17
+ "2":
18
+ - 1
19
+ - 5
20
+ - 11
21
+ - 49
22
+ - 51
23
+ - 53
24
+ - 55
25
+ - 71
26
+ - 98
27
+ "3":
28
+ - 13
29
+ - 23
30
+ - 55
31
+ "4": 3.9.19
32
+ "5": 0.18.5
33
+ "6": 4.45.1
34
+ "8":
35
+ - 5
36
+ "12": 0.18.5
37
+ "13": linux-x86_64
38
+ batch_size:
39
+ value: 3
40
+ epoch:
41
+ value: 6
42
+ lr:
43
+ value: 5e-06
44
+ perturbation:
45
+ value: shuffle_nodeterministic
46
+ seed:
47
+ value: 0
48
+ train_set:
49
+ value: 10M
wandb/run-20241101_201630-e5gt2fir/files/output.log ADDED
@@ -0,0 +1,12 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Traceback (most recent call last):
2
+ File "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py", line 164, in <module>
3
+ dataset = load_dataset('babylm_dataset_test.py', name=dataset_name, trust_remote_code=True)
4
+ File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/datasets/load.py", line 2074, in load_dataset
5
+ builder_instance = load_dataset_builder(
6
+ File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/datasets/load.py", line 1832, in load_dataset_builder
7
+ builder_instance: DatasetBuilder = builder_cls(
8
+ File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/datasets/builder.py", line 342, in __init__
9
+ self.config, self.config_id = self._create_builder_config(
10
+ File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/datasets/builder.py", line 569, in _create_builder_config
11
+ raise ValueError(
12
+ ValueError: BuilderConfig 'babylm_shuffle_nodeterministic_10M_seed0' not found. Available: ['babylm_hop_control_10M_seed0', 'babylm_hop_tokens4_10M_seed0', 'babylm_hop_words4_10M_seed0', 'babylm_reverse_control_10M_seed0', 'babylm_reverse_partial_10M_seed0', 'babylm_reverse_full_10M_seed0', 'babylm_shuffle_control_10M_seed0', 'babylm_shuffle_nondeterministic_10M_seed0', 'babylm_shuffle_deterministic21_10M_seed0', 'babylm_shuffle_deterministic57_10M_seed0', 'babylm_shuffle_deterministic84_10M_seed0', 'babylm_shuffle_local3_10M_seed0', 'babylm_shuffle_local5_10M_seed0', 'babylm_shuffle_local10_10M_seed0', 'babylm_shuffle_even_odd_10M_seed0']
wandb/run-20241101_201630-e5gt2fir/files/wandb-summary.json ADDED
@@ -0,0 +1 @@
 
 
1
+ {"_wandb":{"runtime":0}}
wandb/run-20241101_201630-e5gt2fir/logs/debug-internal.log ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {"time":"2024-11-01T20:16:30.561699495-04:00","level":"INFO","msg":"using version","core version":"0.18.5"}
2
+ {"time":"2024-11-01T20:16:30.561714855-04:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241101_201630-e5gt2fir/logs/debug-core.log"}
3
+ {"time":"2024-11-01T20:16:30.670004167-04:00","level":"INFO","msg":"created new stream","id":"e5gt2fir"}
4
+ {"time":"2024-11-01T20:16:30.670039398-04:00","level":"INFO","msg":"stream: started","id":"e5gt2fir"}
5
+ {"time":"2024-11-01T20:16:30.670068758-04:00","level":"INFO","msg":"sender: started","stream_id":"e5gt2fir"}
6
+ {"time":"2024-11-01T20:16:30.670067278-04:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"e5gt2fir"}}
7
+ {"time":"2024-11-01T20:16:30.670100128-04:00","level":"INFO","msg":"handler: started","stream_id":{"value":"e5gt2fir"}}
8
+ {"time":"2024-11-01T20:16:30.881477183-04:00","level":"INFO","msg":"Starting system monitor"}
9
+ {"time":"2024-11-01T20:16:30.974886093-04:00","level":"INFO","msg":"stream: closing","id":"e5gt2fir"}
10
+ {"time":"2024-11-01T20:16:30.974923473-04:00","level":"INFO","msg":"Stopping system monitor"}
11
+ {"time":"2024-11-01T20:16:30.982511386-04:00","level":"INFO","msg":"Stopped system monitor"}
12
+ {"time":"2024-11-01T20:16:31.563553389-04:00","level":"INFO","msg":"fileTransfer: Close: file transfer manager closed"}
13
+ {"time":"2024-11-01T20:16:31.686184401-04:00","level":"INFO","msg":"handler: closed","stream_id":{"value":"e5gt2fir"}}
14
+ {"time":"2024-11-01T20:16:31.686245181-04:00","level":"INFO","msg":"writer: Close: closed","stream_id":{"value":"e5gt2fir"}}
15
+ {"time":"2024-11-01T20:16:31.686264612-04:00","level":"INFO","msg":"sender: closed","stream_id":"e5gt2fir"}
16
+ {"time":"2024-11-01T20:16:31.686314312-04:00","level":"INFO","msg":"stream: closed","id":"e5gt2fir"}
wandb/run-20241101_201630-e5gt2fir/logs/debug.log ADDED
@@ -0,0 +1,27 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 2024-11-01 20:16:30,556 INFO MainThread:874718 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
2
+ 2024-11-01 20:16:30,557 INFO MainThread:874718 [wandb_setup.py:_flush():79] Configure stats pid to 874718
3
+ 2024-11-01 20:16:30,557 INFO MainThread:874718 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
4
+ 2024-11-01 20:16:30,557 INFO MainThread:874718 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
5
+ 2024-11-01 20:16:30,557 INFO MainThread:874718 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
6
+ 2024-11-01 20:16:30,557 INFO MainThread:874718 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
7
+ 2024-11-01 20:16:30,557 INFO MainThread:874718 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
8
+ 2024-11-01 20:16:30,557 INFO MainThread:874718 [wandb_setup.py:_flush():79] Applying login settings: {}
9
+ 2024-11-01 20:16:30,557 INFO MainThread:874718 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241101_201630-e5gt2fir/logs/debug.log
10
+ 2024-11-01 20:16:30,557 INFO MainThread:874718 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241101_201630-e5gt2fir/logs/debug-internal.log
11
+ 2024-11-01 20:16:30,557 INFO MainThread:874718 [wandb_init.py:init():621] calling init triggers
12
+ 2024-11-01 20:16:30,557 INFO MainThread:874718 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
13
+ config: {}
14
+ 2024-11-01 20:16:30,557 INFO MainThread:874718 [wandb_init.py:init():671] starting backend
15
+ 2024-11-01 20:16:30,557 INFO MainThread:874718 [wandb_init.py:init():675] sending inform_init request
16
+ 2024-11-01 20:16:30,559 INFO MainThread:874718 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
17
+ 2024-11-01 20:16:30,559 INFO MainThread:874718 [wandb_init.py:init():688] backend started and connected
18
+ 2024-11-01 20:16:30,563 INFO MainThread:874718 [wandb_init.py:init():783] updated telemetry
19
+ 2024-11-01 20:16:30,592 INFO MainThread:874718 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
20
+ 2024-11-01 20:16:30,878 INFO MainThread:874718 [wandb_init.py:init():867] starting run threads in backend
21
+ 2024-11-01 20:16:30,965 INFO MainThread:874718 [wandb_run.py:_console_start():2463] atexit reg
22
+ 2024-11-01 20:16:30,965 INFO MainThread:874718 [wandb_run.py:_redirect():2311] redirect: wrap_raw
23
+ 2024-11-01 20:16:30,965 INFO MainThread:874718 [wandb_run.py:_redirect():2376] Wrapping output streams.
24
+ 2024-11-01 20:16:30,965 INFO MainThread:874718 [wandb_run.py:_redirect():2401] Redirects installed.
25
+ 2024-11-01 20:16:30,967 INFO MainThread:874718 [wandb_init.py:init():911] run started, returning control to user process
26
+ 2024-11-01 20:16:30,967 INFO MainThread:874718 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'shuffle_nodeterministic', 'train_set': '10M', 'batch_size': 3, 'epoch': 6, 'seed': 0, 'lr': 5e-06}
27
+ 2024-11-01 20:16:30,974 WARNING MsgRouterThr:874718 [router.py:message_loop():77] message_loop has been closed
wandb/run-20241101_202058-hjyig8so/run-hjyig8so.wandb ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2bb18c91a178f700bc9b2e5d6d0e4bea78a095f878d6c225bc65c1b29e8d0dd1
3
+ size 13297708
wandb/run-20241105_160652-v9udw9ab/files/config.yaml ADDED
@@ -0,0 +1,49 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ _wandb:
2
+ value:
3
+ cli_version: 0.18.5
4
+ m: []
5
+ python_version: 3.9.19
6
+ t:
7
+ "1":
8
+ - 1
9
+ - 5
10
+ - 11
11
+ - 49
12
+ - 51
13
+ - 53
14
+ - 55
15
+ - 71
16
+ - 98
17
+ "2":
18
+ - 1
19
+ - 5
20
+ - 11
21
+ - 49
22
+ - 51
23
+ - 53
24
+ - 55
25
+ - 71
26
+ - 98
27
+ "3":
28
+ - 13
29
+ - 23
30
+ - 55
31
+ "4": 3.9.19
32
+ "5": 0.18.5
33
+ "6": 4.45.1
34
+ "8":
35
+ - 5
36
+ "12": 0.18.5
37
+ "13": linux-x86_64
38
+ batch_size:
39
+ value: 3
40
+ epoch:
41
+ value: 3
42
+ lr:
43
+ value: 5e-06
44
+ perturbation:
45
+ value: shuffle_deterministic21
46
+ seed:
47
+ value: 0
48
+ train_set:
49
+ value: 10M
wandb/run-20241105_160652-v9udw9ab/files/output.log ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ Traceback (most recent call last):
2
+ File "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py", line 165, in <module>
3
+ dataset = load_dataset('babylm_dataset_test.py', name=dataset_name, trust_remote_code=True)
4
+ File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/datasets/load.py", line 2096, in load_dataset
5
+ builder_instance.download_and_prepare(
6
+ File "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/lib/python3.9/site-packages/datasets/builder.py", line 875, in download_and_prepare
7
+ raise OSError(
8
+ OSError: Not enough disk space. Needed: Unknown size (download: Unknown size, generated: Unknown size, post-processed: Unknown size)
wandb/run-20241105_160652-v9udw9ab/files/requirements.txt ADDED
@@ -0,0 +1,147 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ funcsigs==1.0.2
2
+ sentry-sdk==2.17.0
3
+ multiprocess==0.70.16
4
+ numpy==1.26.2
5
+ pluralizer==1.2.0
6
+ debugpy==1.6.7
7
+ nvidia-cudnn-cu11==8.5.0.96
8
+ deepspeed==0.15.2
9
+ data==0.4
10
+ pandas==2.1.3
11
+ tomli==2.0.1
12
+ charset-normalizer==3.3.2
13
+ attrs==24.2.0
14
+ aiosignal==1.3.1
15
+ fsspec==2023.10.0
16
+ nvidia-cusparse-cu11==11.7.4.91
17
+ zipp==3.12.0
18
+ mypy-extensions==1.0.0
19
+ datasets==3.0.1
20
+ joblib==1.3.2
21
+ hjson==3.1.0
22
+ traitlets==5.7.1
23
+ stack-data==0.6.0
24
+ transformers==4.45.1
25
+ sympy==1.11.1
26
+ Pygments==2.15.0
27
+ docker-pycreds==0.4.0
28
+ dill==0.3.8
29
+ wheel==0.44.0
30
+ prompt-toolkit==3.0.30
31
+ parso==0.8.3
32
+ ipykernel==6.23.1
33
+ pyarrow==17.0.0
34
+ certifi==2023.11.17
35
+ nvidia-cufft-cu11==10.9.0.58
36
+ six==1.16.0
37
+ pydantic==2.9.2
38
+ click==8.1.7
39
+ nest-asyncio==1.5.6
40
+ gmpy2==2.1.0
41
+ matplotlib==3.8.2
42
+ scipy==1.11.4
43
+ typing_extensions==4.12.2
44
+ statsmodels==0.14.0
45
+ huggingface-hub==0.25.0
46
+ frozenlist==1.4.1
47
+ gpustat==1.1.1
48
+ nvidia-nvtx-cu11==11.7.91
49
+ safetensors==0.4.5
50
+ stanza==1.9.2
51
+ decorator==5.1.1
52
+ seaborn==0.13.0
53
+ sentencepiece==0.2.0
54
+ PyYAML==6.0.1
55
+ black==24.8.0
56
+ protobuf==4.25.1
57
+ pickleshare==0.7.5
58
+ peft==0.13.0
59
+ triton==2.0.0
60
+ nvidia-cuda-runtime-cu11==11.7.99
61
+ Jinja2==3.1.2
62
+ nvidia-cusolver-cu11==11.4.0.1
63
+ executing==1.2.0
64
+ jupyter_client==8.1.0
65
+ pluggy==1.3.0
66
+ cmake==3.30.3
67
+ pytz==2023.3.post1
68
+ aiohappyeyeballs==2.4.2
69
+ kiwisolver==1.4.5
70
+ py-cpuinfo==9.0.0
71
+ Pillow==10.1.0
72
+ ptyprocess==0.7.0
73
+ importlib_resources==6.4.5
74
+ GitPython==3.1.43
75
+ importlib-metadata==6.0.0
76
+ iniconfig==2.0.0
77
+ scikit-learn==1.3.2
78
+ exceptiongroup==1.1.0
79
+ networkx==2.8.6
80
+ accelerate==1.0.0
81
+ nltk==3.8.1
82
+ shutilwhich==1.1.0
83
+ fonttools==4.45.1
84
+ future==0.18.3
85
+ aiohttp==3.10.6
86
+ wcwidth==0.2.5
87
+ idna==3.6
88
+ filelock==3.12.2
89
+ pathspec==0.12.1
90
+ jupyter_core==5.1.0
91
+ lit==18.1.8
92
+ nvidia-curand-cu11==10.2.10.91
93
+ nvidia-cublas-cu11==11.10.3.66
94
+ nvidia-ml-py==12.560.30
95
+ msgpack==1.1.0
96
+ python-dateutil==2.8.2
97
+ blessed==1.20.0
98
+ packaging==23.0
99
+ gitdb==4.0.11
100
+ yarl==1.13.0
101
+ emoji==2.8.0
102
+ tzdata==2023.3
103
+ cycler==0.12.1
104
+ tornado==6.2
105
+ backcall==0.2.0
106
+ plotnine==0.12.4
107
+ ninja==1.11.1.1
108
+ latex==0.7.0
109
+ wandb==0.18.5
110
+ setproctitle==1.3.3
111
+ threadpoolctl==3.2.0
112
+ requests==2.32.3
113
+ pyparsing==3.1.1
114
+ smmap==5.0.1
115
+ pyzmq==23.0.0
116
+ async-timeout==4.0.3
117
+ annotated-types==0.7.0
118
+ matplotlib-inline==0.1.6
119
+ latexcodec==1.0.0
120
+ ipython==8.0.0
121
+ patsy==0.5.3
122
+ contourpy==1.2.0
123
+ multidict==6.1.0
124
+ mizani==0.9.3
125
+ urllib3==2.1.0
126
+ tokenizers==0.20.0
127
+ MarkupSafe==2.1.2
128
+ pip==24.2
129
+ pexpect==4.8.0
130
+ tqdm==4.66.5
131
+ jedi==0.18.2
132
+ pydantic_core==2.23.4
133
+ tempdir==0.7.1
134
+ mpmath==1.2.1
135
+ setuptools==72.1.0
136
+ pytest==7.4.3
137
+ pure-eval==0.2.2
138
+ psutil==5.9.1
139
+ comm==0.1.2
140
+ nvidia-cuda-cupti-cu11==11.7.101
141
+ nvidia-cuda-nvrtc-cu11==11.7.99
142
+ regex==2023.10.3
143
+ platformdirs==2.5.2
144
+ asttokens==2.2.1
145
+ torch==2.0.0
146
+ nvidia-nccl-cu11==2.14.3
147
+ xxhash==3.5.0
wandb/run-20241105_160652-v9udw9ab/files/wandb-metadata.json ADDED
@@ -0,0 +1,97 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "os": "Linux-5.4.0-162-generic-x86_64-with-glibc2.31",
3
+ "python": "3.9.19",
4
+ "startedAt": "2024-11-05T21:06:52.157764Z",
5
+ "args": [
6
+ "--perturbation",
7
+ "shuffle_deterministic21",
8
+ "--train_set",
9
+ "10M",
10
+ "--batch_size",
11
+ "3",
12
+ "--epoch",
13
+ "3",
14
+ "--seed",
15
+ "0"
16
+ ],
17
+ "program": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py",
18
+ "codePath": "train/train_deep_wandb.py",
19
+ "git": {
20
+ "remote": "git@hf.co:Yaning1001/Impossible_llm.git",
21
+ "commit": "ed716cdcfcdea02b67f7ed0f3504c2b1c8b737c4"
22
+ },
23
+ "email": "yaning1001@gmail.com",
24
+ "root": "/mnt/ssd3/chunhui/yaning/project/impossible_llm/train",
25
+ "host": "mms-large-2",
26
+ "username": "chunhui",
27
+ "executable": "/mnt/ssd3/chunhui/miniconda/envs/impossible_llm/bin/python",
28
+ "codePathLocal": "train_deep_wandb.py",
29
+ "cpu_count": 32,
30
+ "cpu_count_logical": 64,
31
+ "gpu": "NVIDIA RTX A6000",
32
+ "gpu_count": 8,
33
+ "disk": {
34
+ "/": {
35
+ "total": "1888559353856",
36
+ "used": "1792542826496"
37
+ }
38
+ },
39
+ "memory": {
40
+ "total": "202617098240"
41
+ },
42
+ "cpu": {
43
+ "count": 32,
44
+ "countLogical": 64
45
+ },
46
+ "gpu_nvidia": [
47
+ {
48
+ "name": "NVIDIA RTX A6000",
49
+ "memoryTotal": "51527024640",
50
+ "cudaCores": 10752,
51
+ "architecture": "Ampere"
52
+ },
53
+ {
54
+ "name": "NVIDIA RTX A6000",
55
+ "memoryTotal": "51527024640",
56
+ "cudaCores": 10752,
57
+ "architecture": "Ampere"
58
+ },
59
+ {
60
+ "name": "NVIDIA RTX A6000",
61
+ "memoryTotal": "51527024640",
62
+ "cudaCores": 10752,
63
+ "architecture": "Ampere"
64
+ },
65
+ {
66
+ "name": "NVIDIA RTX A6000",
67
+ "memoryTotal": "51527024640",
68
+ "cudaCores": 10752,
69
+ "architecture": "Ampere"
70
+ },
71
+ {
72
+ "name": "NVIDIA RTX A6000",
73
+ "memoryTotal": "51527024640",
74
+ "cudaCores": 10752,
75
+ "architecture": "Ampere"
76
+ },
77
+ {
78
+ "name": "NVIDIA RTX A6000",
79
+ "memoryTotal": "51527024640",
80
+ "cudaCores": 10752,
81
+ "architecture": "Ampere"
82
+ },
83
+ {
84
+ "name": "NVIDIA RTX A6000",
85
+ "memoryTotal": "51527024640",
86
+ "cudaCores": 10752,
87
+ "architecture": "Ampere"
88
+ },
89
+ {
90
+ "name": "NVIDIA RTX A6000",
91
+ "memoryTotal": "51527024640",
92
+ "cudaCores": 10752,
93
+ "architecture": "Ampere"
94
+ }
95
+ ],
96
+ "cudaVersion": "11.8"
97
+ }
wandb/run-20241105_160652-v9udw9ab/files/wandb-summary.json ADDED
@@ -0,0 +1 @@
 
 
1
+ {"_wandb":{"runtime":5}}
wandb/run-20241105_160652-v9udw9ab/logs/debug-internal.log ADDED
@@ -0,0 +1,17 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {"time":"2024-11-05T16:06:52.159922257-05:00","level":"INFO","msg":"using version","core version":"0.18.5"}
2
+ {"time":"2024-11-05T16:06:52.159942657-05:00","level":"INFO","msg":"created symlink","path":"/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241105_160652-v9udw9ab/logs/debug-core.log"}
3
+ {"time":"2024-11-05T16:06:52.268171392-05:00","level":"INFO","msg":"created new stream","id":"v9udw9ab"}
4
+ {"time":"2024-11-05T16:06:52.268215922-05:00","level":"INFO","msg":"stream: started","id":"v9udw9ab"}
5
+ {"time":"2024-11-05T16:06:52.268269323-05:00","level":"INFO","msg":"writer: Do: started","stream_id":{"value":"v9udw9ab"}}
6
+ {"time":"2024-11-05T16:06:52.268270183-05:00","level":"INFO","msg":"handler: started","stream_id":{"value":"v9udw9ab"}}
7
+ {"time":"2024-11-05T16:06:52.268491884-05:00","level":"INFO","msg":"sender: started","stream_id":"v9udw9ab"}
8
+ {"time":"2024-11-05T16:06:52.481746757-05:00","level":"INFO","msg":"Starting system monitor"}
9
+ {"time":"2024-11-05T16:06:57.712126548-05:00","level":"INFO","msg":"stream: closing","id":"v9udw9ab"}
10
+ {"time":"2024-11-05T16:06:57.712223249-05:00","level":"INFO","msg":"Stopping system monitor"}
11
+ {"time":"2024-11-05T16:06:57.713709807-05:00","level":"INFO","msg":"Stopped system monitor"}
12
+ {"time":"2024-11-05T16:06:57.799517789-05:00","level":"ERROR","msg":"sender: sendDefer: failed to build job artifact","error":"failed to write data to file: write /tmp/tmpfile-3332236271: no space left on device"}
13
+ {"time":"2024-11-05T16:06:58.057468605-05:00","level":"INFO","msg":"fileTransfer: Close: file transfer manager closed"}
14
+ {"time":"2024-11-05T16:06:58.178861806-05:00","level":"INFO","msg":"handler: closed","stream_id":{"value":"v9udw9ab"}}
15
+ {"time":"2024-11-05T16:06:58.178915807-05:00","level":"INFO","msg":"writer: Close: closed","stream_id":{"value":"v9udw9ab"}}
16
+ {"time":"2024-11-05T16:06:58.178935557-05:00","level":"INFO","msg":"sender: closed","stream_id":"v9udw9ab"}
17
+ {"time":"2024-11-05T16:06:58.178979897-05:00","level":"INFO","msg":"stream: closed","id":"v9udw9ab"}
wandb/run-20241105_160652-v9udw9ab/logs/debug.log ADDED
@@ -0,0 +1,27 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 2024-11-05 16:06:52,154 INFO MainThread:1771273 [wandb_setup.py:_flush():79] Current SDK version is 0.18.5
2
+ 2024-11-05 16:06:52,154 INFO MainThread:1771273 [wandb_setup.py:_flush():79] Configure stats pid to 1771273
3
+ 2024-11-05 16:06:52,154 INFO MainThread:1771273 [wandb_setup.py:_flush():79] Loading settings from /home/chunhui/.config/wandb/settings
4
+ 2024-11-05 16:06:52,154 INFO MainThread:1771273 [wandb_setup.py:_flush():79] Loading settings from /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/settings
5
+ 2024-11-05 16:06:52,154 INFO MainThread:1771273 [wandb_setup.py:_flush():79] Loading settings from environment variables: {}
6
+ 2024-11-05 16:06:52,154 INFO MainThread:1771273 [wandb_setup.py:_flush():79] Applying setup settings: {'mode': None, '_disable_service': None}
7
+ 2024-11-05 16:06:52,154 INFO MainThread:1771273 [wandb_setup.py:_flush():79] Inferring run settings from compute environment: {'program_relpath': 'train/train_deep_wandb.py', 'program_abspath': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py', 'program': '/mnt/ssd3/chunhui/yaning/project/impossible_llm/train/train_deep_wandb.py'}
8
+ 2024-11-05 16:06:52,154 INFO MainThread:1771273 [wandb_setup.py:_flush():79] Applying login settings: {}
9
+ 2024-11-05 16:06:52,154 INFO MainThread:1771273 [wandb_init.py:_log_setup():534] Logging user logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241105_160652-v9udw9ab/logs/debug.log
10
+ 2024-11-05 16:06:52,155 INFO MainThread:1771273 [wandb_init.py:_log_setup():535] Logging internal logs to /mnt/ssd3/chunhui/yaning/project/impossible_llm/train/wandb/run-20241105_160652-v9udw9ab/logs/debug-internal.log
11
+ 2024-11-05 16:06:52,155 INFO MainThread:1771273 [wandb_init.py:init():621] calling init triggers
12
+ 2024-11-05 16:06:52,155 INFO MainThread:1771273 [wandb_init.py:init():628] wandb.init called with sweep_config: {}
13
+ config: {}
14
+ 2024-11-05 16:06:52,155 INFO MainThread:1771273 [wandb_init.py:init():671] starting backend
15
+ 2024-11-05 16:06:52,155 INFO MainThread:1771273 [wandb_init.py:init():675] sending inform_init request
16
+ 2024-11-05 16:06:52,156 INFO MainThread:1771273 [backend.py:_multiprocessing_setup():104] multiprocessing start_methods=fork,spawn,forkserver, using: spawn
17
+ 2024-11-05 16:06:52,157 INFO MainThread:1771273 [wandb_init.py:init():688] backend started and connected
18
+ 2024-11-05 16:06:52,160 INFO MainThread:1771273 [wandb_init.py:init():783] updated telemetry
19
+ 2024-11-05 16:06:52,189 INFO MainThread:1771273 [wandb_init.py:init():816] communicating run to backend with 90.0 second timeout
20
+ 2024-11-05 16:06:52,478 INFO MainThread:1771273 [wandb_init.py:init():867] starting run threads in backend
21
+ 2024-11-05 16:06:52,566 INFO MainThread:1771273 [wandb_run.py:_console_start():2463] atexit reg
22
+ 2024-11-05 16:06:52,566 INFO MainThread:1771273 [wandb_run.py:_redirect():2311] redirect: wrap_raw
23
+ 2024-11-05 16:06:52,566 INFO MainThread:1771273 [wandb_run.py:_redirect():2376] Wrapping output streams.
24
+ 2024-11-05 16:06:52,566 INFO MainThread:1771273 [wandb_run.py:_redirect():2401] Redirects installed.
25
+ 2024-11-05 16:06:52,568 INFO MainThread:1771273 [wandb_init.py:init():911] run started, returning control to user process
26
+ 2024-11-05 16:06:52,568 INFO MainThread:1771273 [wandb_run.py:_config_callback():1390] config_cb None None {'perturbation': 'shuffle_deterministic21', 'train_set': '10M', 'batch_size': 3, 'epoch': 3, 'seed': 0, 'lr': 5e-06}
27
+ 2024-11-05 16:06:57,712 WARNING MsgRouterThr:1771273 [router.py:message_loop():77] message_loop has been closed
wandb/run-20241105_160652-v9udw9ab/run-v9udw9ab.wandb ADDED
Binary file (2.29 kB). View file
 
wandb/run-20241105_162858-hqnfirxi/files/config.yaml ADDED
@@ -0,0 +1,49 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ _wandb:
2
+ value:
3
+ cli_version: 0.18.5
4
+ m: []
5
+ python_version: 3.9.19
6
+ t:
7
+ "1":
8
+ - 1
9
+ - 5
10
+ - 11
11
+ - 49
12
+ - 51
13
+ - 53
14
+ - 55
15
+ - 71
16
+ - 98
17
+ "2":
18
+ - 1
19
+ - 5
20
+ - 11
21
+ - 49
22
+ - 51
23
+ - 53
24
+ - 55
25
+ - 71
26
+ - 98
27
+ "3":
28
+ - 13
29
+ - 23
30
+ - 55
31
+ "4": 3.9.19
32
+ "5": 0.18.5
33
+ "6": 4.45.1
34
+ "8":
35
+ - 5
36
+ "12": 0.18.5
37
+ "13": linux-x86_64
38
+ batch_size:
39
+ value: 3
40
+ epoch:
41
+ value: 3
42
+ lr:
43
+ value: 5e-06
44
+ perturbation:
45
+ value: shuffle_deterministic57
46
+ seed:
47
+ value: 0
48
+ train_set:
49
+ value: 10M