twanghcmut/backup-foundation-physics / logs /cosmos_wait_final.log
twanghcmut's picture
download
raw
5.5 kB
waiting for a GPU with >= 120 GB free (polling every 120s)
17:58:15 best: GPU0 with 27 GB free (need 120)
18:00:15 best: GPU0 with 27 GB free (need 120)
18:02:15 best: GPU0 with 27 GB free (need 120)
18:04:16 best: GPU0 with 27 GB free (need 120)
18:06:16 best: GPU0 with 138 GB free (need 120)
launching on GPU0
spec: /srv/data/foundation-physics-graph-model/outputs/cosmos_input/spec_depth_vis.json
output: /srv/data/foundation-physics-graph-model/outputs/cosmos_runs/FINAL_depth_vis_720
devices: 0
hf_token: hf_AvnPC... (the gated downloads must use this one)
[07-31 18:06:18|INFO|packages/cosmos-oss/cosmos_oss/init.py:96:_init_log_files] Log saved to /srv/data/foundation-physics-graph-model/outputs/cosmos_runs/FINAL_depth_vis_720/console.log
[07-31 18:06:39|INFO|cosmos_transfer2/_src/imaginaire/utils/checkpoint_db.py:297:download] Downloading checkpoint Wan2.1/vae(685afcaa-4de2-42fe-b7b9-69f7a2dee4d8)
[07-31 18:06:39|INFO|cosmos_transfer2/_src/imaginaire/utils/checkpoint_db.py:165:_hf_download] uvx 'hf>=1.3.5' download nvidia/Cosmos-Predict2.5-2B --repo-type model --revision f176dc95b4a70f53ce01c4b302851595e7322b00 tokenizer.pth
path=/home/quang/.cache/huggingface/hub/models--nvidia--Cosmos-Predict2.5-2B/snapshots/f176dc95b4a70f53ce01c4b302851595e7322b00/tokenizer.pth
[07-31 18:06:42|INFO|cosmos_transfer2/_src/imaginaire/utils/checkpoint_db.py:297:download] Downloading checkpoint Qwen/Qwen2.5-VL-7B-Instruct(7219c6c7-f878-4137-bbdb-76842ea85e70)
[07-31 18:06:42|INFO|cosmos_transfer2/_src/imaginaire/utils/checkpoint_db.py:165:_hf_download] uvx 'hf>=1.3.5' download nvidia/Cosmos-Reason1-7B --repo-type model --revision 3210bec0495fdc7a8d3dbb8d58da5711eab4b423 --include '*'
path=/home/quang/.cache/huggingface/hub/models--nvidia--Cosmos-Reason1-7B/snapshots/3210bec0495fdc7a8d3dbb8d58da5711eab4b423
Traceback (most recent call last):
File "/srv/data/foundation-physics-graph-model/third_party/cosmos-transfer2.5/examples/inference.py", line 94, in <module>
main(args)
File "/srv/data/foundation-physics-graph-model/third_party/cosmos-transfer2.5/examples/inference.py", line 77, in main
inference = Control2WorldInference(args.setup, batch_hint_keys=batch_hint_keys)
File "/srv/data/foundation-physics-graph-model/third_party/cosmos-transfer2.5/cosmos_transfer2/inference.py", line 124, in __init__
self.inference_pipeline = ControlVideo2WorldInference(
File "/srv/data/foundation-physics-graph-model/third_party/cosmos-transfer2.5/cosmos_transfer2/_src/transfer2/inference/inference_pipeline.py", line 133, in __init__
model, config = load_model_predict2(
File "/srv/data/foundation-physics-graph-model/third_party/cosmos-transfer2.5/cosmos_transfer2/_src/predict2/utils/model_loader.py", line 117, in load_model_from_checkpoint
model = instantiate(config.model)
File "/srv/data/foundation-physics-graph-model/third_party/cosmos-transfer2.5/cosmos_transfer2/_src/imaginaire/lazy_config/instantiate.py", line 115, in instantiate
return cls(*args, **instantiate_kwargs)
File "/srv/data/foundation-physics-graph-model/third_party/cosmos-transfer2.5/cosmos_transfer2/_src/transfer2/models/vid2vid_model_control_vace_rectified_flow.py", line 75, in __init__
super().__init__(config, *args, **kwargs)
File "/srv/data/foundation-physics-graph-model/third_party/cosmos-transfer2.5/cosmos_transfer2/_src/predict2/models/text2world_model_rectified_flow.py", line 167, in __init__
self.text_encoder = TextEncoder(self.config.text_encoder_config)
File "/srv/data/foundation-physics-graph-model/third_party/cosmos-transfer2.5/cosmos_transfer2/_src/predict2/text_encoders/text_encoder.py", line 77, in __init__
self.model.to_empty(device=self.device)
File "/srv/data/foundation-physics-graph-model/third_party/cosmos-transfer2.5/.venv/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1207, in to_empty
return self._apply(
File "/srv/data/foundation-physics-graph-model/third_party/cosmos-transfer2.5/.venv/lib/python3.10/site-packages/torch/nn/modules/module.py", line 915, in _apply
module._apply(fn)
File "/srv/data/foundation-physics-graph-model/third_party/cosmos-transfer2.5/.venv/lib/python3.10/site-packages/torch/nn/modules/module.py", line 942, in _apply
param_applied = fn(param)
File "/srv/data/foundation-physics-graph-model/third_party/cosmos-transfer2.5/.venv/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1208, in <lambda>
lambda t: torch.empty_like(t, device=device), recurse=recurse
File "/srv/data/foundation-physics-graph-model/third_party/cosmos-transfer2.5/.venv/lib/python3.10/site-packages/torch/_prims_common/wrappers.py", line 308, in _fn
result = fn(*args, **kwargs)
File "/srv/data/foundation-physics-graph-model/third_party/cosmos-transfer2.5/.venv/lib/python3.10/site-packages/torch/_refs/__init__.py", line 4999, in empty_like
return torch.empty_permuted(
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 1.02 GiB. GPU 0 has a total capacity of 139.81 GiB of which 103.81 MiB is free. Including non-PyTorch memory, this process has 27.65 GiB memory in use. Of the allocated memory 26.85 GiB is allocated by PyTorch, and 211.72 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)

Xet Storage Details

Size:
5.5 kB
·
Xet hash:
3968ad7ab45f5efe12d4e51cb7b51e470fca1a8c63903cdaeb73e63b71ca3320

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.