dynrank-checkpoints / baselines /extract_v2_final.log
zhc12's picture
Upload baselines/extract_v2_final.log with huggingface_hub
b4b0a93 verified
Raw
History Blame Contribute Delete
4.05 kB
Loading model from /home/xiyaofeng/.cache/huggingface/hub/models--jeffwan--llama-7b-hf/snapshots/82eb0e6908390680598ca3ec1d77adfc5e1b24aa...
`torch_dtype` is deprecated! Use `dtype` instead!
2026-04-21 06:50:11.162455: I tensorflow/core/util/port.cc:153] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`.
2026-04-21 06:50:11.187966: E external/local_xla/xla/stream_executor/cuda/cuda_fft.cc:467] Unable to register cuFFT factory: Attempting to register factory for plugin cuFFT when one has already been registered
WARNING: All log messages before absl::InitializeLog() is called are written to STDERR
E0000 00:00:1776754211.209964 7450 cuda_dnn.cc:8579] Unable to register cuDNN factory: Attempting to register factory for plugin cuDNN when one has already been registered
E0000 00:00:1776754211.216595 7450 cuda_blas.cc:1407] Unable to register cuBLAS factory: Attempting to register factory for plugin cuBLAS when one has already been registered
W0000 00:00:1776754211.246581 7450 computation_placer.cc:177] computation placer already registered. Please check linkage and avoid linking the same target more than once.
W0000 00:00:1776754211.246606 7450 computation_placer.cc:177] computation placer already registered. Please check linkage and avoid linking the same target more than once.
W0000 00:00:1776754211.246610 7450 computation_placer.cc:177] computation placer already registered. Please check linkage and avoid linking the same target more than once.
W0000 00:00:1776754211.246613 7450 computation_placer.cc:177] computation placer already registered. Please check linkage and avoid linking the same target more than once.
2026-04-21 06:50:11.251103: I tensorflow/core/platform/cpu_feature_guard.cc:210] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations.
To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags.
AttributeError: 'MessageFactory' object has no attribute 'GetPrototype'
AttributeError: 'MessageFactory' object has no attribute 'GetPrototype'
AttributeError: 'MessageFactory' object has no attribute 'GetPrototype'
AttributeError: 'MessageFactory' object has no attribute 'GetPrototype'
AttributeError: 'MessageFactory' object has no attribute 'GetPrototype'
Loading checkpoint shards: 0%| | 0/2 [00:00<?, ?it/s] Loading checkpoint shards: 50%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 1/2 [00:01<00:01, 1.98s/it] Loading checkpoint shards: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 2/2 [00:02<00:00, 1.19s/it] Loading checkpoint shards: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 2/2 [00:02<00:00, 1.31s/it]
Model has 32 layers, 224 sublayers
Collecting calibration data (256 samples, seqlen=2048)...
Collecting covariances...
Calibration: 16/256 samples
Calibration: 32/256 samples
Calibration: 48/256 samples
Calibration: 64/256 samples
Calibration: 80/256 samples
Calibration: 96/256 samples
Calibration: 112/256 samples
Calibration: 128/256 samples
Calibration: 144/256 samples
Calibration: 160/256 samples
Calibration: 176/256 samples
Calibration: 192/256 samples
Calibration: 208/256 samples
Calibration: 224/256 samples
Calibration: 240/256 samples
Calibration: 256/256 samples
============================================================
Extracting svdllm_v2 profile at ratio=0.4
Done in 955.5s
Rank stats: min=309, max=2299, mean=1469.3, unique=176
Saved to outputs/baselines/svdllm_v2_r04.json
Budget: 3884330752 / 3884572672 (100.0%)
============================================================
Extracting svdllm_v2 profile at ratio=0.6
Done in 1103.0s
Rank stats: min=1, max=1941, mean=1030.0, unique=176
Saved to outputs/baselines/svdllm_v2_r06.json
Budget: 2721378560 / 2590064640 (105.1%)
All profiles saved to outputs/baselines/