microsoft/Fara1.5-4B | MXFP8
#43
by INC4AI - opened
Pipeline Failure Report
Model: microsoft/Fara1.5-4B
Quantization Scheme: MXFP8
Failed Phase: quantize
Run ID: Fara1.5-4B-AutoRound-MXFP8-Tuning
Error Category: out_of_memory
Full Error Log
11:56:04 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/microsoft/Fara1.5-4B/bb54c0c419a77eb156ff70bf887c290a17b66cd2/preprocessor_config.json "HTTP/1.1 200 OK"
11:56:04 [INFO] Starting quantization...
[transformers] `loss_type=None` was set in the config but it is unrecognized. Using the default loss: `ForCausalLMLoss`.
[38;20m2026-07-27 11:56:04 INFO utils.py L1026: Ignored layers: lm_head, lm_head[0m
[33;1m2026-07-27 11:56:04 WARNING utils.py L541: reset `quant_lm_head` to false as quantizing lm_head with tied weights has not been supported currently[0m
[38;20m2026-07-27 11:56:04 INFO data_driven.py L772: start to cache block inputs[0m
[38;20m2026-07-27 11:56:04 INFO mllm.py L83: Using MLLM template: qwen3_5[0m
[38;20m2026-07-27 11:56:04 INFO calib_dataset.py L977: Preprocessing calibration dataset in a subprocess to avoid memory leaks...[0m
11:56:05 [INFO] HTTP Request: HEAD https://huggingface.co/datasets/NeelNanda/pile-10k/resolve/main/README.md "HTTP/1.1 307 Temporary Redirect"
11:56:05 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/datasets/NeelNanda/pile-10k/127bfedcd5047750df5ccf3a12979a47bfa0bafa/README.md "HTTP/1.1 200 OK"
11:56:05 [INFO] HTTP Request: HEAD https://huggingface.co/datasets/NeelNanda/pile-10k/resolve/127bfedcd5047750df5ccf3a12979a47bfa0bafa/pile-10k.py "HTTP/1.1 404 Not Found"
11:56:06 [INFO] HTTP Request: HEAD https://s3.amazonaws.com/datasets.huggingface.co/datasets/datasets/NeelNanda/pile-10k/NeelNanda/pile-10k.py "HTTP/1.1 404 Not Found"
11:56:06 [INFO] HTTP Request: GET https://huggingface.co/api/datasets/NeelNanda/pile-10k/revision/127bfedcd5047750df5ccf3a12979a47bfa0bafa "HTTP/1.1 200 OK"
11:56:06 [INFO] HTTP Request: HEAD https://huggingface.co/datasets/NeelNanda/pile-10k/resolve/127bfedcd5047750df5ccf3a12979a47bfa0bafa/.huggingface.yaml "HTTP/1.1 404 Not Found"
11:56:06 [INFO] HTTP Request: GET https://datasets-server.huggingface.co/info?dataset=NeelNanda/pile-10k "HTTP/1.1 200 OK"
11:56:06 [INFO] HTTP Request: GET https://huggingface.co/api/datasets/NeelNanda/pile-10k/tree/127bfedcd5047750df5ccf3a12979a47bfa0bafa/data?recursive=true&expand=false "HTTP/1.1 200 OK"
11:56:06 [INFO] HTTP Request: GET https://huggingface.co/api/datasets/NeelNanda/pile-10k/tree/127bfedcd5047750df5ccf3a12979a47bfa0bafa?recursive=false&expand=false "HTTP/1.1 200 OK"
11:56:07 [INFO] HTTP Request: HEAD https://huggingface.co/datasets/NeelNanda/pile-10k/resolve/main/README.md "HTTP/1.1 307 Temporary Redirect"
11:56:07 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/datasets/NeelNanda/pile-10k/127bfedcd5047750df5ccf3a12979a47bfa0bafa/README.md "HTTP/1.1 200 OK"
11:56:07 [INFO] HTTP Request: HEAD https://huggingface.co/datasets/NeelNanda/pile-10k/resolve/127bfedcd5047750df5ccf3a12979a47bfa0bafa/pile-10k.py "HTTP/1.1 404 Not Found"
11:56:07 [INFO] HTTP Request: HEAD https://s3.amazonaws.com/datasets.huggingface.co/datasets/datasets/NeelNanda/pile-10k/NeelNanda/pile-10k.py "HTTP/1.1 404 Not Found"
11:56:07 [INFO] HTTP Request: GET https://huggingface.co/api/datasets/NeelNanda/pile-10k/revision/127bfedcd5047750df5ccf3a12979a47bfa0bafa "HTTP/1.1 200 OK"
11:56:08 [INFO] HTTP Request: HEAD https://huggingface.co/datasets/NeelNanda/pile-10k/resolve/127bfedcd5047750df5ccf3a12979a47bfa0bafa/.huggingface.yaml "HTTP/1.1 404 Not Found"
11:56:08 [INFO] HTTP Request: GET https://datasets-server.huggingface.co/info?dataset=NeelNanda/pile-10k "HTTP/1.1 200 OK"
11:56:08 [INFO] HTTP Request: GET https://huggingface.co/api/datasets/NeelNanda/pile-10k/tree/127bfedcd5047750df5ccf3a12979a47bfa0bafa/data?recursive=true&expand=false "HTTP/1.1 200 OK"
11:56:08 [INFO] HTTP Request: GET https://huggingface.co/api/datasets/NeelNanda/pile-10k/tree/127bfedcd5047750df5ccf3a12979a47bfa0bafa?recursive=false&expand=false "HTTP/1.1 200 OK"
[38;20m2026-07-27 11:56:08 INFO data_driven.py L795: caching done[0m
0%| | 0/32 [00:00<?, ?it/s]
Quantizing model.language_model.layers.0: 0%| | 0/32 [00:00<?, ?it/s]11:56:10 [ERROR] Quantization failed: CUDA out of memory. Tried to allocate 160.00 MiB. GPU 0 has a total capacity of 23.64 GiB of which 16.81 MiB is free. Process 3120243 has 23.62 GiB memory in use. Of the allocated memory 22.98 GiB is allocated by PyTorch, and 173.52 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)
Traceback (most recent call last):
File "/root/_work/1/s/auto_quant/phases/quantize.py", line 479, in <module>
quantize(args)
File "/root/_work/1/s/auto_quant/phases/quantize.py", line 370, in quantize
autoround.quantize()
File "/root/.venv/lib/python3.12/site-packages/auto_round/compressors/data_driven.py", line 837, in quantize
self._quantize_blocks(
File "/root/.venv/lib/python3.12/site-packages/auto_round/compressors/data_driven.py", line 659, in _quantize_blocks
self.pipeline.block_quantizer.quantize_block(ctx)
File "/root/.venv/lib/python3.12/site-packages/auto_round/algorithms/quantization/sign_round/quantizer.py", line 230, in quantize_block
pred_output = ctx.forward_block_batch(indices, device=device, cache_device=loss_device)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/auto_round/algorithms/pipeline.py", line 529, in forward_block_batch
return self.io.forward_block_batch(indices, device=device, cache_device=cache_device)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/auto_round/algorithms/pipeline.py", line 240, in forward_block_batch
output = self._run_block(block, quantizer, input_ids, input_others, device)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/auto_round/algorithms/pipeline.py", line 247, in _run_block
return quantizer._resolve_block_forward()(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/auto_round/compressors/utils.py", line 209, in block_forward
output = block(**input_others)
^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/transformers/modeling_layers.py", line 110, in __call__
return super().__call__(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/torch/nn/modules/module.py", line 1739, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/torch/nn/modules/module.py", line 1750, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/transformers/models/qwen3_5/modeling_qwen3_5.py", line 810, in forward
hidden_states = self.mlp(hidden_states)
^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/torch/nn/modules/module.py", line 1739, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/torch/nn/modules/module.py", line 1750, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/transformers/models/qwen3_5/modeling_qwen3_5.py", line 736, in forward
down_proj = self.down_proj(self.act_fn(self.gate_proj(x)) * self.up_proj(x))
^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/torch/nn/modules/module.py", line 1739, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/torch/nn/modules/module.py", line 1750, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/auto_round/wrapper.py", line 533, in forward
x, _, _ = self._qdq_act(
^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/auto_round/wrapper.py", line 304, in _qdq_act
x, scale, zp = self.act_quant_func(
^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/auto_round/data_type/mxfp.py", line 176, in quant_mx
tensor = quant_element(tensor, ebits, mbits, max_norm, mantissa_rounding)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/auto_round/data_type/mxfp.py", line 67, in quant_element
tensor = torch.sign(tensor) * (floor_ste(abs_tensor + 0.5) - mask_tensor)
^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/root/.venv/lib/python3.12/site-packages/auto_round/data_type/utils.py", line 333, in floor_ste
return (x.floor() - x).detach() + x
~~~~~~~~~~^~~
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 160.00 MiB. GPU 0 has a total capacity of 23.64 GiB of which 16.81 MiB is free. Process 3120243 has 23.62 GiB memory in use. Of the allocated memory 22.98 GiB is allocated by PyTorch, and 173.52 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)
Quantizing model.language_model.layers.0: 0%| | 0/32 [00:01<?, ?it/s]
Auto-generated by error_analysis pipeline. cc @lvkaokao