sklearn-pmml-model: Uncontrolled Memory Allocation via SparseArray/@n (DoS)
Target: https://github.com/iamDecode/sklearn-pmml-model (PyPI: sklearn-pmml-model)
Platform: huntr (Model File Vulnerabilities, PMML format)
Class: CWE-789 (Memory Allocation with Excessive Size Value) / CWE-400 (Uncontrolled Resource Consumption)
Status: Gated PoC β access granted to protectai-bot for triage.
Summary
sklearn_pmml_model.base.parse_sparse_array() builds an in-memory Python list
sized directly from the attacker-controlled n attribute of a <SparseArray>
element in a PMML file, with no upper bound or sanity check against the
document's actual content:
# sklearn_pmml_model/base.py
def parse_sparse_array(array):
...
values = [0] * int(array.get('n')) # <-- unbounded allocation
indices = [int(i) - 1 for i in array.find('Indices').text.split(' ')]
...
Loading any PMML model that contains a SupportVectorMachineModel with a
VectorInstance whose vector is encoded as a SparseArray reaches this line
via sklearn_pmml_model.svm._base.get_vectors() -> parse_array(). This is
hit through the library's own public, documented API
(auto_detect_estimator(), PMMLSVC, PMMLSVR, PMMLNuSVC, PMMLNuSVR),
which is precisely the workflow the README recommends for loading a "trained
model exported to PMML" from any source β i.e., attacker-supplied model
files are the intended threat model for this library.
A ~1KB PMML file is enough to make the process attempt to allocate a multi-terabyte Python list, which either:
- raises
MemoryErrorinsideparse_sparse_arrayand aborts the caller (denial of service for that request/process), or - (with a smaller-but-still-huge
n) causes multi-second/minute stalls while CPython allocates and fills the list, tying up the loading thread/process.
There is no relationship enforced between n and the actual number of
Indices/Entries provided β the allocation happens before those are
even parsed, so a minimal, otherwise mostly-empty SparseArray triggers it.
Attacker input -> sink chain
- Attacker supplies a PMML file to any application that loads models via
sklearn_pmml_model.auto_detect.auto_detect_estimator(path)(the library's documented top-level API) or directly viasklearn_pmml_model.svm.PMMLSVR/PMMLSVC(pmml=path). auto_detect_estimatorsees<SupportVectorMachineModeland target type, and dispatches toPMMLSVR/PMMLSVC.PMMLBaseSVM.__init__(sklearn_pmml_model/svm/_base.py) iteratesVectorDictionary/VectorInstanceentries and callsget_vectors(vector_dictionary, s)for each support vector id.get_vectors()finds the vector'sSparseArray/REAL-SparseArrayelement and callsparse_array(array).parse_array()dispatchesSparseArraytags toparse_sparse_array(array).parse_sparse_array()executesvalues = [0] * int(array.get('n'))using the raw, attacker-controllednXML attribute value with no bound check β sink reached.
Proof of Concept
Files in this repo:
malicious_svr.pmmlβ 1031-byte PMML file: a minimalSupportVectorMachineModelwhose single support vector is a<SparseArray type="real" n="999999999999">.trigger_svm_api.pyβ loads the file viasklearn_pmml_model.svm.PMMLSVR(pmml=...).trigger_auto_detect_api.pyβ loads the file via the documented top-levelsklearn_pmml_model.auto_detect.auto_detect_estimator(...).
Reproduction (against an unmodified, freshly cloned/installed checkout)
git clone https://github.com/iamDecode/sklearn-pmml-model.git
python3 -m venv venv && source venv/bin/activate
pip install -e sklearn-pmml-model
# Bound virtual memory to 3GB purely so the demo fails fast/safely on the
# demonstrator's own machine; the bug itself has no dependency on any limit.
(ulimit -v 3000000; python3 trigger_auto_detect_api.py malicious_svr.pmml)
Observed (real run, this environment, sklearn-pmml-model @ master, April 2026 checkout)
Traceback (most recent call last):
File ".../trigger_auto_detect_api.py", line 7, in <module>
clf = auto_detect_estimator(sys.argv[1])
File ".../sklearn_pmml_model/auto_detect/base.py", line 40, in auto_detect_estimator
return auto_detect_regressor(pmml, **kwargs)
File ".../sklearn_pmml_model/auto_detect/base.py", line 129, in auto_detect_regressor
return parse(f)
File ".../sklearn_pmml_model/auto_detect/base.py", line 123, in parse
return reg(pmml, **kwargs)
File ".../sklearn_pmml_model/svm/_classes.py", line 295, in __init__
PMMLBaseSVM.__init__(self)
File ".../sklearn_pmml_model/svm/_base.py", line 56, in __init__
get_vectors(vector_dictionary, s) for s in self.support_
File ".../sklearn_pmml_model/svm/_base.py", line 109, in get_vectors
return np.array(parse_array(array))
File ".../sklearn_pmml_model/base.py", line 330, in parse_array
return parse_sparse_array(array)
File ".../sklearn_pmml_model/base.py", line 369, in parse_sparse_array
values = [0] * int(array.get('n'))
MemoryError raised as expected via auto_detect_estimator(). Traceback:
Without an external memory limit, the same file will make the hosting
process attempt to allocate ~8TB (999999999999 * 8 bytes for the pointer
array alone) until the OS OOM-kills the process or the allocation fails β
in either case the process handling the model-load crashes.
Impact
Any service that loads user- or third-party-supplied PMML files with
sklearn-pmml-model (e.g. a model-serving/inference API, an AutoML
pipeline accepting model uploads, a batch scoring job) can be crashed by
a single ~1KB file, with no valid support-vector data required. This is a
denial-of-service primitive delivered entirely through the model file β
the same class of "malicious model file" bug huntr's Model File
Vulnerability program targets for other formats (pickle, HDF5/Keras,
PyTorch, etc.), here affecting the PMML format via sklearn-pmml-model.
Fix suggestion
Bound n to a sane maximum (or to the number of Indices actually
present) before allocating, and/or catch MemoryError/OverflowError
around PMML parsing and re-raise as a controlled Exception, e.g.:
n = int(array.get('n'))
num_indices = len(array.find('Indices').text.split())
if n <= 0 or n > MAX_REASONABLE_SPARSE_ARRAY_SIZE or n < num_indices:
raise Exception('Invalid or excessive SparseArray size.')
values = [0] * n
Dedup check
- No GitHub issues, PRs, or Security Advisories in
iamDecode/sklearn-pmml-modelreferenceSparseArray, memory exhaustion, or denial of service (checked 2026-07-06 via GitHub search API and the repo's Security Advisories page β zero results). - No CVE/GHSA exists for
sklearn-pmml-modelin the GitHub Advisory Database as of 2026-07-06.
Reporter
Enigma Partners Global β authorized security research, huntr Model File Vulnerability program (target: PMML).