YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
sklearn-pmml-model: Incorrect SparseArray Indices/Entries pairing in parse_sparse_array() -> IndexError DoS or silent model-output corruption
Target repo: https://github.com/iamDecode/sklearn-pmml-model
Version confirmed: v1.0.8 (PyPI), verified byte-identical to current GitHub master sklearn_pmml_model/base.py via diff against a freshly fetched copy of raw.githubusercontent.com.
CWE: CWE-129 (Improper Validation of Array Index) / incorrect array index usage.
Reported: 2026-07-12
Summary
parse_sparse_array() in sklearn_pmml_model/base.py mis-pairs the <Indices>
and <Entries> elements of a PMML <SparseArray>. Per the PMML/DMG spec,
position i of <Indices> names the destination slot for position i of
<Entries> (i.e. the correct code is values[indices[i]] = entries[i], e.g.
via zip/enumerate). The actual code instead reuses the value of each
index as if it were also a valid position into entries:
for index in indices:
values[index] = entries[index]
This is wrong for essentially any non-trivial sparse array (any case where
the index values are not exactly the sequence 0,1,2,...,len(entries)-1),
which is the entire point of using a SparseArray instead of a dense
Array in the first place.
This bug is reached through the library's own public, documented API
(sklearn_pmml_model.auto_detect.auto_detect_estimator(), and the
PMMLSVR/PMMLSVC classes that use it), the same call path as an
already-filed, separate SparseArray @n uncontrolled-allocation finding.
This is a distinct defect: different root cause (indexing/pairing logic,
not pre-allocation size), different line (base.py ~line 395 vs. the @n
allocation site), and not fixed by that finding's suggested remediation
(bounding/capping @n) β this bug fires or silently corrupts data
regardless of how small @n is.
Root cause
File: sklearn_pmml_model/base.py, function parse_sparse_array()
(approx. lines 368-396 of current master):
values = [0] * int(array.get('n'))
indices = [int(i) - 1 for i in array.find('Indices').text.split(' ')]
...
entries = [float(x) for x in element.text.split(' ')] # (or int, per type)
...
for index in indices:
values[index] = entries[index]
indices holds the 0-based (converted from PMML's 1-based) target
positions named by <Indices>. entries is a separate, independently
sized list parsed from <Entries> / <NUM-Entries> / <REAL-Entries> /
<INT-Entries>. The code indexes entries using the index value itself
rather than the position within the indices list, so:
- If any index value is
>= len(entries),entries[index]raises an uncaughtIndexError. - If index values are within range of
len(entries)but not exactly0, 1, 2, ..., len(entries)-1in order, the wrong entries silently land in the wrong slots β no exception, but corrupted values.
Two security-relevant, independently reachable outcomes:
- IndexError crash / DoS β whenever any index value is
>= len(entries)(the normal case for genuinely sparse data: a large feature space with few non-zero entries), parsing raises an uncaughtIndexErrorand aborts the caller. No huge-allocation trick needed βncan be a completely ordinary, small value. - Silent model-output corruption β whenever index values happen to stay in range but are not in strictly sequential order, the resulting support-vector (or coefficient) values are silently wrong. An attacker-crafted PMML file loads "successfully" but yields corrupted predictions at inference time, undermining the integrity of the loaded model with no visible error.
Proof of Concept
Three minimal (~1KB) SupportVectorMachineModel PMML files, each with a
single support vector encoded as a <SparseArray>, loaded via the
library's public, documented API,
sklearn_pmml_model.auto_detect.auto_detect_estimator().
1. wrong_indices.pmml β IndexError DoS
<SparseArray type="real" n="5">
<Indices>1 3 5</Indices>
<Entries>10 20 30</Entries>
</SparseArray>
Real, captured output loading this file:
Traceback (most recent call last):
File "<string>", line 6, in <module>
clf = auto_detect_estimator('/home/kali/hunt-workspace/pmml-index-bug/wrong_indices.pmml')
File "/home/kali/pmml-verify-build/sklearn-pmml-model/sklearn_pmml_model/auto_detect/base.py", line 40, in auto_detect_estimator
return auto_detect_regressor(pmml, **kwargs)
File "/home/kali/pmml-verify-build/sklearn-pmml-model/sklearn_pmml_model/auto_detect/base.py", line 129, in auto_detect_regressor
return parse(f)
File "/home/kali/pmml-verify-build/sklearn-pmml-model/sklearn_pmml_model/auto_detect/base.py", line 123, in parse
return reg(pmml, **kwargs)
File "/home/kali/pmml-verify-build/sklearn-pmml-model/sklearn_pmml_model/svm/_classes.py", line 295, in __init__
PMMLBaseSVM.__init__(self)
File "/home/kali/pmml-verify-build/sklearn-pmml-model/sklearn_pmml_model/svm/_base.py", line 56, in __init__
get_vectors(vector_dictionary, s) for s in self.support_
File "/home/kali/pmml-verify-build/sklearn-pmml-model/sklearn_pmml_model/svm/_base.py", line 109, in get_vectors
return np.array(parse_array(array))
File "/home/kali/pmml-verify-build/sklearn-pmml-model/sklearn_pmml_model/base.py", line 330, in parse_array
return parse_sparse_array(array)
File "/home/kali/pmml-verify-build/sklearn-pmml-model/sklearn_pmml_model/base.py", line 395, in parse_sparse_array
values[index] = entries[index]
IndexError: list index out of range
Exception: <class 'IndexError'> list index out of range
2. silent_corruption.pmml β silent model-output corruption
<SparseArray type="real" n="3">
<Indices>2 3 1</Indices>
<Entries>10 20 30</Entries>
</SparseArray>
Real, captured output loading this file:
Loaded OK. support_vectors_ = [[10. 20. 30.]]
Expected per PMML spec (Indices 1-based positional pairing with Entries): [[30, 10, 20]]
No exception is raised, but the loaded support-vector values are wrong, silently corrupting any predictions made with this model.
3. sequential_control.pmml β negative control
<SparseArray type="real" n="3">
<Indices>1 2 3</Indices>
<Entries>10 20 30</Entries>
</SparseArray>
Real, captured output loading this file:
Negative control (sequential Indices=1,2,3): support_vectors_ = [[10. 20. 30.]] -- correct, no bug visible here
This case loads correctly, because when the index value happens to equal
its position, the buggy entries[index] lookup accidentally coincides
with the correct entries[position] lookup. This is why ordinary/unit-test
PMML fixtures β which tend to use sequential toy data β would not have
surfaced this bug.
Verification environment
- Real execution against a fresh checkout of
iamDecode/sklearn-pmml-model(v1.0.8) at/home/kali/pmml-verify-build/sklearn-pmml-model. - Confirmed byte-identical to current GitHub
master'ssklearn_pmml_model/base.pyviadiffagainst a freshly fetched copy fromraw.githubusercontent.com(included here aslatest_base.pyfor identity confirmation). - All three PoC files loaded through the library's own public,
documented API entry point,
auto_detect_estimator()β no internal/private functions called directly.
Dedup check
GitHub code/issue search API against iamDecode/sklearn-pmml-model for
"SparseArray" and "IndexError" both returned total_count: 0.
repos/.../security-advisories returned []. Checked 2026-07-12.
This finding is a distinct defect from an already-filed SparseArray @n
uncontrolled-allocation finding against the same project: different root
cause (Indices/Entries pairing logic vs. unbounded pre-allocation of the
values list), different line, and not remediated by that finding's fix
(bounding @n) since this bug is independent of the size of @n.
Suggested fix
Pair indices and entries positionally, per the PMML spec, e.g.:
for pos, index in enumerate(indices):
values[index] = entries[pos]
(equivalently, for index, entry in zip(indices, entries): values[index] = entry),
and validate len(indices) == len(entries) and that every index is within
bounds of values, raising a clear ValueError/Exception rather than an
uncaught IndexError for malformed input.
Files in this repo
wrong_indices.pmmlβ PoC 1 (IndexError DoS)silent_corruption.pmmlβ PoC 2 (silent corruption)sequential_control.pmmlβ PoC 3 (negative control)latest_base.pyβ freshly fetched copy of current GitHub mastersklearn_pmml_model/base.py, used to confirm the verified build matches the currently released code.