coolblaze03's picture
docs: state the byte-exact certification ceiling β€” CPython 3.12 compiles the same AST to different bytecode depending on comprehension line-wrapping (measured: 80.62% of 516 stdlib modules preserve the fingerprint under a formatting-only rewrite). Costs the reported scores 0 rows; applies to any system scored this way.
3486ce0 verified
|
Raw
History Blame Contribute Delete
20.1 kB

PyBytecode v3 β€” full evaluation

The model card is deliberately short. This file is the long version: every number, the interval around it, the set it was measured on, and the command that reproduces it. Nothing here is softened relative to the card; the card is a subset.

Companion files: ORACLE-LIMITS.md for what the check can and cannot do, DATA-CARD-training-corpus.md for training-data lineage.


1. What "certified" means

The oracle is sound but incomplete:

verified   = PROVABLY correct. Recompiling the prediction yields a code object byte-identical to
             the target's. Identical code object => identical behaviour. No false positives.
unverified = UNKNOWN, not wrong. A correct decompilation that compiles differently β€” a `while`
             where the original had a `for`, a differently-ordered but equivalent boolean β€”
             does not certify.

Three consequences:

  • Reported accuracy is a lower bound on correctness. Everything counted here is proven; some of what is not counted is right too.
  • Best-of-N buys real accuracy, not a better guess. Sample until something certifies, then stop. The verifier decides when to stop, so extra budget cannot mislead you.
  • The check is exact. Comparison runs over the real code object, recursively, including docstrings and co_exceptiontable β€” not over a textual disassembly. That is not decoration: an earlier representation omitted the exception table's end, which made a try: body and a try/else: body indistinguishable, and it once certified a wrong prediction.

Strict L1 oracle throughout this file. Every figure carries a 95% interval and names its set.


2. Headline results β€” csn-3.12-licensed, n=600 across 117 repositories

certified 95% CI
PyBytecode v3, greedy 506 / 600 = 84.33% [80.48, 87.94]
PyBytecode v3, verified best-of-32 562 / 600 = 93.67% [90.86, 96.08]
Untuned Qwen2.5-Coder-1.5B-Instruct, greedy 4 / 600 = 0.67% [0.16, 1.35]

Intervals are repo-clustered (bootstrap over repositories, 10,000 resamples), because rows from one repository share an author and a house style and are not independent. Design effects 1.65 / 1.79 / 0.84.

The benchmark ships with the model, licence-resolved per row at the exact commit, so these numbers are recomputable from files in this repository rather than taken on faith.

The untuned base scores ~0, which is what makes the benchmark trustworthy

The same base model, given the same prompt, the same decoding budget and the same oracle, certifies 4 of 600 here and 0 of 400 on our earlier benchmark. Under looser oracles it is clearer still: the base produces syntactically valid Python on 52% of rows and reaches AST-identity on 0%. It paraphrases the disassembly into plausible-looking code with invented helpers rather than inverting it.

A strong general code model cannot guess its way to a single point on this task. The benchmark is not solvable by pattern-matching, and everything v3 scores was learned from the fine-tune. The paired difference is not close to the noise floor: exact McNemar on 502 discordant pairs in our favour and 0 against gives p < 1e-12.

The benchmark rebuild reproduced the result

The published 600-row set was built from scratch: different rows (only 85 of 600 appear in the earlier 400-row set), 117 repositories instead of 24, and a 1% per-repo cap replacing a 15% concentration. It lands in the same place β€” certified@32 93.67% against 93.25%, greedy 84.33% against 83.75%. The difference of +0.42pp carries a 95% interval of [βˆ’4.38, +5.22], so it is well inside noise, and the two size profiles are near-identical (median 59 vs 62 rep lines; 93.3% vs 94.0% of rows under 200), which is the composition factor that would otherwise move the number. Standardised to the older set's size mix the new set reads 94.48%. A result that survives a rebuild of the set it was measured on is worth more than the same result quoted twice.

Which numbers came from where

Set Published? Why
csn-3.12-licensed (600) yes every row licence-resolved at its commit
pybytecode-mbpp-3.12 (383) yes MBPP, CC-BY-4.0
CSN-3.12 (400) no 45% of its rows are not redistributable
Held-out (279) no per-row attribution was not retained

Our own results are measured on both the published set and the earlier pair. The figures in Comparison with other systems exist only on the earlier pair, because that is where those runs were done; they are labelled there because you cannot re-run them from what we shipped.


3. How much sampling budget to spend

The best-of-32 budget is 1 greedy decode + 31 sampled candidates at temperature 0.8, with verified early stop β€” sampling halts on the first candidate that certifies, so the mean cost is far below 32. Across the 94 greedy failures it took a mean of 15.16 samples and recovered 56:

budget certified
@1 (greedy alone) 506 / 600 84.33%
@2 525 / 600 87.50%
@4 546 / 600 91.00%
@8 555 / 600 92.50%
@16 558 / 600 93.00%
@32 562 / 600 93.67%

Most of the recovery arrives in the first few samples: @4 already captures two thirds of what @32 gets. Because early stop depends only on the index of the first passing sample, this curve is exact rather than an estimate β€” the same cached generations give every point.


4. Accuracy by input size β€” the full curve

The size axis is rep lines: the number of lines in the disassembly handed to the model. One line measures it:

from harness.pybytecode_core.rep import disassemble_v2
rep_lines = disassemble_v2(code_object).count("\n")

On the published 600-row benchmark

Recomputable from the files in this repo with harness/size_curve.py:

rep lines rows v3 greedy 95% CI v3 best-of-32 95% CI
0–49 224 95.98% [93.01, 98.51] 98.21% [96.31, 99.58]
50–99 224 89.73% [85.17, 93.93] 98.21% [96.26, 99.57]
100–199 112 65.18% [55.36, 74.14] 85.71% [77.57, 92.98]
200–299 27 51.85% β€” 77.78% β€”
300–399 5 60.00% β€” 100.00% β€”
400–599 5 0.00% β€” 0.00% β€”
600+ 3 0.00% β€” 0.00% β€”

Intervals are repo-clustered, and omitted below 30 rows / 10 repositories rather than printed at a width that would imply precision we do not have. The 300–399 bucket reading above the one below it is n=5 noise, not a recovery. Median input in this benchmark is 59 rep lines; p90 is 167, max 1,622.

On this set the top two buckets stay at 0% even with 32 samples β€” above ~400 rep lines, sampling did not rescue a single one of the 8 rows. 96.64% of greedy certifications and 95.37% of best-of-32 certifications come from units under 200 rep lines.

The four-row grouping on the model card is this same table pooled: under 100 rep lines is 416 / 448 = 92.86% greedy and 440 / 448 = 98.21% at 32; 200–399 is 17 / 32 = 53.12% and 26 / 32 = 81.25%; 400+ is 0 / 8 at both budgets.

Pooled over both earlier benchmarks (n=679)

Pooled so the thin upper buckets carry as many rows as we can give them. This is also the only table where PyLingual appears per bucket:

rep lines rows v3 greedy v3 best-of-32 untuned base PyLingual k=32
0–49 311 94.21% 98.71% 0.96% 99.04%
50–99 229 88.65% 96.94% 0.00% 95.63%
100–199 105 78.10% 88.57% 0.00% 84.76%
200–299 19 47.37% 84.21% 0.00% 84.21%
300–399 11 18.18% 54.55% 0.00% 81.82%
400–599 2 0.00% 50.00% 0.00% 100.00%
600+ 2 0.00% 0.00% 0.00% 50.00%
  • The greedy knee is at ~200 rep lines: 78.1% β†’ 47.4% β†’ 18.2% β†’ 0% across four consecutive buckets.
  • Best-of-32 postpones the knee to ~300; it does not remove it. Sampling 32 times buys roughly one bucket of headroom.
  • Almost every point we score comes from small units. On the earlier CSN-3.12 set (n=400), 97.61% of greedy certifications and 95.98% of best-of-32 certifications are units under 200 rep lines; on the published 600-row set it is 96.64% of greedy certifications. The headline accuracy is a statement about small units.
  • The top two buckets are n=2 each and carry nothing on their own. What carries the conclusion is the monotone decline through n=311/229/105/19/11 below them.

5. Comparison with other systems

PyLingual is another system that performs this task, by symbolic reconstruction rather than generation. We ran it on the same sets, under the same oracle, at a matched budget.

CSN-3.12, n=400, 24 repos certified 95% CI (repo-clustered)
untuned base, greedy 0 / 400 = 0.00% β€”
v3 greedy 335 / 400 = 83.75% [77.84, 88.89]
v3 certified@32 373 / 400 = 93.25% [88.82, 96.89]
PyLingual k=32 377 / 400 = 94.25% [91.06, 96.80]
union 393 / 400 = 98.25% [96.81, 99.51]
Held-out, n=279 certified 95% CI (Wilson)
untuned base, greedy 3 / 279 = 1.08% [0.37, 3.11]
v3 greedy 254 / 279 = 91.04% [87.11, 93.86]
v3 certified@32 272 / 279 = 97.49% [94.91, 98.78]
PyLingual k=32 267 / 279 = 95.70% [92.63, 97.52]
union 275 / 279 = 98.57% [96.37, 99.44]

Neither difference is statistically significant (exact paired McNemar, p = 0.6177 on CSN and p = 0.2266 on held-out). We claim no accuracy advantage, and none is claimed against us. On cost the comparison runs against us: ~1.5B parameters against their ~320M, and slower wall-clock.

Docstring recovery on the 115 held-out rows carrying a real docstring: 115 / 115 = 100% (PyLingual: 113/115 = 98.26%, under an oracle that does not score docstrings at all).

The held-out interval is Wilson, not clustered: every row of that set carries the same placeholder repository value, so its clustering cannot be assessed. The CSN design effects of 2.34–2.69 show why that matters β€” on a set where one repository supplied 15% of rows, a naive interval would have been far too narrow.

These figures were measured on our earlier benchmark pair, which is not published. PyLingual has not been run on the 600-row licensed benchmark released alongside this model.

The two approaches fail on different inputs

The more interesting measured result is not which number is larger, but that the two systems lose on different inputs, along an axis that turns out to be unit size:

rep lines rows PyBytecode greedy PyBytecode @32 PyLingual k=32
100–199 105 78.10% 88.57% 84.76%
200–299 19 47.37% 84.21% 84.21%
300–399 11 18.18% 54.55% 81.82%
400–599 2 0.00% 50.00% 100.0%

A generative model degrades as its input grows; a symbolic reconstruction does not. Two things follow, both useful:

  • Large units are not intrinsically unsolvable. A symbolic system handles them at 81.82% where we score 18.18%. What our curve measures is a capability limit of a 1.5B model, not a property of the task β€” which is a statement about where to spend the next round of scale.
  • The union exceeds either system, 98.25% and 98.57% above. Because verification is sound, the union costs nothing to take: run either, keep the answer that certifies, fall through to the other. Neither system can certify a wrong answer, so combining them cannot mislead you.

6. Limits, in full

  • Large units β€” this is the measured limit. Certification falls from 94% under 50 representation lines to 18% at 300–399 and 0% above 400. The curve in Β§4 is the honest specification of what this model does. It was trained on functions; a whole module is a different regime in both size and shape, and we have not evaluated one.

  • Python 3.12 only. A scope statement, not the explanation for the size behaviour above: trained and measured on 3.12, and the oracle refuses cross-minor input by design. Every number in this file is in-version, so nothing here is contaminated by a version mismatch.

  • It did not decompile real malware. On the one wild PyInstaller-packed sample in scope, the entry-point module produced nothing certifiable β€” the prediction did not even compile. 3 of 9 units in that sample certified, and those three were PyInstaller's own bootstrap plus a bundled stdlib module. Extraction and representation work (1,510/1,510 code objects recovered); end-to-end decompilation of real malware logic is not demonstrated.

    That entry point was both 3.13 and 491 representation lines, so size and version are confounded in it and it is evidence for neither on its own. It is reported because you should know the result, not because it diagnoses anything. The benign in-version curve in Β§4 is the instrument, and it already shows the model failing at that size on clean 3.12 input.

  • Annotated functions are handicapped: the training corpus was compiled with PEP-563 stringised annotations inherited, real .pyc files are not, and the benchmarks deliberately do not hand the model its training-time distribution back.

The oracle's limits, where you will meet them

Full detail in ORACLE-LIMITS.md.

  • 0.33% false-reject floor on foreign .pyc. On 600 wild install-time .pyc, 598 certify; 2 do not, because CPython 3.12.3 and 3.12.13 generate different code for the same source. No normalisation fixes that without abandoning the byte-identical guarantee. It fails to a false reject β€” you are told "unknown" about a correct answer, never "verified" about a wrong one.

  • Optimization level must match the producer's. A .pyc built with -O is a different code object. The wrong level does not degrade gracefully, it collapses to ~24%. Try all three; it costs three compiles. But note that at optimize>=1 docstrings are absent from the .pyc, so docstring recovery becomes unprovable against such a file β€” a limit of the artifact, not of the oracle.

  • The 100% pre-flight the harness prints proves nothing about soundness. It grades each reference against itself, i.e. compares compile(x) with compile(x), so any deterministic function of the source scores 100% β€” including a stub that ignores the bytecode entirely. It detects a broken harness and nothing more. The soundness evidence is the mutation test (0 true survivors in 1,239 mutants) and 18 targeted blind-spot probes (18/18).

  • Formatting the AST does not capture can still change the bytecode, so byte-exact certification has a ceiling on some code. CPython 3.12 inlines comprehensions (PEP 709) and emits a cleanup entry in the exception table for the inlined scope. How the comprehension is line-wrapped changes that entry. Two source texts with an identical AST, identical ast.unparse output and identical behaviour therefore compile to different code objects:

    def f(xs): return sum([1 for a in range(len(xs)) for b in range(a, len(xs)) if xs[b] == a])
    
    def f(xs): return sum([1 for a in range(len(xs)) for b in range(a, len(   # same AST,
        xs)) if xs[b] == a])                                                 # different bytecode
    

    Measured on this interpreter: over 516 standard-library modules, rewriting each through ast.unparse(ast.parse(src)) β€” a formatting-only change β€” preserves the compiled fingerprint for 416, 80.62%. Roughly one wild module in five compiles differently after a rewrite that changes no code. (Byte-identical text round-trip is 0.6%, as expected, since unparse normalises formatting by design. 80.62% is the figure that matters, because certification compares code objects, not text.)

    On the published benchmarks, an AST-identical re-wrap changes the fingerprint for 26 of 600 rows of csn-3.12-licensed (4.33%) and 20 of 383 rows of the MBPP set (5.22%). 12.8% and 19.1% of their rows contain a comprehension at all.

    This is a property of byte-exact certification, not of this model, and it applies to every system scored this way β€” including PyLingual in Β§5. A decompiler that recovers correct source but wraps it differently is recorded as failing on those inputs. It fails safe: the verdict is "not proven", never "verified" for a wrong answer.

    It costs the numbers in this file nothing measurable. Both benchmarks' reference text is ast.unparse output and the model reproduces that canonical single-line form, so of the 94 greedy failures on csn-3.12-licensed, 0 have an AST identical to the reference. The ceiling limits what certification can prove about differently-formatted source; it does not inflate or deflate the reported scores.

    Two consequences worth stating plainly. First, both published benchmarks are internally self-consistent: the reference source and the .pyc are compiled from the same normalised string β€” verified 600/600 and 383/383 β€” which is what makes the oracle sound. Second, that normalised string is not the upstream text. For the MBPP set, 1 of 383 rows (src/00117.py, MBPP task_id 757) compiles to different bytecode than the original MBPP solution, because that solution wraps a nested comprehension across lines. The row remains a valid decompilation task and its rewrite is declared in the dataset's provenance.transform and NOTICES.md; it is recorded here rather than quietly re-derived.

  • Untested: cross-minor 3.13, Nuitka, non-CPython builds, obfuscated bytecode.


7. Reproduce every number

No model, no GPU, no network, no third-party decompiler β€” CPython 3.12 standard library only:

cd harness
python3 grade.py --bench ../benchmarks/csn-3.12-licensed/bench.jsonl --self-test-only --out /tmp/st.json
python3 analyze_scores.py --bench ../benchmarks/csn-3.12-licensed/bench.jsonl \
    --greedy ../generations/gen_v3_csn600.jsonl \
    --base   ../generations/gen_base_csn600.jsonl \
    --out /tmp/scores.json --rows-out /tmp/rows.jsonl
python3 size_curve.py --bench ../benchmarks/csn-3.12-licensed/bench.jsonl \
    --greedy ../generations/gen_v3_csn600.jsonl \
    --base   ../generations/gen_base_csn600.jsonl \
    --out /tmp/size_curve.json

Per-row verdicts for all 600 rows β€” repo, function, commit SHA, SPDX, tuned verdict, base verdict β€” land in rows.jsonl. Every aggregate in Β§2 and Β§4 is recomputable from that file.

The figures in Β§5 are the exception: they were produced on the unpublished earlier benchmark pair and against a GPL-3.0 third-party decompiler we do not vendor, so they cannot be re-run from what ships here.


8. Training

Base Qwen/Qwen2.5-Coder-1.5B-Instruct (Apache-2.0)
Method LoRA r=16, Ξ±=32, all attention + MLP projections, 1 epoch, lr 2e-4, merged into the base
Rows 48,196 pairs of Python 3.12 disassembly β†’ source
Corpus codeparrot/github-code-clean, per-row filtered to 7 permissive licences; all GPL/LGPL/AGPL/MPL/EPL dropped before extraction
Decontamination shard-disjoint, repo-disjoint, and identifier-blind fingerprint-disjoint

The training corpus is not distributed. Per-row attribution was not retained during extraction, so the corpus cannot be redistributed without dropping required MIT/BSD/Apache notices. That is a property of the artifact, not of the licences β€” nothing in it is unlicensed or reciprocally licensed. Full lineage, including what was dropped and why: DATA-CARD-training-corpus.md.