PRA-MEN · neural decompiler for Windows and MS-DOS
Pramen (Czech for source, spring) turns the assembly of a single function back into C. It targets the toolchains of the Windows / DOS world rather than Linux/GCC:
| variant | target |
|---|---|
msvc-x64-Od / msvc-x64-O2 |
Windows x64, MSVC, no optimization / /O2 |
msvc-x86-Od / msvc-x86-O2 |
Windows x86 (32-bit cdecl), MSVC |
dos16-small-od |
MS-DOS 16-bit real mode, small model, Open Watcom, no optimization |
dos16-large-ox |
MS-DOS 16-bit real mode, large model (far calls), Open Watcom -ox |
It is a 223 M parameter decoder-only transformer (GPT-2 style) trained from scratch – no natural language, just assembly and C, with its own 16 384-token tokenizer.
The output is a proposal. It usually compiles and often has the right structure, but names are invented (they are not in the binary), types are guessed and the logic has to be checked.
Results
Benchmark: HumanEval-C from LLM4Binary/decompile-eval (CC0), recompiled with the same toolchains as the training data. A task counts as solved when the decompiled C compiles and passes the original tests (re-executability). Greedy decoding, max 600 new tokens, 167 tasks for MSVC, 163 for Watcom.
| variant | re-executability |
|---|---|
| msvc-x64-Od | 40.7 % |
| msvc-x64-O2 | 13.2 % |
| msvc-x86-Od | 32.3 % |
| dos16-small-od | 30.1 % |
| dos16-large-ox | 11.7 % |
Optimized code is much harder. The most common failures are wrong parameter types (pointer vs. integer,
float vs. int) and invented helper names.
examples.json contains 12 classic functions × 6 variants with Pramen's outputs and test results
(47 / 72 pass; unoptimized 10 / 12 for each target) – they are also shown in the web UI.
Quick start
pip install huggingface_hub
hf download radix16/pramen --local-dir pramen # or: git clone https://huggingface.co/radix16/pramen
cd pramen
pip install -r requirements.txt # torch, tokenizers, safetensors
# command line (model is found in the current folder)
python -m pramen example-count_words-msvc-x64-Od.asm --variant msvc-x64-Od
# web UI with examples → http://127.0.0.1:8780
python -m pramen.web
A GPU is recommended (a few seconds per function). On CPU it works but is slow – generation does not use a KV cache yet.
from pramen.normalize import normalize
from pramen.runtime import Pramen
pramen = Pramen(".") # folder with config.json, model.safetensors, tokenizer.json
asm, fmt = normalize(open("func.asm").read(), "auto", "msvc-x86-Od")
print(pramen.decompile(asm, "msvc-x86-Od")["c"])
Getting the assembly
- MSVC:
dumpbin /disasm:nobytes program.obj– copy the block of one function. - Open Watcom:
wdis -a program.obj. - Ghidra: copy instructions from the Listing window (header, XREFs and bytes are dropped). Pramen was trained
on numeric stack offsets (
[EBP + -0x8]), so turn off Markup Stack Variable References (Edit → Tool Options → Listing Fields → Operands Field).
pramen.normalize converts all of these into the exact form used in training (upper-case mnemonics and
registers, 0x hex, LAB_n: labels). Feeding raw text that is formatted differently lowers the quality.
Why not GGUF / llama.cpp / LM Studio
The tokenizer uses its own pre-tokenization (hex literals like 0x10 and identifiers like LAB_2 stay whole).
llama.cpp only supports a fixed set of pre-tokenizers; with gpt-2 it splits 0x10 into 0, x, 10
and the model then repeats assembly instead of writing C. The weights themselves work in llama.cpp when fed
token IDs, but text prompts do not – so use the Python package here.
Prompt format
<|asm|>{variant}
{normalized assembly}
<|c|>
{C code}<|endoftext|>
Training
- Data: 5.26 M assembly → C pairs (2.33 B tokens). C functions come from
ExeBench (
train_synth_compilable,train_real_compilable), each compiled six ways with MSVC 19.36 (dumpbin) and Open Watcom V2 (wdis), deduplicated (exact + near). - Additional C/C++ source code (3.5 B tokens), mixed in so the model learns C syntax and idioms, from permissively licensed GitHub repositories (codeparrot/github-code-clean), The Stack, and several open-source projects (cJSON, fmt, nlohmann/json, and the GPL-2.0 source releases of Doom and Quake I–III).
- Architecture: 16 layers, 1 024 hidden, 16 heads, MLP 4 096, context 4 096, learned positions, tied embeddings.
Limitations
- one function at a time, up to ~3 500 input tokens (roughly 400–600 lines of assembly)
- globals and structs are only visible as addresses and offsets; the model guesses their meaning
- variable and function names are invented
- trained only on code produced by MSVC 19.36 and Open Watcom V2; other compilers (Borland, Turbo C, GCC/MinGW, …) are out of distribution
License and data
Model weights and code: Apache-2.0. The training functions keep the licenses of their original GitHub repositories (as collected by ExeBench); the additional source code includes GPL-2.0 code (Doom, Quake). Use the outputs accordingly – decompiled code of third-party software is still subject to its original license.
Česky
Pramen = zdroj. Neuronový dekompilátor, který z assembleru jedné funkce napíše C. Míří na Windows (MSVC x86/x64)
a MS-DOS (Open Watcom 16 bit). Model má 223 M parametrů a je natrénovaný od nuly jen na assembleru a C. Výstup je
návrh: obvykle se přeloží a mívá správnou strukturu, ale jména jsou vymyšlená, typy odhadnuté a logiku je potřeba
ověřit. Spuštění: pip install -r requirements.txt, pak python -m pramen.web (webové rozhraní s ukázkami na
http://127.0.0.1:8780) nebo python -m pramen soubor.asm --variant msvc-x86-Od.
- Downloads last month
- 128