Spaces:
Running on Zero
Running on Zero
File size: 22,441 Bytes
9f29ecf fd7251d 9f29ecf fd7251d 9f29ecf 8817f63 9f29ecf fd7251d 9f29ecf cba4874 fd7251d e4464d0 7247dcd e4464d0 fd7251d f5c6cec d80154b f5c6cec 0e69d49 8052148 fd7251d 8052148 fd7251d 8052148 fd7251d 8052148 0e69d49 346a1ad 8052148 fd7251d d113a7f fd7251d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 | ---
title: PDF Injection Detector (MiMo-7B)
emoji: π
colorFrom: yellow
colorTo: gray
sdk: gradio
sdk_version: 5.50.0
app_file: app.py
pinned: false
license: mit
short_description: MiMo-7B finds payloads hidden inside PDF files
---
# PDF Injection Detector β MiMo-7B
Upload a PDF. It is rendered to text with the extractor that built the project corpus, the regions
carrying structural markers are ranked, and **MiMo-7B-RL** reads the most promising ones and says
whether a payload is hidden there β naming the family and quoting the substring that convinced it.
It is a coursework artefact built on a synthetic corpus of 1,100 PDFs carrying harmless
EICAR/AMTSO/WICAR/RANSIM test markers. **It is not a general malware scanner.**
## How to use it
Live Space: [BentoUniAcc/Mimo_Injection_detector](https://huggingface.co/spaces/BentoUniAcc/Mimo_Injection_detector).
Every screenshot below is of that Space as it actually runs β nothing is mocked up.
### 1. Open the Space

The page comes up before anything is downloaded. The header states which runtime is live and roughly
what a region costs on it β **GPU, ~4.2 s per region** above. Read the disclaimer paragraph: the
verdict, the family and the treatment are all fallible, and the family is wrong more often than it
is right.
### 2. Give it a PDF
Either drag a file onto the **PDF** box, or press one of the **Examples** buttons down the left
edge β one per attack family, plus a clean control. An example just drops its file into the same
uploader, so it is triaged by exactly the same code path as your own upload.

Note the uploader shows `Example_5.pdf`, not the family name. The button is labelled; the file is
not. Nothing about the document announces its own answer, and a filename never reaches the model.
### 3. Read the plan, then set the controls
Under the button, the app says what it found and what it intends to read *before* any model runs β
that panel is broken down in [the next section](#where-the-plan-panels-numbers-come-from).

| Control | What it does | When to change it |
|---|---|---|
| **Runtime** | `gpu` (4-bit NF4, Part B's configuration, ~4 s/region, capped by your daily ZeroGPU quota) or `cpu` (Q4_K_M GGUF, ~2 min/region, no quota). Hidden when only one backend is installed β as on this Space. | Switch to `cpu` when the GPU quota runs out. |
| **Regions per batch** | How many regions one run of the model reads. It only re-cuts the same list, so a lower number means more batches, not less work. | Leave it at the maximum. A run reserves the same GPU time whatever this is, so lowering it inspects less of the file for the same quota. |
| **Batch to check** | Which batch this run spends. Each label states how many regions, whether they are marker or sweep, where in the skeleton they sit, and the estimated time. | Batch 1 is the most marker-dense, so start there and spend further runs on later batches if you want the rest read. |
| **Sweep the rest of the document too** | On, the batches tile the whole skeleton. Off, only marker regions are read β which is the exact shape Part B measured. | Turn it off to stay strictly inside the measured configuration; leave it on for coverage of text no marker points at. |
| **Also show the nearest files in the corpus** | Runs Part A's embedding lookup after the scan. Costs a one-off 550 MB model download. | Turn it off to save the download. |
### 4. Press **Check this batch**
The button locks itself for the duration and any previous result is cleared, so nothing on screen
can belong to a different file. The first scan after a restart also fetches the weights, which is
the slow part.

One press is one run of the model over one batch. Press it again β on a different batch β to spend
another run; the app never scans the whole file behind your back.
### 5. Read the four tabs
- **Report** β the verdict for this batch. If MiMo flagged anything it says so, names the family it
guessed, and prints the containment advice for that family. It always states how many regions in
other batches are still unread, and how many answers could not be parsed (those count as *not
injected*, exactly as Part B scored them).
- **Regions read** β one row per region: where in the skeleton, `marker` or `sweep`, which
signatures the regex saw, MiMo's verdict, the family, and the substring quoted as evidence.
- **Nearest corpus files** β the five most similar corpus documents to the first flagged region.
Precision@5 is 35.6%, so read it as *resemblance*, never identification.
- **What MiMo actually said** β the untouched generation per region, and which prompt route
produced it. This is where to look when a verdict seems wrong.
#### The nearest-files tab is a recommender, and the score is cosine similarity
That tab is the project's recommendation system, and it is content-based rather than
collaborative β there are no users to learn from, only documents. The flagged region is embedded
with Part A's winning model (`nomic-ai/nomic-embed-text-v1.5`, the `search_document: ` prefix,
768 dimensions) into the same space as the 1,100 corpus files, and the five most similar are
returned.
Both the index and the query are **unit-normalised**, so the dot product *is* the cosine
similarity, and the whole lookup is a single 1,100 Γ 768 matrixβvector product β no approximate
index or clustering is needed at this size, and the ranking is exact. The number in the *cosine
similarity* column is that value: 1.0 is identical direction, 0 is unrelated. It measures how
alike two payload windows *read*, which is not the same as being the same attack β hence
precision@5 of 35.6% against a 6.8% random baseline, and the word *resemblance* rather than
*identification*. `check_provenance()` asserts the model, prefix, dimension, normalisation and
source column against Part A's own results file before any lookup runs, because a query embedded
differently lands in a different space and returns confident nonsense with no error anywhere.
### Quotas, and what "MiMo could not run" means

A free visitor gets roughly **five minutes of ZeroGPU per day**, and the scheduler reserves a run's
full requested duration up front rather than what it turns out to use β so two or three batches can
exhaust the day. The quota is counted per visitor, not per Space: this message means your allowance,
not a broken app. The report says exactly how long until it resets. Where a CPU backend is installed,
switching the runtime to `cpu` keeps working with no quota at all; on this Space it is not installed
(see *Two runtimes* below), so the options are to wait, or to sign in to Hugging Face for a larger
allowance.
## Where the plan panel's numbers come from
This block appears under the button as soon as a file is loaded, and it is written **before any
model runs** β it is pure text extraction and regex, computed in `corpus_text.py` on every upload,
every slider move and every checkbox change.

**Line 1 β the extraction.**
- **`12,729 bytes` on disk** β the length of the uploaded file, unmodified.
- **`5,255-character skeleton`** β the result of `build_skeleton()`. Every `stream β¦ endstream`
body in the PDF is inflated with `zlib` where it is FlateDecode, capped at 4,096 characters, and
kept if more than 60% of its bytes are printable ASCII. Control and binary residue is replaced
with spaces and runs of four or more spaces collapse to three. So the number is smaller than the
file for a compressed text document and much smaller for one full of images.
- **`4 binary stream(s) dropped`** β streams that failed the 60%-printable test, i.e. genuine
binary such as an image. Each is replaced in the skeleton by the literal `<<BINARY n bytes>>`
rather than deleted, so its position is preserved. This count is exactly how many times that
happened.
- If the file is very large the line also says **truncated**: past a 120,000-character budget the
skeleton keeps 45% from the head and the rest from the tail, because the corpus generator inserts
payloads at both ends and a plain head truncation would lose most of them.
**Line 2 β the triage.** `ANY_MARKER_RE` β a single alternation of all twelve families' structural
patterns plus the framework and test-payload strings β is run over the whole skeleton.
- **`1 region(s) carry a marker`** β every match becomes a Β±1,500-character window centred on the
hit, and overlapping windows are merged so two markers 200 characters apart are read once rather
than twice. This is the count after merging, so it is regions, not raw matches.
- **`the remaining 1 cover the rest of the document`** β the sweep. The skeleton is tiled into
3,000-character windows in document order, and a tile is dropped if a marker region already
covers half of it or more. This is what stops "clean" meaning "the model never looked there".
With the sweep checkbox off, this part is absent and only the marker regions exist.
- **`2 in total, cut into 1 batch(es) of at most 8`** β the two lists are concatenated, marker
regions first in density order, and the result is sliced into batches of the size the *Regions
per batch* slider is set to. Move the slider and this line re-computes immediately; the regions
themselves do not change, only how they are grouped.
- **`Structural signatures in the raw file: javascript_injection, ransomware_simulation`** β which
families' regexes matched, run by `detect_markers()` against the **raw bytes**, not the skeleton.
Two names appear here because the generator gives several families a JavaScript launcher, so a
ransomware sample legitimately carries a JavaScript signature too.
Two things this panel deliberately does not do. It does not say `javascript_injection` on its own
as though that were a finding β plenty of harmless PDFs contain a `/JS (` token, and the word
*signature* is doing real work in that sentence. And the ordering it produces decides **reading
order only**: batch 1 is the most marker-dense, not the guilty one. The verdict below it is MiMo's
alone, which is what the italic line under the counts is there to say.
## The three repos this is built on
| Repo | What this Space takes from it |
|---|---|
| [Generated_Injected_PDFs_HARMLESS](https://huggingface.co/datasets/Cyber-security-final-project/Generated_Injected_PDFs_HARMLESS) | The 12 injection families and their structural signatures β the definition of what an attack looks like |
| [HARMLESS_Synthetic_Injected_PDFs_EDA](https://huggingface.co/datasets/Cyber-security-final-project/HARMLESS_Synthetic_Injected_PDFs_EDA) | `build_skeleton`, `mask_leaks` and `payload_window` β how a PDF becomes the text a model reads |
| [Evaluation_of_OpenSource_Modelsβ¦](https://huggingface.co/datasets/Cyber-security-final-project/Evaluation_of_OpenSource_Models_for_PDF_Injection_Recognition) | Part A's embedding index and winning configuration; Part B's prompt, prefill and parser |
The prompt, the MiMo prefill, the brace-counting parser and the whole text-extraction path are
quoted **verbatim** from those notebooks. That is the correctness argument for quoting Part B's
scores here at all: change how the text is extracted or how the question is asked, and the
published numbers stop describing this program.
## The examples
Thirteen example documents ship with the Space β **one per injection family, plus one clean
control** β taken from the generation repo. They appear as a rail of buttons down the left edge,
labelled by attack type: press **ransomware simulation** and that document loads into the uploader
and is triaged exactly as an uploaded file would be. There is no separate code path for examples
and nothing about them is pre-computed.
**The files themselves stay anonymous.** On disk they are `Example_0` β¦ `Example_12`, so the
uploader shows a neutral filename and the document never announces its own answer. The buttons are
labelled because a demo you can drive deliberately is more useful than a guessing game β and it
costs nothing on the model side, since a filename never reaches the prompt. Only extracted text
does.
Two further precautions on the files: the numbering is shuffled with a fixed seed so `Example_0`
is not simply the alphabetically-first family, and every file sits in a **12β15 KB band** so file
size does not reveal which one is the clean control. An earlier draft had the clean file at 3 KB
against 200 KB+ injected ones, which gave it away completely.
Button labels are read from `examples/manifest.json`, written by the same script that copies the
PDFs β so a button cannot end up pointing at the wrong family.
Being small also matters practically: each example produces 2β8 regions, so **every one fits in a
single batch** β one example is one GPU run covering the whole document.
`examples/Example_Key.txt` records which is which, for whoever is marking this. It also notes that
several files legitimately carry two structural signatures (the generator gives some families a
JavaScript launcher), so `ransomware_simulation` and `dde_template_injection` also match
`javascript_injection`, and `object_action_injection` also matches `shellcode_embedded_exe`.
The clean control is there for the harder half of the claim: watching the app *not* flag something.
## Why MiMo, when Gemma scored higher
Part B's actual winner is **Gemma-2-9B at F1 0.969**, and this app runs **MiMo-7B at F1 0.945**
instead. That is a hosting decision, not a disagreement with the evaluation, and it comes down to
three things a free Space cannot absorb:
- **Gemma is gated.** It needs a Hugging Face account with Google's licence accepted, plus a read
token. On a public Space that means the first thing a new visitor meets is a 403, or the app has
to ask strangers to paste a token. MiMo downloads for anyone, with no account at all.
- **Gemma is 2.6Γ slower** β 10.95 s per window against MiMo's 4.18 s, measured in Part B on the
same T4. On ZeroGPU a single grant is capped at 300 s and the whole scan plus the model load must
fit inside it, so the slower model means roughly a third as many regions per run.
- **Free ZeroGPU is about five minutes per day.** At Gemma's rate that is a couple of batches for a
whole day; at MiMo's it is meaningfully more. The 9B model is also ~6 GB in 4-bit against MiMo's
~5 GB, which is not decisive on its own but points the same way.
What it costs: **0.024 F1**, and family-naming accuracy drops from 63% to 43%. Both figures are
stated in the interface rather than quietly rounded away. If the Space were ever moved to dedicated
paid hardware, switching to Gemma would be a change to `mimo.py`'s repo constants and a token
secret β the prompt, prefill and parser are shared and would not need touching.
## Two runtimes, chosen at startup
A Space's hardware is not this code's decision, so `mimo.py` carries both paths and picks one at
import time. Which is live is printed at the top of the interface and again at the foot of every
report.
**`gpu` β ZeroGPU, and Part B's own configuration.** The original BF16 checkpoint of
`XiaomiMiMo/MiMo-7B-RL` quantised to 4-bit NF4 by `bitsandbytes`, greedy, 200 new tokens, batched
at 8. Nothing about the arithmetic differs from the run that produced F1 0.945, so that figure
describes this configuration. About 4 seconds per region, which is why the slider goes to 24.
The whole scan happens inside a single `@spaces.GPU` call, because ZeroGPU grants and reclaims the
device around each one and paying the model-load cost per region would dominate everything else.
**`cpu` β a free CPU Space, where `bitsandbytes` cannot run at all** (it requires CUDA). The same
base model runs as
[`quantflex/MiMo-7B-RL-nomtp-Q4_K_M.gguf`](https://huggingface.co/quantflex/MiMo-7B-RL-nomtp-GGUF)
(4.7 GB) through `llama.cpp`. That build has MiMo's multi-token-prediction layers removed, because
`llama.cpp` cannot load them β MTP is a speculative-decoding accelerator that the ordinary forward
pass does not use, so greedy output should be unaffected, but it is a real difference. **On this
path, read F1 0.945 as the figure for the configuration Part B measured, not for the one running.**
About two minutes per region on 2 vCPUs, so the slider stops at 6.
The prompt, the prefill, the decoding parameters and the parser are byte-identical on both paths.
**On this Space the CPU path is dormant, and that is not a choice.** `llama-cpp-python` cannot be
installed here by either available route: the prebuilt wheels are tagged `linux_x86_64` but linked
against **musl** while a Space runs on glibc (`libc.musl-x86_64.so.1: cannot open shared object
file`), and PyPI ships no binary wheel at all β compiling the sdist exceeded the Space build limit
with `Job timeout`. So `requirements.txt` installs neither, the runtime picker hides itself when
only one backend is live, and the code path stays in `mimo.py` for any machine that does have a
working llama.cpp. The runtime that remains is the one that reproduces Part B exactly, so what was
lost is a fallback rather than a capability.
**Where both are live, the picker matters.** A free Hugging Face account gets
roughly **five minutes of ZeroGPU per day**, and the scheduler reserves a run's full requested
duration up front rather than what it turns out to use β so a couple of batches can exhaust the
day. When that happens the Space is not broken: switch the runtime to `cpu` and it keeps working,
slowly and without any quota. That is also why one GPU grant is kept short (110s requested, 165s
reserved) and a batch tops out at 8 regions.
## Triage, and why it exists
**It triages instead of scanning everything.** MiMo reads a 3,000-character window in roughly two
minutes on 2 vCPUs, and a real PDF has dozens of windows. So the same marker alternation that
located the payload in the corpus is run over the whole skeleton, every hit becomes a candidate
window with the identical Β±1,500-character shape, overlapping ones are merged, and the most
marker-dense go to the model first. **The ranking decides reading order, never the verdict.** The
report always says how many marked regions were left unread, so "clean" never overstates itself.
A file with no marker anywhere yields exactly one marker candidate β the head of the document β
which is byte-identical to what the corpus builder produced for a *clean* file.
### Batches, and the sweep
Marker regions alone leave most of a file unread: the triage only knows the twelve families this
project generated, so a payload shaped like none of them produces no marker and would sit in text
the model never saw while the report said "clean". So after the marker regions, the rest of the
skeleton is tiled into windows of the same size, and the whole list is cut into **batches sized to
fit one run of the model** β one ZeroGPU grant, or a tolerable wait on CPU. You pick which batch to
spend a run on, and the report always states how much is still unread.
**The sweep regions do not inherit Part B's accuracy, and the app says so.** Part B only ever
showed MiMo marker-centred windows or the head of a document. Handed an arbitrary mid-file content
stream β a page of font-positioning operators β MiMo frequently does not answer at all: it carries
on copying the input after the prefill, and the answer parses as unrecoverable, which scores as
*not injected*. Those regions buy coverage of text that would otherwise never be looked at; a
*clean* verdict on one is close to no evidence. The regions table labels every row `marker` or
`sweep`, the report counts the sweep parse failures separately and explains them, and the sweep can
be switched off to keep the app strictly inside the shape Part B measured.
**Nothing is downloaded until it is needed.** The page comes up first; the 4.7 GB GGUF, the
550 MB embedding model and the 3 MB index arrive on the first scan and are cached after that.
## The numbers
On the 1,100-document corpus Part B measured, MiMo-7B-RL scored:
| | |
|---|---|
| F1 | **0.945** |
| precision | 0.988 |
| recall | 0.906 |
| names the family correctly | 43.3% of files it caught |
| false alarms | 10 of 200 clean files (5%) |
| unparsable answers | 155 of 1,100 |
**A detector that calls every file malicious scores F1 0.900 on this corpus**, because 82% of it is
injected. Read 0.945 against 0.900, not against zero β it is a 5% relative improvement on doing no
work at all. Gemma-2-9B scored 0.969 and is the actual Part B winner; MiMo is used here because it
is ungated, needs no token, and is 2.6Γ faster, which on a CPU is the difference between usable
and not.
Two limits worth stating plainly:
- **The family is a suggestion, not a verdict** β right 43% of the time. The nearest known corpus
files are shown beside it so the two can disagree in public.
- **The nearest-file lookup is weak on purpose to report.** Part A's winning embedder reaches
precision@5 of 35.6% against a 6.8% random baseline: fewer than 2 of the 5 files returned are the
same kind of attack. Far better than chance, and not good. It is labelled *resemblance*, never
*identification*.
## The files
| File | What it does |
|---|---|
| `app.py` | The Gradio interface and the document-level report. No detection logic. |
| `corpus_text.py` | PDF bytes β skeleton β candidate windows. Everything above `Triage` is verbatim from the EDA notebook. |
| `mimo.py` | The prompt, the prefill, the parser (verbatim from Part B) and the llama.cpp runtime. |
| `neighbours.py` | Part A's embedding index and the nearest-neighbour lookup, with a provenance assertion. |
`neighbours.check_provenance()` asserts the embedder repo, prefix, dimension, normalisation and
input column against Part A's own `part_a_results.json` before any lookup runs β a query embedded
with the wrong model lands in a different space and returns meaningless neighbours silently, with
no error anywhere.
## Secrets
None. Every model used here is ungated.
|