BentoUniAcc's picture
Replace static placeholder with the Gradio PDF injection detector
1c16318 verified
|
Raw
History Blame Contribute Delete
5.12 kB

A newer version of the Gradio SDK is available: 6.22.0

Upgrade
metadata
title: PDF Injection Detector
emoji: πŸ”
colorFrom: yellow
colorTo: gray
sdk: gradio
app_file: app.py
suggested_hardware: zero-a10g
pinned: false
license: mit
short_description: Finds payloads hidden inside PDF files

PDF Injection Detector

Upload a PDF. The app renders it to text with the same pipeline that built the evaluation corpus, reads it window by window with a language model, and reports whether something has been hidden inside it β€” which part of the file, what kind of thing it is, and the exact text that gave it away.

It is a coursework artefact built on a synthetic corpus of 1,100 PDFs, and it is honest about what that is worth. It is not a general malware scanner.

How to use it

  1. Pick a model. MiMo-7B is the default and needs no account.
  2. Press Start. The app downloads the embedding index and the model weights, then quantises to 4-bit. Nothing is downloaded before this β€” the page comes up first.
  3. Upload a PDF and press Check this PDF. The app tells you how many sections it will read and roughly how long that takes before you commit to it.

Results arrive in three tabs: the report, every section with what the model said about it, and the nearest files in the corpus this system was measured on.

The four models

Model F1 Names the family Speed Token Note
MiMo-7B (default) 0.945 0.433 ~4.2 s/window not needed Reasoning-trained; needs a prefilled think-block to answer at all
Gemma-2-9B 0.969 0.630 ~10.9 s/window required (gated) The winner: 3% false alarms, 6 of 200 clean files
Qwen2.5-7B 0.957 0.524 ~10.9 s/window not needed The most precise (0.995) and the worst at making itself understood
Phi-4-mini 0.900 β€” ~2.1 s/window not needed Ties the always-malicious constant; flags 95% of clean files

MiMo is the default rather than the winner because it is the only one of the four that anybody can download: Gemma is gated by Google, and defaulting to it would mean a new visitor's first experience is a licence refusal. Picking Gemma reveals a token box; paste a read token from huggingface.co/settings/tokens and press Send token. It is used for that session only and is never stored.

Phi-4-mini is in the list because Part B measured it, not because it should be used. Its F1 of 0.900 is exactly what a detector that flags every file without reading it scores.

What the numbers mean

A detector that calls every file malicious scores F1 0.900 on this corpus, because 82% of the corpus is injected. Every score above should be read against that number, not against zero. On that basis Gemma's 0.969 is a 7.7% relative improvement, and Phi's 0.900 is none at all.

Two more limits worth stating plainly:

  • The family is a suggestion, not a verdict. The best model names the right family 63% of the time. The app shows the nearest known files beside it so the two can disagree in public.
  • The false-alarm rate was measured per window, not per document. A long PDF is scanned in dozens of windows, and each is a fresh chance to raise a flag. The report counts how many windows agreed, which is a partial answer, not a full one.

The corpus is synthetic: harmless EICAR/AMTSO/WICAR/RANSIM test markers injected into ordinary PDFs. Real malware does not announce itself the same way.

Running it

Hardware. This needs a GPU. 4-bit quantisation goes through bitsandbytes, which requires CUDA, so the free CPU tier cannot serve any of these models β€” the app says so rather than failing obscurely. ZeroGPU is supported: the scan is cut into slices of 8 windows, each asking for the GPU separately, because a long document is well past a single allocation.

Secrets. None are required. Setting HF_TOKEN as a Space secret lets Gemma load without anyone pasting a token; the app also checks Colab's secret store and the CLI login cache.

The files

File What it does
app.py The Gradio interface. Holds no detection logic.
pipeline.py PDF bytes to text, then into overlapping 3,000-character windows.
detector.py The prompt, the parser, and the model registry.
analysis.py Per-window verdicts into one document verdict.
embedder.py The Part A embedding index and the nearest-neighbour lookup.

Everything above window_document in pipeline.py, and the prompt and parser in detector.py, are lifted verbatim from the notebooks that built and measured the corpus. That is the correctness argument for quoting those scores here at all: change how the text is extracted or how the question is asked, and the published numbers stop describing this program.

One deliberate exception: the salvage regexes in detector.parse_response are wider than the ones Part B scored with, because the originals discarded any answer that was not strict JSON. That recovers verdicts the evaluation threw away, and it means the F1 figures above describe the original parser, not this one.