Faucon 2 🦅 — a scam screenshot detector that shows where it looks
Faucon ("falcon") scores screenshots for scams (fake crypto giveaways, fake airdrops and presales, fake "free Nitro / Robux / gift card" offers, wallet-draining "sync / rectification" pages, fake casinos, fake investment platforms…). Instead of answering in one glance, it looks at the whole image, then zooms twice where it wants, remembers what it saw, and tells you what kind of scam it suspects — and you can see every zoom.
Red box = zoom 1, blue box = zoom 2, both chosen by the model (position and zoom level); the zoomed regions are shown on the right. Generated example with fictional names and a fictional look-alike domain.
How it thinks
Everything happens inside one ONNX file — no extra logic, no external model:
- First look — the whole screenshot → a first opinion, a first suspicion and a "how sure am I" signal.
- Zoom 1 — it picks where to look (attention over the image, guided by what it suspects) and how much to zoom (×2 for a block, up to ×3 for small text), then reads that region from a double-resolution copy of the image (real extra pixels).
- Memory — a small recurrent memory keeps what it saw and where it looked, so the second zoom goes somewhere else (another clue, the rest of the sentence, the button…).
- Zoom 2 — same, guided by the memory.
- Decision — after each look it may change its mind; it also learned how likely it is to stop at each step, and the final score weighs the looks accordingly (easy images: the first look counts most).
- Why — after its last look it names the kind of scam it suspects (6 types, see below).
Nobody tells it where to look: the zooms are learned only from being right or wrong.
Quick start
pip install onnxruntime pillow numpy
python predict.py screenshot.png more.png # prints verdicts AND saves zoom pictures in glimpses/
SCAM 0.958 1.9 looks zooms: x2.1 bottom center, x2.1 center suspects: fake game gift / gift card fake_nitro_dm.jpg
clean 0.023 2.1 looks zooms: x2.2 center, x2.2 center friend_chat.png
Every image gets a picture in glimpses/ like the one above: where it zoomed, what it saw, its score, how many
looks counted and what it suspects. Use --no-images to only print verdicts, --threshold 0.5 to catch more
scams (for human review).
ONNX outputs
Input image: float32 NCHW, dynamic batch / height / width.
| Output | Shape | Meaning |
|---|---|---|
logits |
(N, 2) | [clean, scam] → softmax, index 1 = scam probability |
glimpse_centers |
(N, 2, 2) | (x, y) in [-1, 1] of zoom 1 and zoom 2 |
glimpse_windows |
(N, 2) | side of each zoom box as a fraction of the image side (0.5 = ×2, 0.33 = ×3) |
reason_probs |
(N, 6) | suspected scam type after the last look: fake crypto giveaway · fake airdrop / presale · account / wallet theft · fake game gift / gift card · fake casino / prize · fake investment |
looks |
(N,) | how many looks counted (1 = the first look was enough, up to 3) |
Preprocessing (see predict.py): RGB with transparency on white → longest side ≤ 768 px (bicubic) → closest
aspect-ratio bucket (height/width ∈ {0.5, 2/3, 1, 1.5, 2}) resized at double resolution (≈ 448 × 448 pixels, sides
multiple of 16, bilinear) → (x / 255 - 0.5) / 0.5.
Model
| Architecture | small CNN trained from scratch (own self-supervised pretraining on ~167 k unlabeled images, no external pretrained model) + learned attention, zoom position and zoom level, recurrent memory, learned stopping, scam-type head |
| Parameters | ~1.35 M (the whole "thinking" part: ~185 k) |
| Looks per image | 3 (whole image + 2 zooms; zooms computed at reduced size to stay cheap) |
| Speed | ~20–30 ms per image on a laptop CPU (Ryzen 3 5400U), no GPU needed |
| File | Faucon.onnx, 5.5 MB, opset 17 (GridSample inside the graph) |
| Recommended threshold | 0.6 (≈ 0.5 % false positives on everyday images) · 0.5 to catch more scams with human review |
Trained entirely on a laptop CPU 💻
Every step — the self-supervised pretraining, the detector and the "thinking" part — was trained on my own laptop CPU: an AMD Ryzen 3 5400U (4 cores), 7.3 GB of RAM, no GPU, no cloud. That was not a given: before optimizing, a single detector training took over 3 hours, and the laptop often has only ~1.5 GB of free memory. It took a lot of research and measuring to make it work:
- progressive resolution (128 → 160 → 192 → 224 px) and a memory layout faster on CPU: a full detector training went from ~3 h 10 to ~1 h 15;
- our own lightweight self-supervised pretraining instead of a large external model (~2 h 35 for ~167 k images);
- a small model by design (~1.35 M parameters) and zooms computed at reduced size, so two zooms cost about as much as one;
- training only the thinking part on top of a frozen detector (plus one slowly-unfrozen block), with a reminder of what the detector already knew, so it learns to think without getting worse;
- hard negative mining (searching ~150 k images for the ones earlier versions got wrong) instead of throwing more data at it;
- many measured dead ends along the way (compiling the model, a C++ image pipeline… both were slower here).
The "thinking" training of Faucon 2 alone took about 3 hours for 30 epochs. The result runs the same way: on any CPU, ~20–30 ms per image, no GPU needed.
Evaluation (real images, never seen during training)
The main benchmark uses 87 recent real scam pages (captured by public URL scanners in October 2026, reviewed by hand, no domain, title or page template shared with the training data — checked with perceptual hashing) and 231 legitimate pages in the same capture format (36 official sites: exchanges, wallets, game stores, shops, banks, news…), plus 2,000 everyday clean images (art, web/app UIs, games, photos).
| Threshold 0.6 (recommended) | Threshold 0.5 | |
|---|---|---|
| Everyday clean images flagged | 10 / 2,000 (0.5 %) | 13 / 2,000 (0.65 %) |
| Recent real scams caught | 59 / 87 (68 %) | 66 / 87 (76 %) |
| Legitimate web pages flagged | 26 / 231 (11 %) | 30 / 231 (13 %) |
| Older web scams (2020–22 style) caught | 11 / 36 | 15 / 36 |
| Legitimate X profiles flagged | 0 / 30 | 1 / 30 |
| Discord moderation samples (80 clean, 1 scam) | 0 false positives, scam missed | 0 false positives, scam caught |
Compared at the same false-positive rate (≤ 9 / 2,000 everyday images): Faucon 2 catches 58 / 87 recent scams with 24 / 231 legitimate pages flagged; Faucon 1 caught 45 / 87 (13 / 231 flagged); a classic small CNN without zooms trained on older data caught 11 / 87. Test sets are small: differences of 1–2 images are within noise.
Known limitations
- It finds clues it does not fully trust yet. The second zoom often lands on the decisive clue (e.g. the "Claim" button of a fake airdrop), but the learned trust in zoomed regions is still low, so zooms change the final score modestly — they mostly confirm or soften the first look. This is the focus of the next version.
- About 1 legitimate web page in 9 with a scam-like layout (crypto exchanges with bonuses, referral pages, shops with "claim" buttons, donation forms) is flagged at 0.6. Use it to flag for human review, not to punish.
- The suspected scam type is right about two times out of three on training-style scams and can be wrong (e.g. a fake airdrop described as a "game gift"); it is shown only when the score is ≥ 0.3.
- New scam styles appear every week and may be missed; not designed for phishing pages identical to the real site (scam only in the URL), text-only messages or videos (for GIFs, score key frames and take the maximum).
Training data (summary)
- Self-supervised pretraining (no labels) on ~167,000 images: synthetic websites (WebSight), photos (COCO, mini-ImageNet), mobile app screens (RICO), gameplay frames, memes, desktop UI screenshots (WildGUI), plus the labeled set below without its labels.
- Supervised training on ~1,200 scam and ~23,400 clean images: real scam pages captured by public URL scanners (reviewed by hand), synthetic scam screenshots each with a legitimate twin, legitimate pages from official domains, real legitimate websites, ~5,900 hard negatives (legitimate images that earlier versions wrongly found suspicious, reviewed by hand) and moderation samples from the author's own Discord bot.
The training data and training code are not released.
Intended use
- ✅ Flagging screenshots for human review in community moderation; research on explainable "looking closer" models.
- ❌ Automatic punishment of users; commercial use; deciding on its own whether a website is safe.
Previous version
Faucon 1 (first preview: one fixed ×2 zoom, no memory) is kept in Faucon-1/ with its own
predict.py.
License
CC BY-NC 4.0. Several pretraining and training sources only allow non-commercial / research use, hence this license.
Version
Faucon 2 — October 2026. Next steps: teaching it to trust a zoom when the zoom is sure and the first look hesitated, and stopping early on easy images to save compute.
Citation
If you use Faucon in your work, please cite it (attribution is required by the CC BY-NC 4.0 license):
@misc{charlet2026faucon,
author = {Charlet, Th{\'e}o},
title = {Faucon: a scam screenshot detector that learns where to look closer},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/RDTvlokip/Faucon}}
}
Théo CHARLET
TSSR Graduate (IT Systems & Networks Technician) — AI/ML Specialization
Creator of AG-BPE (Attention-Guided Byte-Pair Encoding)
🚀 Seeking internship opportunities
- Downloads last month
- 17
