enlitqhtened's picture
Upload folder using huggingface_hub
bda422a verified
|
Raw
History Blame Contribute Delete
1.83 kB
---
license: apache-2.0
tags:
- object-detection
- comics
- onnx
---
# DASS comic character detector — ONNX (`dcm` weights)
ONNX conversion of **DASS** (*Domain-Adaptive Self-Supervised Pre-training for
Face & Body Detection in Drawings*) by Barış Batuhan Topal et al., using the
`xl_dcm_finetuned_stage3` weights — the variant fine-tuned on **DCM772**
(western comics) rather than manga.
- Upstream: https://github.com/barisbatuhan/DASS_Det_Inference
- Licence: **Apache-2.0**, unchanged.
- **No retraining and no fine-tuning.** This is a format conversion only:
the authors' PyTorch checkpoint exported to ONNX (opset 17). The state dict
loaded with 0 missing and 0 unexpected keys.
## Inputs / outputs
| | shape | notes |
|---|---|---|
| input `images` | `[1, 3, 1024, 1024]` | **raw 0–255**, NCHW, RGB. No `/255`, no mean/std. |
| output 0 | `[1, 21504, 5]` | **faces** — `(cx, cy, w, h, score)` |
| output 1 | `[1, 21504, 5]` | **bodies** — same layout |
Two things worth knowing, both verified by probing the graph rather than assumed:
1. **Outputs are already decoded.** Coordinates arrive in 0–1024 pixel space and
scores are already through a sigmoid. Do **not** apply the usual YOLOX
grid/stride decode — it will corrupt every box.
2. **Input is unnormalised.** Dividing by 255 drives every score toward zero,
which looks like "the model finds nothing" rather than a preprocessing bug.
Letterbox rather than stretch (pad value 114), and run NMS **per head** — a face
sits inside its body, so a shared NMS makes them suppress each other.
## Citation
```
@article{topal2022domain,
title={Domain-Adaptive Self-Supervised Pre-training for Face & Body Detection in Drawings},
author={Topal, Bar{\i}\c{s} Batuhan and Yuret, Deniz and G{\"u}ng{\"o}r, Tunga},
year={2022}
}
```