metadata
license: apache-2.0
tags:
- object-detection
- comics
- onnx
DASS comic character detector — ONNX (dcm weights)
ONNX conversion of DASS (Domain-Adaptive Self-Supervised Pre-training for
Face & Body Detection in Drawings) by Barış Batuhan Topal et al., using the
xl_dcm_finetuned_stage3 weights — the variant fine-tuned on DCM772
(western comics) rather than manga.
- Upstream: https://github.com/barisbatuhan/DASS_Det_Inference
- Licence: Apache-2.0, unchanged.
- No retraining and no fine-tuning. This is a format conversion only: the authors' PyTorch checkpoint exported to ONNX (opset 17). The state dict loaded with 0 missing and 0 unexpected keys.
Inputs / outputs
| shape | notes | |
|---|---|---|
input images |
[1, 3, 1024, 1024] |
raw 0–255, NCHW, RGB. No /255, no mean/std. |
| output 0 | [1, 21504, 5] |
faces — (cx, cy, w, h, score) |
| output 1 | [1, 21504, 5] |
bodies — same layout |
Two things worth knowing, both verified by probing the graph rather than assumed:
- Outputs are already decoded. Coordinates arrive in 0–1024 pixel space and scores are already through a sigmoid. Do not apply the usual YOLOX grid/stride decode — it will corrupt every box.
- Input is unnormalised. Dividing by 255 drives every score toward zero, which looks like "the model finds nothing" rather than a preprocessing bug.
Letterbox rather than stretch (pad value 114), and run NMS per head — a face sits inside its body, so a shared NMS makes them suppress each other.
Citation
@article{topal2022domain,
title={Domain-Adaptive Self-Supervised Pre-training for Face & Body Detection in Drawings},
author={Topal, Bar{\i}\c{s} Batuhan and Yuret, Deniz and G{\"u}ng{\"o}r, Tunga},
year={2022}
}