DASS comic character detector β€” ONNX (dcm weights)

ONNX conversion of DASS (Domain-Adaptive Self-Supervised Pre-training for Face & Body Detection in Drawings) by Barış Batuhan Topal et al., using the xl_dcm_finetuned_stage3 weights β€” the variant fine-tuned on DCM772 (western comics) rather than manga.

  • Upstream: https://github.com/barisbatuhan/DASS_Det_Inference
  • Licence: Apache-2.0, unchanged.
  • No retraining and no fine-tuning. This is a format conversion only: the authors' PyTorch checkpoint exported to ONNX (opset 17). The state dict loaded with 0 missing and 0 unexpected keys.

Inputs / outputs

shape notes
input images [1, 3, 1024, 1024] raw 0–255, NCHW, RGB. No /255, no mean/std.
output 0 [1, 21504, 5] faces β€” (cx, cy, w, h, score)
output 1 [1, 21504, 5] bodies β€” same layout

Two things worth knowing, both verified by probing the graph rather than assumed:

  1. Outputs are already decoded. Coordinates arrive in 0–1024 pixel space and scores are already through a sigmoid. Do not apply the usual YOLOX grid/stride decode β€” it will corrupt every box.
  2. Input is unnormalised. Dividing by 255 drives every score toward zero, which looks like "the model finds nothing" rather than a preprocessing bug.

Letterbox rather than stretch (pad value 114), and run NMS per head β€” a face sits inside its body, so a shared NMS makes them suppress each other.

Citation

@article{topal2022domain,
  title={Domain-Adaptive Self-Supervised Pre-training for Face & Body Detection in Drawings},
  author={Topal, Bar{\i}\c{s} Batuhan and Yuret, Deniz and G{\"u}ng{\"o}r, Tunga},
  year={2022}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support