Manga Speech-Bubble Segmentation Weights

Converted SafeTensors weights for a single-class manga and comic speech-bubble instance-segmentation model.

This repository distributes model weights only. It does not include the model architecture, preprocessing code, or an inference implementation.

Model details

Property Value
Task Instance segmentation
Model family YOLO11n-style segmentation
Number of classes 1
Class bubble (class ID 0)
Recommended inference size 1600 ร— 1600
Inference strides 8, 16, 32
Default confidence threshold in the reference implementation 0.25
Default NMS IoU threshold in the reference implementation 0.7
Maximum detections in the reference implementation 300

The model predicts speech-bubble bounding boxes and instance masks. Input images are expected to be RGB; the reference inference implementation uses letterboxing for model input.

Repository contents

.
โ””โ”€โ”€ model.safetensors
  • model.safetensors โ€” converted model weights in SafeTensors format.

The file contains tensors with the net.stages.* names expected by the corresponding standalone PyTorch implementation. It is not a drop-in replacement for the original best.pt checkpoint and should not be loaded directly with the standard Ultralytics YOLO(...) interface.

Loading the weights

Install PyTorch and SafeTensors:

pip install torch safetensors

Download the file from this repository and load its tensors:

from safetensors.torch import load_file

state_dict = load_file("model.safetensors", device="cpu")

print(f"Loaded {len(state_dict)} tensors")
for name, tensor in list(state_dict.items())[:5]:
    print(name, tuple(tensor.shape), tensor.dtype)

Loading the tensor dictionary is not sufficient to perform inference. You must provide a compatible model implementation that defines the expected architecture and tensor names.

Conversion

The weights were converted from the EMA state dictionary in the upstream best.pt checkpoint.

The conversion process:

  1. Loaded the checkpoint on CPU and selected its EMA model state dictionary.
  2. Renamed keys from the source model.* namespace to the net.stages.* namespace.
  3. Converted FP16 tensors to FP32.
  4. Made tensors contiguous and saved them in SafeTensors format.

This is a format and key-name conversion; it does not involve additional training, fine-tuning, or quantization.

Evaluation

The original model card reports the following results at epoch 44:

Metric Box detection Mask segmentation
Precision 97.55% 97.66%
Recall 97.03% 97.15%
mAP@50 99.10% 99.13%
mAP@50โ€“95 96.67% 94.69%

These are metrics reported by the upstream project. They are not an independent evaluation of this converted file or of any particular inference implementation.

Source and attribution

Source checkpoint and upstream model card:

The upstream model card describes the model as a YOLO11n segmentation model fine-tuned for manga/comic speech bubbles and lists the following datasets:

Please consult the upstream model card and the dataset terms for additional training-data information and acknowledgements.

License

The original model and this conversion are distributed under Apache-2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
2.86M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for iamokish/manga-bubble-segmentation-pytorch

Finetuned
(2)
this model