Manga Speech-Bubble Segmentation Weights
Converted SafeTensors weights for a single-class manga and comic speech-bubble instance-segmentation model.
This repository distributes model weights only. It does not include the model architecture, preprocessing code, or an inference implementation.
Model details
| Property | Value |
|---|---|
| Task | Instance segmentation |
| Model family | YOLO11n-style segmentation |
| Number of classes | 1 |
| Class | bubble (class ID 0) |
| Recommended inference size | 1600 ร 1600 |
| Inference strides | 8, 16, 32 |
| Default confidence threshold in the reference implementation | 0.25 |
| Default NMS IoU threshold in the reference implementation | 0.7 |
| Maximum detections in the reference implementation | 300 |
The model predicts speech-bubble bounding boxes and instance masks. Input images are expected to be RGB; the reference inference implementation uses letterboxing for model input.
Repository contents
.
โโโ model.safetensors
model.safetensorsโ converted model weights in SafeTensors format.
The file contains tensors with the net.stages.* names expected by the
corresponding standalone PyTorch implementation. It is not a drop-in
replacement for the original best.pt checkpoint and should not be loaded
directly with the standard Ultralytics YOLO(...) interface.
Loading the weights
Install PyTorch and SafeTensors:
pip install torch safetensors
Download the file from this repository and load its tensors:
from safetensors.torch import load_file
state_dict = load_file("model.safetensors", device="cpu")
print(f"Loaded {len(state_dict)} tensors")
for name, tensor in list(state_dict.items())[:5]:
print(name, tuple(tensor.shape), tensor.dtype)
Loading the tensor dictionary is not sufficient to perform inference. You must provide a compatible model implementation that defines the expected architecture and tensor names.
Conversion
The weights were converted from the EMA state dictionary in the upstream
best.pt checkpoint.
The conversion process:
- Loaded the checkpoint on CPU and selected its EMA model state dictionary.
- Renamed keys from the source
model.*namespace to thenet.stages.*namespace. - Converted FP16 tensors to FP32.
- Made tensors contiguous and saved them in SafeTensors format.
This is a format and key-name conversion; it does not involve additional training, fine-tuning, or quantization.
Evaluation
The original model card reports the following results at epoch 44:
| Metric | Box detection | Mask segmentation |
|---|---|---|
| Precision | 97.55% | 97.66% |
| Recall | 97.03% | 97.15% |
| mAP@50 | 99.10% | 99.13% |
| mAP@50โ95 | 96.67% | 94.69% |
These are metrics reported by the upstream project. They are not an independent evaluation of this converted file or of any particular inference implementation.
Source and attribution
Source checkpoint and upstream model card:
The upstream model card describes the model as a YOLO11n segmentation model fine-tuned for manga/comic speech bubbles and lists the following datasets:
Please consult the upstream model card and the dataset terms for additional training-data information and acknowledgements.
License
The original model and this conversion are distributed under Apache-2.0.
Model tree for iamokish/manga-bubble-segmentation-pytorch
Base model
huyvux3005/manga109-segmentation-bubble