mit48px-ocr / README.md
Mayo
Add complete model card metadata
205395b unverified
|
Raw
History Blame Contribute Delete
1.56 kB
---
license: gpl-3.0
library_name: candle
pipeline_tag: image-to-text
tags:
- candle
- ocr
- image-to-text
- manga
- comic
- multilingual
- safetensors
---
# MIT 48px OCR
A SafeTensors conversion of the 48-pixel OCR model used by [`zyddnys/manga-image-translator`](https://github.com/zyddnys/manga-image-translator) and BallonsTranslator. The model recognizes cropped comic text lines and also predicts foreground and background colors.
## Model details
- Input text height: 48 pixels
- Maximum configured width: 8100 pixels
- Visual backbone: ConvNeXt feature extractor
- Sequence model: four-layer Transformer encoder and five-layer Transformer decoder
- Embedding dimension: 320
- Attention heads: 4
- Default beam size: 5
- Default maximum sequence length: 255
## Files
- `model.safetensors`: converted model weights
- `config.json`: architecture and decoding configuration
- `alphabet-all-v7.txt`: tokenizer alphabet and special-token vocabulary
Token IDs are `0` for padding, `1` for beginning-of-sequence, and `2` for end-of-sequence. `<SP>` represents a space.
## Intended use and limitations
The model expects already detected, cropped, and normalized comic text regions. It does not locate text on a page. Recognition quality depends strongly on crop quality, text scale, language coverage in the supplied alphabet, and image degradation. Training data details and evaluation metrics are not included with this conversion.
## License
GPL-3.0, following the upstream manga-image-translator and BallonsTranslator implementations.