manga-ocr-base-coreml
Model Summary
This is an unofficial Core ML conversion of kha-white/manga-ocr-base (a Japanese OCR model specialized for manga speech-bubble text), for running directly on iOS/macOS. All credit for the original model goes to its author.
Architecture
A ViT encoder (image -> features) + a 2-layer BERT decoder (features -> text, autoregressive, character-level). Since the decoder is only 2 layers deep, a simple design without a KV cache still runs at practical speed.
| File | Content | Size |
|---|---|---|
manga-ocr-base_encoder_fp16.mlpackage |
Image encoder (fixed input: 224x224 RGB) | ~164MB |
manga-ocr-base_decoder_seq32_fp16.mlpackage |
Text decoder (sequence length 32) | ~46MB |
manga-ocr-base_decoder_seq64_fp16.mlpackage |
Text decoder (sequence length 64, for longer lines) | ~46MB |
Usage
1. Image preprocessing
Resize and normalize the image to 224x224. Use transformers'
ViTImageProcessor (kha-white/manga-ocr-base), or an equivalent
implementation.
2. Run the encoder
encoder_hidden_states = encoder.predict({"pixel_values": pixel_values})["encoder_hidden_states"]
3. Run the decoder (greedy decoding, no KV cache)
decoder_start_token_id = 2
eos_token_id = 3
SEQ_LEN = 32 # match the decoder file you're using
generated = [decoder_start_token_id]
for step in range(SEQ_LEN - 1):
padded = generated + [0] * (SEQ_LEN - len(generated)) # pad with pad_token_id=0
logits = decoder.predict({
"decoder_input_ids": np.array([padded], dtype=np.int32),
"encoder_hidden_states": encoder_hidden_states.astype(np.float16),
})["logits"]
next_id = int(np.argmax(logits[0, len(generated) - 1]))
generated.append(next_id)
if next_id == eos_token_id:
break
text = tokenizer.decode(generated, skip_special_tokens=True)
Uses a character-level tokenizer (BertJapaneseTokenizer, based on
cl-tohoku/bert-base-japanese-char-v2).
Accuracy
Verified on 2 generated test images ("瑠璃色の空", "ありがとうございます") that
the Core ML output token sequence (via the procedure above) exactly matches
PyTorch's model.generate() output token sequence.
Notes
- The original model (
kha-white/manga-ocr-base) is distributed only aspytorch_model.bin(pickle format); this converted version consists only ofsafetensors. - This is a community conversion, not an official release from the original author.
Security
Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.
モデルの概要
kha-white/manga-ocr-base(漫画のフキダシ文字認識に特化した日本語OCRモデル)を、iOS/macOS (Core ML) で直接動かせるように変換したものです。元モデルの著作権はその作者に帰属します。
構成
ViTエンコーダー(画像→特徴量)+ 2層BERTデコーダー(特徴量→テキスト、自己回帰・文字単位)という構成です。 デコーダーが2層と浅いため、KVキャッシュなしのシンプルな設計で実用的な速度が出ます。
| ファイル | 内容 | サイズ |
|---|---|---|
manga-ocr-base_encoder_fp16.mlpackage |
画像エンコーダー(固定入力: 224×224 RGB) | 約164MB |
manga-ocr-base_decoder_seq32_fp16.mlpackage |
テキストデコーダー(系列長32) | 約46MB |
manga-ocr-base_decoder_seq64_fp16.mlpackage |
テキストデコーダー(系列長64、長めのセリフ向け) | 約46MB |
使い方
1. 画像の前処理
画像を224×224にリサイズ・正規化します。transformersのViTImageProcessor(kha-white/manga-ocr-base)、
または同等の前処理を実装してください。
2. エンコーダーの実行
encoder_hidden_states = encoder.predict({"pixel_values": pixel_values})["encoder_hidden_states"]
3. デコーダーの実行(グリーディデコード、KVキャッシュなし)
decoder_start_token_id = 2
eos_token_id = 3
SEQ_LEN = 32 # 使用するデコーダーのファイルに合わせる
generated = [decoder_start_token_id]
for step in range(SEQ_LEN - 1):
padded = generated + [0] * (SEQ_LEN - len(generated)) # pad_token_id=0 で埋める
logits = decoder.predict({
"decoder_input_ids": np.array([padded], dtype=np.int32),
"encoder_hidden_states": encoder_hidden_states.astype(np.float16),
})["logits"]
next_id = int(np.argmax(logits[0, len(generated) - 1]))
generated.append(next_id)
if next_id == eos_token_id:
break
text = tokenizer.decode(generated, skip_special_tokens=True)
文字単位トークナイザ(BertJapaneseTokenizer, cl-tohoku/bert-base-japanese-char-v2ベース)を使用します。
精度検証
生成したテスト画像2種(「瑠璃色の空」「ありがとうございます」)で、PyTorchのmodel.generate()の
出力トークン列と、上記手順によるCore ML版の出力トークン列が完全に一致することを確認しています。
備考
- 元モデル(
kha-white/manga-ocr-base)はpytorch_model.bin(pickle形式)のみで配布されていますが、 本変換版はsafetensorsのみで構成されています。 - 本変換は非公式のコミュニティ版です。
セキュリティー
model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。
- Downloads last month
- 19
Model tree for masahiroid/manga-ocr-base-coreml
Base model
kha-white/manga-ocr-base