table-transformer-detection-mlx

English | 日本語

Model Summary

This is an unofficial MLX conversion of microsoft/table-transformer-detection (a table-region detector for document images, using a ResNet18 backbone and a DETR-style encoder-decoder). All credit for the original model goes to its authors (Microsoft).

This cannot be loaded with mlx-vlm

A ResNet18 + DETR-style encoder-decoder with learned object queries isn't in mlx-vlm's list of supported architectures, so this was reimplemented from scratch for MLX and requires the bundled table_transformer_mlx.py.

Limitation: fixed square input only, no padding

This implementation is simplified by assuming pixel_mask is always fully valid (no padding). Concretely, it only works correctly when input images are always resized to a fixed 800x800 square (not when batching images of different aspect ratios with letterbox-style padding). This simplification lets the implementation skip the attention-mask handling that DETR-family models normally require.

Usage

import mlx.core as mx
from mlx.utils import tree_unflatten
from PIL import Image
import numpy as np

from table_transformer_mlx import TableTransformerMLX, IMAGE_SIZE

model = TableTransformerMLX()
weights = mx.load("model.safetensors")
model.update(tree_unflatten(list(weights.items())))
mx.eval(model.parameters())

image = Image.open("document.png").convert("RGB").resize((IMAGE_SIZE, IMAGE_SIZE))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = mx.array(pixel_values[None].astype(np.float16))  # (1, 800, 800, 3), NHWC

logits, boxes = model(pixel_values)  # logits: (1, 15, 3), boxes: (1, 15, 4) normalized cxcywh

id2label is the same as the original model: {0: "table", 1: "table rotated"} (class 2 means "no object"). boxes are (center_x, center_y, width, height) normalized to [0, 1].

Accuracy

Compared against the PyTorch fp32 reference on a COCO validation image:

Precision Logits cosine sim. Boxes cosine sim. Label agreement
MLX fp32 1.0 1.0 100%
MLX fp16 (this release) 0.9999995 0.9999999 100%

Specs

Item Value
Base model microsoft/table-transformer-detection (ResNet18 + DETR, 28.8M params)
Precision float16
Input 800x800, NHWC, fixed size, no padding
Framework MLX (from-scratch table_transformer_mlx.py)

Notes

  • This is a community conversion, not an official release from Microsoft.
  • Security audit uses model-audit-lite (see SECURITY.md for details).

Security

Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.


モデルの概要

microsoft/table-transformer-detection (ResNet18バックボーン + DETR系エンコーダー・デコーダーによる、文書画像中の表領域検出モデル)の MLX版です。元モデルの著作権はその作者(Microsoft)に帰属します。

mlx-vlmでは読み込めません

ResNet18 + DETR系エンコーダー・デコーダー + 学習済みobject queryという構成はmlx-vlmの対応 アーキテクチャ一覧に含まれていないため、MLXでの実装をゼロから書き起こして変換しています。 同梱のtable_transformer_mlx.pyが必要です。

制約:パディング無しの正方形固定入力のみ対応

本実装は**pixel_maskが全て有効(パディング無し)であることを前提**に簡略化しています。 具体的には、入力を常に800x800の正方形にリサイズして使う場合(バッチ処理でアスペクト比の 異なる画像を混在させ、レターボックス的にパディングする使い方はしない場合)にのみ正しく動作します。 これにより、DETR系実装で通常必要になるattention maskの処理をすべて省略できています。

使い方

import mlx.core as mx
from mlx.utils import tree_unflatten
from PIL import Image
import numpy as np

from table_transformer_mlx import TableTransformerMLX, IMAGE_SIZE

model = TableTransformerMLX()
weights = mx.load("model.safetensors")
model.update(tree_unflatten(list(weights.items())))
mx.eval(model.parameters())

image = Image.open("document.png").convert("RGB").resize((IMAGE_SIZE, IMAGE_SIZE))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = mx.array(pixel_values[None].astype(np.float16))  # (1, 800, 800, 3), NHWC

logits, boxes = model(pixel_values)  # logits: (1, 15, 3), boxes: (1, 15, 4) normalized cxcywh

id2labelは元モデルと同じ: {0: "table", 1: "table rotated"}(クラス2は「該当なし」)。 boxesは(center_x, center_y, width, height)を0〜1に正規化した値。

精度検証

PyTorch fp32リファレンスと、COCO検証画像1枚で比較:

精度 Logitsコサイン類似度 Boxesコサイン類似度 ラベル一致率
MLX fp32 1.0 1.0 100%
MLX fp16(本リリース) 0.9999995 0.9999999 100%

Specs

Item Value
ベースモデル microsoft/table-transformer-detection(ResNet18 + DETR、28.8M params)
精度 float16
入力 800x800、NHWC、固定サイズ、パディング無し
フレームワーク MLX(ゼロから実装したtable_transformer_mlx.py)

備考

  • 本変換は非公式のコミュニティ版です。
  • セキュリティー監査にはmodel-audit-liteを 使用しています(詳細はSECURITY.md)。

セキュリティー

model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
28.8M params
Tensor type
F16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for masahiroid/table-transformer-detection-mlx

Finetuned
(16)
this model