--- license: apache-2.0 tags: - executorch - xnnpack - pte - on-device - image-to-text - text-recognition --- # docTR CRNN-MobileNetV3-Small — ExecuTorch Text **recognition**: reads one word crop. The other half of the pair is [DB-MobileNetV3-Large](https://huggingface.co/mlboydaisuke/docTR-DB-MobileNetV3-Large-ExecuTorch), which finds the crops. - **Source**: [mindee/doctr](https://github.com/mindee/doctr) `crnn_mobilenet_v3_small`, 2.1M parameters - **License**: Apache-2.0 - **Input**: `[1, 3, 32, 128]`, normalised with docTR's own mean (0.694, 0.695, 0.693) and std (0.299, 0.296, 0.301) - **Output**: `[1, 32, 127]` — 32 timesteps over a 126-symbol alphabet plus the CTC blank Logits, not probabilities: greedy CTC only needs the argmax per timestep, so a softmax would be thrown away. The alphabet is `0123456789`, `a-z`, `A-Z`, punctuation, `°£€¥¢฿`, and accented Latin — read it out of `crnn_mobilenet_v3_small().cfg["vocab"]` rather than retyping it. ## Verification (Mac arm64, 2026-08-23) | build | size | latency | worst corr | |---|---|---|---| | XNNPACK fp32 | 8.8 MB | 3.2 ms | 1.000000 | | XNNPACK fp16 | 6.5 MB | 4.0 ms | 0.997143 | | Core ML fp32 | 6.2 MB | **0.3 ms** | 0.999530 | Eager fp32 for the same input is 121 ms. **Read end to end** with the detector, on a London street photograph (`convert/check_doctr.py`): ``` boxes: 10 read: TIUZABXPRESS | Chrisiophers | Place | STREE! | LAG | BAR! ``` The signs say PIZZA EXPRESS and Christopher's Place. ## Cropping it docTR's own `Resize` for this model is `preserve_aspect_ratio=True, symmetric_pad=False`: the crop keeps its shape, sits at the top-left, and the rest of the strip is black. The check does the same. That is the principled choice rather than a measured improvement. Over four street photographs neither it nor a plain squeeze won: one picture reads `Chrisiophers` squeezed and `Clridabers` padded, another reads `TRAFFIC` padded and `TRAEFIC` squeezed. ## Not shipped, and why - **int8**: worst corr **0.883**, against a 0.95 gate, for 6.1 MB against 8.8. Recognition is the half where a wrong character is the whole answer, so this one is not close enough. ## Conversion ```bash python convert/export_doctr.py recognise ``` docTR's `forward` runs CTC decoding in numpy whether or not it was asked to, which `torch.export` refuses. The model carries an `exportable` flag that skips it and returns raw logits; the wrapper sets that. **Feed it contiguous tensors.** ExecuTorch reads a tensor in memory order rather than by its strides, and a crop built with `np.transpose` is not contiguous. The detector half of this pair scored 1.000000 against eager on random input and then returned a map that thresholded to zero boxes on a real photograph, entirely because of that. (conversion scripts: [executorch-models](https://github.com/john-rocky/executorch-models))