Instructions to use flowxai/docformner with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use flowxai/docformner with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="flowxai/docformner")# Load model directly from transformers import AutoProcessor, AutoModelForTokenClassification processor = AutoProcessor.from_pretrained("flowxai/docformner") model = AutoModelForTokenClassification.from_pretrained("flowxai/docformner", device_map="auto") - Notebooks
- Google Colab
- Kaggle
DocFormNER
DocFormNER is a small, on-device cross NER model from the FlowX OpenNER family. Developed by FlowX.AI. Runs 100% on-premise / air-gapped, so no data leaves your boundary.
What it does
- Task: token-classification
- Base model:
microsoft/layoutlmv3-base - Entity types (9): APP_NO, BORROWER, INCOME, INTEREST_RATE, LENDER, LOAN_AMOUNT, LTV, PROPERTY_ADDR, TERM_YEARS
- Held-out F1: 1.0000
- Runtime: CPU, Apple Silicon, one GPU, or browser/edge via ONNX (INT8). ~100-160 ms/doc on CPU.
Why a small model
Fine-tuned encoders match or beat frontier LLMs on structured, convention-bound extraction, at a fraction of the latency and cost, with zero data egress. Identifiers are validated by checksum (IBAN mod-97, card Luhn, ISIN/LEI, container ISO-6346, VIN, national IDs), a correctness guarantee general LLMs lack. See the FlowX OpenNER benchmark for measured results.
Document-AI preview. DocFormNER is a LayoutLMv3 model that reads text + 2D layout + the page image, so it works on scanned / photographed forms (mortgage packets, bills of lading, ACORD forms). This checkpoint is trained on synthetic rendered forms as a walking skeleton; for production, fine-tune on real OCR'd documents (FUNSD / DocLayNet / CORD or your own scans) with the same
{words, bboxes, labels, image}shape.
Usage
from transformers import AutoTokenizer, AutoModelForTokenClassification
tok = AutoTokenizer.from_pretrained("flowxai/docformner")
model = AutoModelForTokenClassification.from_pretrained("flowxai/docformner")
License & attribution
Licensed under the Apache License 2.0. Copyright 2026 FlowX.AI (https://flowx.ai). See the NOTICE file. Trained on synthetic, checksum-validated data.
Part of the FlowX OpenNER model family. Synthetic-data F1 reflects an in-distribution synthetic distribution; validate on real documents before production use.
- Downloads last month
- 1
Model tree for flowxai/docformner
Base model
microsoft/layoutlmv3-base