metadata
license: cc-by-nc-sa-4.0
base_model: microsoft/layoutlmv3-base
datasets:
- nielsr/funsd-layoutlmv3
metrics:
- f1
TrustDoc Field Extractor
LayoutLMv3-base fine-tuned for token classification (HEADER/QUESTION/ANSWER field extraction on forms), part of the TrustDoc project -- a document AI trust layer with calibrated confidence for human-in-the-loop review.
- Base model: microsoft/layoutlmv3-base (CC BY-NC-SA 4.0 -- non-commercial use only, inherited by this fine-tune)
- Training data: nielsr/funsd-layoutlmv3 (149 train / 50 test forms)
- Test set micro-avg F1: 0.862
Limitations
FUNSD is a small (199-document) benchmark of noisy scanned forms -- strong results here don't guarantee generalization to arbitrary document layouts. See the TrustDoc repo for the full pipeline (OCR -> classify -> extract -> calibrate) and its limitations.