RVL-CDIP document classifiers, with and without identification codes

Seven image classifiers and six linear SVM baselines trained on RVL-CDIP (16 document categories), each in four versions that differ in the labels and in the images they were trained on, and three text and layout models (LayoutLM, BERT, RoBERTa) in the two corrected-label versions:

folder labels training images
labels-original_codes-kept original RVL-CDIP labels de-identified pages
labels-original_codes-removed original RVL-CDIP labels de-identified pages with the identification codes whitened out
labels-corrected_codes-kept corrected labels de-identified pages
labels-corrected_codes-removed corrected labels de-identified pages with the identification codes whitened out

Identification codes are the Bates numbers stamped on most pages; they are associated with the category label and act as a shortcut (see stefan-hf/rvlcdip-id-codes). The codes were located with stefan-hf/yolov8n-rvlcdip-idcodes. All models were trained on a de-identified copy of the corpus, in which personal data had been replaced with synthetic values, so none of them was trained on the original personal data.

The codes-removed models cannot use the codes, at a cost of 0.0–0.3 points in-domain on the corrected labels. On RVL-CDIP-N, which was collected from other sources, the codes-kept models are 1–6 points more accurate (second table below), so neither version is better on every measure.

Models and accuracy

Test accuracy (%) of the released checkpoint on the test split of its own training condition; in parentheses, the mean over three training seeds. The released checkpoint is always seed 42, the first run, chosen in advance and not by score. Models trained on the original labels are scored on the original test labels (39,999 pages) and models trained on the corrected labels on the corrected test labels (36,446 pages), so the two halves of the table are not directly comparable.

model folder original, codes kept original, codes removed corrected, codes kept corrected, codes removed size
AlexNet alexnet 89.21 (89.30) 88.41 (88.52) 92.85 (92.82) 92.73 (92.73) 228 MB
GoogLeNet googlenet 89.34 (89.25) 88.52 (88.50) 93.19 (93.15) 92.91 (92.97) 23 MB
ResNet-50 resnet50 91.01 (90.85) 90.21 (90.20) 94.26 (94.31) 94.28 (94.19) 94 MB
ResNeXt-50 (32x4d) resnext50 91.48 (91.33) 90.48 (90.63) 94.64 (94.69) 94.56 (94.52) 92 MB
SqueezeNet 1.0 squeezenet 88.70 (88.66) 87.74 (87.69) 92.46 (92.59) 92.50 (92.46) 3 MB
VGG-16 vgg16 91.18 (91.27) 90.74 (90.81) 94.61 (94.62) 94.46 (94.42) 537 MB
DiT-base dit_base 92.69 (92.76) 92.20 (92.19) 95.94 (95.92) 95.74 (95.81) 343 MB

Corrected-label models on both test conditions and on RVL-CDIP-N (1,002 newly collected pages):

model kept β†’ kept kept β†’ removed removed β†’ kept removed β†’ removed RVL-CDIP-N, kept RVL-CDIP-N, removed
AlexNet 92.85 91.54 92.37 92.73 73.6 69.7
GoogLeNet 93.19 92.62 92.72 92.91 77.9 71.5
ResNet-50 94.26 93.71 94.19 94.28 79.2 77.0
ResNeXt-50 (32x4d) 94.64 94.10 94.45 94.56 80.8 76.2
SqueezeNet 1.0 92.46 91.57 92.32 92.50 77.0 75.0
VGG-16 94.61 93.75 94.25 94.46 78.5 77.8
DiT-base 95.94 95.37 95.71 95.74 87.1 83.9

"kept β†’ removed" is a model trained with codes and tested on pages without them.

Usage

rvlcdip_models.py in this repository builds the architecture, loads the weights, and applies the training-time preprocessing (grayscale, bilinear resize to 224 Γ— 224, replicate to three channels, normalize).

from huggingface_hub import hf_hub_download
from PIL import Image
import importlib.util, torch

spec = importlib.util.spec_from_file_location(
    "rvlcdip_models", hf_hub_download("stefan-hf/rvlcdip-classifiers", "rvlcdip_models.py"))
lib = importlib.util.module_from_spec(spec); spec.loader.exec_module(lib)

model, cfg = lib.load("resnet50", "labels-corrected_codes-removed")
x = lib.preprocess(Image.open("page.png"), cfg).unsqueeze(0)
with torch.no_grad():
    print(cfg["id2label"][str(model(x).argmax(1).item())])

The DiT-base folders are also in the transformers layout and load with AutoModelForImageClassification.from_pretrained(..., subfolder="dit_base/labels-corrected_codes-removed"). Each folder's config.json (rvlcdip_config.json for DiT) records the label map, the normalization, the selected epoch, and the accuracies above.

Linear SVM baselines

Six linear baselines trained on the same four configurations (one run each; liblinear is deterministic, so there are no seeds). The TF-IDF models read the OCR text of the page, not the image, and the fusion models read both. Test accuracy (%) on the test split of the model's own training condition:

model folder input original, codes kept original, codes removed corrected, codes kept corrected, codes removed size
TF-IDF + linear SVM svm_tfidf OCR text 88.95 88.78 93.40 93.35 4.1 MB
TF-IDF + handwriting tokens + linear SVM svm_tfidf_hw OCR text 89.77 89.72 94.65 94.64 4.1 MB
Linear SVM on CLIP ViT-L/14-336 embeddings svm_clip_vitl14_336 image 87.47 87.44 93.77 93.59 0.1 MB
TF-IDF + linear SVM, Tesseract text svm_tfidf_t411 OCR text (Tesseract) 79.81 79.71 84.68 84.51 4.0 MB
TF-IDF + CLIP embeddings, one linear SVM svm_fusion OCR text and image 92.35 92.27 96.76 96.70 4.1 MB
TF-IDF + handwriting tokens + CLIP embeddings, one linear SVM svm_fusion_hw OCR text and image 92.43 92.36 96.86 96.78 4.1 MB
  • TF-IDF. Word uni- and bigrams, 50,000 features, sublinear term frequency, one-vs-rest LinearSVC, with C chosen on validation accuracy.
  • Handwriting tokens. The same model with two pseudo-tokens in front of the text: the number of words on the page in logarithmic bins, and the share of words that Textract marks as handwriting.
  • CLIP embeddings. The projected image embedding of the grayscale page from the frozen openai/clip-vit-large-patch14-336, L2-normalized and standardized, then the same SVM.
  • Tesseract text. The TF-IDF model trained and tested on Tesseract 4.1.1 (--psm 3) text in place of Amazon Textract text. All other text models here expect Textract text.
  • Fusion. One SVM on the TF-IDF features and the CLIP features of the same page, concatenated. The standardized CLIP block is divided by √768, so that its rows have about the same norm as the TF-IDF rows, and multiplied by a weight alpha; alpha and C are chosen together on validation accuracy and recorded in config.json. svm_fusion_hw adds the handwriting tokens to the text. With the corrected labels, the fusion models reach 95.1–95.6% on RVL-CDIP-N, against 85.7–86.1% for svm_tfidf, 92.1% for svm_tfidf_hw, and 88.3–88.7% for svm_clip_vitl14_336.

The weights are stored as safetensors and the vocabulary as JSON; rvlcdip_svms.py rebuilds the models with scikit-learn and no pickled objects.

from huggingface_hub import hf_hub_download
import importlib.util

spec = importlib.util.spec_from_file_location(
    "rvlcdip_svms", hf_hub_download("stefan-hf/rvlcdip-classifiers", "rvlcdip_svms.py"))
svms = importlib.util.module_from_spec(spec); spec.loader.exec_module(svms)

text_svm = svms.load_tfidf_svm("svm_tfidf", "labels-corrected_codes-removed")
print(text_svm.predict(["Dear Mr. Smith, thank you for your letter of March 3 ..."]))

image_svm = svms.load_clip_svm("labels-corrected_codes-removed")   # downloads CLIP
# image_svm.predict([PIL.Image.open("page.png")])

fusion_svm = svms.load_fusion_svm("svm_fusion", "labels-corrected_codes-removed")   # downloads CLIP
# fusion_svm.predict(["text of the page ..."], [PIL.Image.open("page.png")])

For svm_tfidf_hw and svm_fusion_hw, put svms.hw_prefix(n_words, n_handwritten_words) (or svms.hw_prefix_from_textract(response_text)) in front of each page's text. svm_tfidf_t411 loads with load_tfidf_svm like the other text models. The decision values are uncalibrated one-vs-rest margins.

Text and layout models

Three fine-tuned transformers that read the OCR output of the page (Amazon Textract) and not the image: BERT and RoBERTa read the text, and LayoutLM reads the words together with their bounding boxes. Only the two corrected-label versions are included at present. For these models codes-removed means that the words of the identification codes were deleted from the OCR output the model was trained on. Test accuracy (%) of the released seed-42 checkpoint on the corrected test labels (36,446 pages) under both test conditions, and on RVL-CDIP-N:

model folder input kept β†’ kept kept β†’ removed removed β†’ kept removed β†’ removed RVL-CDIP-N, kept RVL-CDIP-N, removed size
LayoutLM (v1, base) layoutlmv1 OCR words and boxes 96.66 96.40 96.58 96.55 96.6 96.7 451 MB
BERT-base (uncased) bert_base OCR text 95.61 95.24 95.35 95.31 91.8 91.4 439 MB
RoBERTa-base roberta_base OCR text 96.02 95.68 95.85 95.77 92.3 92.9 502 MB

Over three training seeds, LayoutLM averages 96.67 (kept β†’ kept) and 96.57 (removed β†’ removed).

Each folder is in the transformers layout, with the tokenizer next to the weights, and loads with from_pretrained(..., subfolder=...). The example reads a page from stefan-hf/rvlcdip-redact, whose textract_* configurations contain the text, words, and boxes these models were trained on:

import torch
from datasets import load_dataset
from transformers import AutoModelForSequenceClassification, AutoTokenizer

repo = "stefan-hf/rvlcdip-classifiers"
page = next(iter(load_dataset("stefan-hf/rvlcdip-redact", "textract_redact_noid", split="test", streaming=True)))

# BERT or RoBERTa: the plain text of the page, first 512 tokens
folder = "roberta_base/labels-corrected_codes-removed"
tok = AutoTokenizer.from_pretrained(repo, subfolder=folder)
model = AutoModelForSequenceClassification.from_pretrained(repo, subfolder=folder).eval()
text = page["text"] + "\n" if page["text"] else ""
with torch.no_grad():
    logits = model(**tok(text, truncation=True, max_length=512, return_tensors="pt")).logits
print(model.config.id2label[logits.argmax(1).item()], "| label:", page["label_fixed"])

# LayoutLM: the words and their boxes (0-1000), first 512 tokens
folder = "layoutlmv1/labels-corrected_codes-removed"
tok = AutoTokenizer.from_pretrained(repo, subfolder=folder)
model = AutoModelForSequenceClassification.from_pretrained(repo, subfolder=folder).eval()
words, boxes = page["tokens"] or [""], page["bboxes"] or [[0, 0, 0, 0]]
enc = tok(words, is_split_into_words=True, truncation=True, max_length=512, return_tensors="pt")
bbox = torch.tensor([[[0, 0, 0, 0] if w is None else boxes[w] for w in enc.word_ids(0)]])
with torch.no_grad():
    logits = model(**enc, bbox=bbox).logits
print(model.config.id2label[logits.argmax(1).item()], "| label:", page["label_fixed"])

The text files used in training end with a newline on most pages, which the RoBERTa tokenizer turns into a token, so the example appends one; BERT's tokenizer ignores it. Each folder's rvlcdip_config.json records the base model, the selected epoch, and the accuracies above.

Training

  • Data. RVL-CDIP train split (319,999 pages; 296,605 with the corrected labels), model selected by validation accuracy after every epoch.
  • CNNs. ImageNet-pretrained torchvision models, SGD (learning rate 0.001, momentum 0.9, polynomial decay), batch 64, 30 epochs, cross-entropy.
  • DiT-base. microsoft/dit-base, AdamW (learning rate 3e-5, cosine schedule with warm-up), batch 32, 10 epochs.
  • LayoutLM, BERT, RoBERTa. microsoft/layoutlm-base-uncased, google-bert/bert-base-uncased, and FacebookAI/roberta-base, AdamW (learning rate 5e-5, linear warm-up over the first 10% of steps, then linear decay), gradient clipping at 1.0, batch 16, 5 epochs, full precision, the first 512 tokens of the page. All of these runs selected the last epoch.
  • Input. Grayscale page resized to 224 Γ— 224. At this size the identification codes are two or three pixels tall.

Each exported file was checked by reloading it through rvlcdip_models.py and comparing its predictions on 512 test pages with those of the original checkpoint (agreement 99.8–100%). The SVM exports were checked the same way through rvlcdip_svms.py (identical predictions). For the fusion models that check used the image embeddings stored at training time; computed again from the page images, the predictions agreed on 63 or 64 of 64 pages per model. The text and layout exports were reloaded with from_pretrained and reproduced the logits of the original checkpoints on 512 test pages.

Limitations

  • Trained and evaluated on tobacco-litigation scans only; accuracy on other document sources is lower (see the RVL-CDIP-N column).
  • The codes-kept models can rely on the identification codes and lose accuracy when the codes are absent.
  • RVL-CDIP has substantial overlap between its train and test splits, which inflates all test accuracies here.
  • The text and layout models were trained on Amazon Textract output and were not tested with other OCR engines; the TF-IDF model loses about 9 points when trained and tested on Tesseract text.
  • GoogLeNet must be built with transform_input=True and without auxiliary heads, as the loader does.

License

The weights, configuration files, and loader code in this repository are released under the MIT license.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Datasets used to train stefan-hf/rvlcdip-classifiers