RVL-CDIP document classifiers, with and without identification codes
Seven image classifiers and six linear SVM baselines trained on RVL-CDIP (16 document categories), each in four versions that differ in the labels and in the images they were trained on, and three text and layout models (LayoutLM, BERT, RoBERTa) in the two corrected-label versions:
| folder | labels | training images |
|---|---|---|
labels-original_codes-kept |
original RVL-CDIP labels | de-identified pages |
labels-original_codes-removed |
original RVL-CDIP labels | de-identified pages with the identification codes whitened out |
labels-corrected_codes-kept |
corrected labels | de-identified pages |
labels-corrected_codes-removed |
corrected labels | de-identified pages with the identification codes whitened out |
Identification codes are the Bates numbers stamped on most pages; they are associated with the
category label and act as a shortcut (see stefan-hf/rvlcdip-id-codes).
The codes were located with stefan-hf/yolov8n-rvlcdip-idcodes.
All models were trained on a de-identified copy of the corpus, in which personal data had been
replaced with synthetic values, so none of them was trained on the original personal data.
The codes-removed models cannot use the codes, at a cost of 0.0β0.3 points in-domain on the
corrected labels. On RVL-CDIP-N, which was collected from other sources, the codes-kept models
are 1β6 points more accurate (second table below), so neither version is better on every measure.
Models and accuracy
Test accuracy (%) of the released checkpoint on the test split of its own training condition; in parentheses, the mean over three training seeds. The released checkpoint is always seed 42, the first run, chosen in advance and not by score. Models trained on the original labels are scored on the original test labels (39,999 pages) and models trained on the corrected labels on the corrected test labels (36,446 pages), so the two halves of the table are not directly comparable.
| model | folder | original, codes kept | original, codes removed | corrected, codes kept | corrected, codes removed | size |
|---|---|---|---|---|---|---|
| AlexNet | alexnet |
89.21 (89.30) | 88.41 (88.52) | 92.85 (92.82) | 92.73 (92.73) | 228 MB |
| GoogLeNet | googlenet |
89.34 (89.25) | 88.52 (88.50) | 93.19 (93.15) | 92.91 (92.97) | 23 MB |
| ResNet-50 | resnet50 |
91.01 (90.85) | 90.21 (90.20) | 94.26 (94.31) | 94.28 (94.19) | 94 MB |
| ResNeXt-50 (32x4d) | resnext50 |
91.48 (91.33) | 90.48 (90.63) | 94.64 (94.69) | 94.56 (94.52) | 92 MB |
| SqueezeNet 1.0 | squeezenet |
88.70 (88.66) | 87.74 (87.69) | 92.46 (92.59) | 92.50 (92.46) | 3 MB |
| VGG-16 | vgg16 |
91.18 (91.27) | 90.74 (90.81) | 94.61 (94.62) | 94.46 (94.42) | 537 MB |
| DiT-base | dit_base |
92.69 (92.76) | 92.20 (92.19) | 95.94 (95.92) | 95.74 (95.81) | 343 MB |
Corrected-label models on both test conditions and on RVL-CDIP-N (1,002 newly collected pages):
| model | kept β kept | kept β removed | removed β kept | removed β removed | RVL-CDIP-N, kept | RVL-CDIP-N, removed |
|---|---|---|---|---|---|---|
| AlexNet | 92.85 | 91.54 | 92.37 | 92.73 | 73.6 | 69.7 |
| GoogLeNet | 93.19 | 92.62 | 92.72 | 92.91 | 77.9 | 71.5 |
| ResNet-50 | 94.26 | 93.71 | 94.19 | 94.28 | 79.2 | 77.0 |
| ResNeXt-50 (32x4d) | 94.64 | 94.10 | 94.45 | 94.56 | 80.8 | 76.2 |
| SqueezeNet 1.0 | 92.46 | 91.57 | 92.32 | 92.50 | 77.0 | 75.0 |
| VGG-16 | 94.61 | 93.75 | 94.25 | 94.46 | 78.5 | 77.8 |
| DiT-base | 95.94 | 95.37 | 95.71 | 95.74 | 87.1 | 83.9 |
"kept β removed" is a model trained with codes and tested on pages without them.
Usage
rvlcdip_models.py in this repository builds the architecture, loads the weights, and applies the
training-time preprocessing (grayscale, bilinear resize to 224 Γ 224, replicate to three channels,
normalize).
from huggingface_hub import hf_hub_download
from PIL import Image
import importlib.util, torch
spec = importlib.util.spec_from_file_location(
"rvlcdip_models", hf_hub_download("stefan-hf/rvlcdip-classifiers", "rvlcdip_models.py"))
lib = importlib.util.module_from_spec(spec); spec.loader.exec_module(lib)
model, cfg = lib.load("resnet50", "labels-corrected_codes-removed")
x = lib.preprocess(Image.open("page.png"), cfg).unsqueeze(0)
with torch.no_grad():
print(cfg["id2label"][str(model(x).argmax(1).item())])
The DiT-base folders are also in the transformers layout and load with
AutoModelForImageClassification.from_pretrained(..., subfolder="dit_base/labels-corrected_codes-removed").
Each folder's config.json (rvlcdip_config.json for DiT) records the label map, the
normalization, the selected epoch, and the accuracies above.
Linear SVM baselines
Six linear baselines trained on the same four configurations (one run each; liblinear is deterministic, so there are no seeds). The TF-IDF models read the OCR text of the page, not the image, and the fusion models read both. Test accuracy (%) on the test split of the model's own training condition:
| model | folder | input | original, codes kept | original, codes removed | corrected, codes kept | corrected, codes removed | size |
|---|---|---|---|---|---|---|---|
| TF-IDF + linear SVM | svm_tfidf |
OCR text | 88.95 | 88.78 | 93.40 | 93.35 | 4.1 MB |
| TF-IDF + handwriting tokens + linear SVM | svm_tfidf_hw |
OCR text | 89.77 | 89.72 | 94.65 | 94.64 | 4.1 MB |
| Linear SVM on CLIP ViT-L/14-336 embeddings | svm_clip_vitl14_336 |
image | 87.47 | 87.44 | 93.77 | 93.59 | 0.1 MB |
| TF-IDF + linear SVM, Tesseract text | svm_tfidf_t411 |
OCR text (Tesseract) | 79.81 | 79.71 | 84.68 | 84.51 | 4.0 MB |
| TF-IDF + CLIP embeddings, one linear SVM | svm_fusion |
OCR text and image | 92.35 | 92.27 | 96.76 | 96.70 | 4.1 MB |
| TF-IDF + handwriting tokens + CLIP embeddings, one linear SVM | svm_fusion_hw |
OCR text and image | 92.43 | 92.36 | 96.86 | 96.78 | 4.1 MB |
- TF-IDF. Word uni- and bigrams, 50,000 features, sublinear term frequency, one-vs-rest
LinearSVC, withCchosen on validation accuracy. - Handwriting tokens. The same model with two pseudo-tokens in front of the text: the number of words on the page in logarithmic bins, and the share of words that Textract marks as handwriting.
- CLIP embeddings. The projected image embedding of the grayscale page from the frozen
openai/clip-vit-large-patch14-336, L2-normalized and standardized, then the same SVM. - Tesseract text. The TF-IDF model trained and tested on Tesseract 4.1.1 (
--psm 3) text in place of Amazon Textract text. All other text models here expect Textract text. - Fusion. One SVM on the TF-IDF features and the CLIP features of the same page, concatenated.
The standardized CLIP block is divided by β768, so that its rows have about the same norm as the
TF-IDF rows, and multiplied by a weight
alpha;alphaandCare chosen together on validation accuracy and recorded inconfig.json.svm_fusion_hwadds the handwriting tokens to the text. With the corrected labels, the fusion models reach 95.1β95.6% on RVL-CDIP-N, against 85.7β86.1% forsvm_tfidf, 92.1% forsvm_tfidf_hw, and 88.3β88.7% forsvm_clip_vitl14_336.
The weights are stored as safetensors and the vocabulary as JSON; rvlcdip_svms.py rebuilds the
models with scikit-learn and no pickled objects.
from huggingface_hub import hf_hub_download
import importlib.util
spec = importlib.util.spec_from_file_location(
"rvlcdip_svms", hf_hub_download("stefan-hf/rvlcdip-classifiers", "rvlcdip_svms.py"))
svms = importlib.util.module_from_spec(spec); spec.loader.exec_module(svms)
text_svm = svms.load_tfidf_svm("svm_tfidf", "labels-corrected_codes-removed")
print(text_svm.predict(["Dear Mr. Smith, thank you for your letter of March 3 ..."]))
image_svm = svms.load_clip_svm("labels-corrected_codes-removed") # downloads CLIP
# image_svm.predict([PIL.Image.open("page.png")])
fusion_svm = svms.load_fusion_svm("svm_fusion", "labels-corrected_codes-removed") # downloads CLIP
# fusion_svm.predict(["text of the page ..."], [PIL.Image.open("page.png")])
For svm_tfidf_hw and svm_fusion_hw, put svms.hw_prefix(n_words, n_handwritten_words) (or
svms.hw_prefix_from_textract(response_text)) in front of each page's text. svm_tfidf_t411
loads with load_tfidf_svm like the other text models. The decision values are uncalibrated
one-vs-rest margins.
Text and layout models
Three fine-tuned transformers that read the OCR output of the page (Amazon Textract) and not the
image: BERT and RoBERTa read the text, and LayoutLM reads the words together with their bounding
boxes. Only the two corrected-label versions are included at present. For these models
codes-removed means that the words of the identification codes were deleted from the OCR output
the model was trained on. Test accuracy (%) of the released seed-42 checkpoint on the corrected
test labels (36,446 pages) under both test conditions, and on RVL-CDIP-N:
| model | folder | input | kept β kept | kept β removed | removed β kept | removed β removed | RVL-CDIP-N, kept | RVL-CDIP-N, removed | size |
|---|---|---|---|---|---|---|---|---|---|
| LayoutLM (v1, base) | layoutlmv1 |
OCR words and boxes | 96.66 | 96.40 | 96.58 | 96.55 | 96.6 | 96.7 | 451 MB |
| BERT-base (uncased) | bert_base |
OCR text | 95.61 | 95.24 | 95.35 | 95.31 | 91.8 | 91.4 | 439 MB |
| RoBERTa-base | roberta_base |
OCR text | 96.02 | 95.68 | 95.85 | 95.77 | 92.3 | 92.9 | 502 MB |
Over three training seeds, LayoutLM averages 96.67 (kept β kept) and 96.57 (removed β removed).
Each folder is in the transformers layout, with the tokenizer next to the weights, and loads with
from_pretrained(..., subfolder=...). The example reads a page from
stefan-hf/rvlcdip-redact, whose
textract_* configurations contain the text, words, and boxes these models were trained on:
import torch
from datasets import load_dataset
from transformers import AutoModelForSequenceClassification, AutoTokenizer
repo = "stefan-hf/rvlcdip-classifiers"
page = next(iter(load_dataset("stefan-hf/rvlcdip-redact", "textract_redact_noid", split="test", streaming=True)))
# BERT or RoBERTa: the plain text of the page, first 512 tokens
folder = "roberta_base/labels-corrected_codes-removed"
tok = AutoTokenizer.from_pretrained(repo, subfolder=folder)
model = AutoModelForSequenceClassification.from_pretrained(repo, subfolder=folder).eval()
text = page["text"] + "\n" if page["text"] else ""
with torch.no_grad():
logits = model(**tok(text, truncation=True, max_length=512, return_tensors="pt")).logits
print(model.config.id2label[logits.argmax(1).item()], "| label:", page["label_fixed"])
# LayoutLM: the words and their boxes (0-1000), first 512 tokens
folder = "layoutlmv1/labels-corrected_codes-removed"
tok = AutoTokenizer.from_pretrained(repo, subfolder=folder)
model = AutoModelForSequenceClassification.from_pretrained(repo, subfolder=folder).eval()
words, boxes = page["tokens"] or [""], page["bboxes"] or [[0, 0, 0, 0]]
enc = tok(words, is_split_into_words=True, truncation=True, max_length=512, return_tensors="pt")
bbox = torch.tensor([[[0, 0, 0, 0] if w is None else boxes[w] for w in enc.word_ids(0)]])
with torch.no_grad():
logits = model(**enc, bbox=bbox).logits
print(model.config.id2label[logits.argmax(1).item()], "| label:", page["label_fixed"])
The text files used in training end with a newline on most pages, which the RoBERTa tokenizer
turns into a token, so the example appends one; BERT's tokenizer ignores it. Each folder's
rvlcdip_config.json records the base model, the selected epoch, and the accuracies above.
Training
- Data. RVL-CDIP train split (319,999 pages; 296,605 with the corrected labels), model selected by validation accuracy after every epoch.
- CNNs. ImageNet-pretrained torchvision models, SGD (learning rate 0.001, momentum 0.9, polynomial decay), batch 64, 30 epochs, cross-entropy.
- DiT-base.
microsoft/dit-base, AdamW (learning rate 3e-5, cosine schedule with warm-up), batch 32, 10 epochs. - LayoutLM, BERT, RoBERTa.
microsoft/layoutlm-base-uncased,google-bert/bert-base-uncased, andFacebookAI/roberta-base, AdamW (learning rate 5e-5, linear warm-up over the first 10% of steps, then linear decay), gradient clipping at 1.0, batch 16, 5 epochs, full precision, the first 512 tokens of the page. All of these runs selected the last epoch. - Input. Grayscale page resized to 224 Γ 224. At this size the identification codes are two or three pixels tall.
Each exported file was checked by reloading it through rvlcdip_models.py and comparing its
predictions on 512 test pages with those of the original checkpoint (agreement 99.8β100%).
The SVM exports were checked the same way through rvlcdip_svms.py (identical predictions). For
the fusion models that check used the image embeddings stored at training time; computed again
from the page images, the predictions agreed on 63 or 64 of 64 pages per model. The text and layout
exports were reloaded with from_pretrained and reproduced the logits of the original checkpoints
on 512 test pages.
Limitations
- Trained and evaluated on tobacco-litigation scans only; accuracy on other document sources is lower (see the RVL-CDIP-N column).
- The
codes-keptmodels can rely on the identification codes and lose accuracy when the codes are absent. - RVL-CDIP has substantial overlap between its train and test splits, which inflates all test accuracies here.
- The text and layout models were trained on Amazon Textract output and were not tested with other OCR engines; the TF-IDF model loses about 9 points when trained and tested on Tesseract text.
- GoogLeNet must be built with
transform_input=Trueand without auxiliary heads, as the loader does.
License
The weights, configuration files, and loader code in this repository are released under the MIT license.