Impresso newspaper image classifier (multi-facet L0-L3)

Classifies image regions cropped from historical newspapers along four independent facets, with one shared google/siglip2-so400m-patch14-384 backbone and one linear classifier of 28 outputs: each facet is a fixed slice of the logits (one linear head per facet, stored side by side). Part of the Impresso project.

  • L0: visual content: Image, Not Image
  • L1: technique: Not photograph, Photograph
  • L2: communication goal: Advertising, Decorative, Entertainment, Informative or Illustrative
  • L3: type: Caricature or Editorial cartoon or Humoristic drawing, Comic strip, Game, Graph, Human representations - Fashion visual, Human representations - Portrait, Human representations - Scene, Illustrated story, Map - Geological map, Map - Geopolitical map, Map - Physical map or Roadmap, Map - Plan, Map - Weather map, Non-figurative visual content, Object, Ornament or Illustrated Title, Other, Scenery or Landscape, Technical drawing, Weather infographic

Every facet is decided by argmax, including L0. L1-L3 apply only to crops that L0 classifies as Image. They were trained on image rows only, so their output on a non-image is meaningless and must be discarded.

Test results

Held-out test split of impresso-project/newspaper-image-classification. Each facet is scored on its own eligible rows: L1-L3 on gold Image rows.

facet classes accuracy macro-F1
L0: visual content 2 0.9945 0.9895
L1: technique 2 0.9647 0.9646
L2: communication goal 4 0.9163 0.8709
L3: type 20 0.8092 0.8082

L0 as a filter: image recall (retention) 0.9974, not-image precision 0.9858, not-image recall 0.9789, AUROC 0.9991.

Classes with no test support (not evaluated): l3/Map - Geological map, l3/Map - Plan.

Training

Recipe

  • What trains: backbone frozen; only the 28-way classifier trains (32,284 parameters).
  • Classifier input: mean of all output tokens (SigLIP's attention-pooling head is not used).
  • Optimizer: AdamW, learning rate 0.001, weight decay 0.05.
  • Schedule: cosine schedule with 5% warmup, 20 epochs, batch 64, fp16 mixed precision.
  • Loss: sum of per-facet cross-entropies (λ = 1), L1-L3 masked on Not Image rows; label smoothing 0.
  • Data: augmentation none, sampling none, seed 42.

The full resolved config is in resolved_config.yaml.

Benchmark of the recipe

The recipe was chosen on a multi-seed benchmark: trained on the train split only, best validation epoch kept, seeds 42, 123, 456 (wandb group siglip2-b-hf-bs64-lr1e3). Mean ± std over seeds:

metric validation test
mean macro-F1 (L1-L3) 0.8584 ± 0.0015 0.8638 ± 0.0022
mean accuracy (L1-L3) 0.8984 ± 0.0004 0.8905 ± 0.0018
L0 macro-F1 0.9867 ± 0.0007 0.9895 ± 0.0000
L1 macro-F1 0.9673 ± 0.0013 0.9659 ± 0.0013
L2 macro-F1 0.8434 ± 0.0009 0.8635 ± 0.0026
L3 macro-F1 0.7645 ± 0.0057 0.7619 ± 0.0067

This model

  • Final run: train + validation splits merged (7812 rows), a fixed 20 epochs (the budget from the multi-seed benchmark), no model selection on held-out data. Test was untouched until this evaluation.
  • Data: impresso-project/newspaper-image-classification @ f92e9c1b383b0404b61630e507f58932bb66d70b (config default).
  • Code: impresso-newspaper-image-classification @ 045f1da.
  • Weights SHA-256: b253c6d42810d299eef4634fd1f3c113d49ab8bb7fa178ac48229b4995f99ed2.

Input contract

Aspect-preserving LANCZOS resize and center padding to a 384x384 square (pad colour [0, 0, 0]), rescale to [0, 1], then normalize with mean [0.5, 0.5, 0.5] / std [0.5, 0.5, 0.5]. The stock image processor does not pad, so pad first (pad_to_square below, identical to the training code); on a padded square its resize is a no-op.

Usage

A standard transformers image classifier: no custom code, no trust_remote_code.

import torch
from PIL import Image
from transformers import AutoImageProcessor, AutoModelForImageClassification

repo = "impresso-project/image-classification-siglip2"
model = AutoModelForImageClassification.from_pretrained(repo).eval()
processor = AutoImageProcessor.from_pretrained(repo)
cfg = model.config
prep = cfg.impresso_preprocessing


def pad_to_square(img, size, color):
    # Aspect-preserving LANCZOS resize, then center padding (the training contract).
    scale = min(size / img.width, size / img.height)
    w, h = max(1, int(img.width * scale)), max(1, int(img.height * scale))
    out = Image.new("RGB", (size, size), tuple(color))
    out.paste(img.resize((w, h), Image.Resampling.LANCZOS), ((size - w) // 2, (size - h) // 2))
    return out


img = pad_to_square(Image.open("crop.jpg").convert("RGB"), prep["image_size"], prep["pad_color"])
with torch.no_grad():
    logits = model(**processor(img, return_tensors="pt")).logits[0]

# One softmax per facet over its slice of the 28 logits ({'l0': [0, 2], 'l1': [2, 4], 'l2': [4, 8], 'l3': [8, 28]}).
pred = {}
for facet, (a, b) in cfg.facet_slices.items():
    probs = logits[a:b].softmax(-1)
    pred[facet] = (cfg.facets[facet][int(probs.argmax())], float(probs.max()))
if pred["l0"][0] != "Image":
    pred = {"l0": pred["l0"]}  # L1-L3 are meaningless for a non-image

Do not use pipeline("image-classification"): it applies one softmax over all 28 labels. For Image / Not Image filtering only, read the l0 slice: every facet comes from the same backbone pass, so skipping the others saves nothing.

Files

model.safetensors (weights), config.json (architecture, label vocabularies facets and their logit slices facet_slices, input contract impresso_preprocessing), preprocessor_config.json, resolved_config.yaml (training config + data/code revisions), test_results.json, model.safetensors.sha256.

Limitations

  • Trained on Impresso's annotated newspaper crops; quality on other collections, periods or scan pipelines is not measured.
  • Small classes have few test examples, so their per-class scores are noisy; classes without test support are not evaluated at all.
  • L2 (communication goal) labels can be inherited from L3 correspondences, so agreement with them does not independently validate those rules.
  • The per-facet softmax probabilities are not calibrated (no temperature scaling): use them to rank, not as true probabilities.

License

The weights follow the backbone licence (apache-2.0). Rights on the training images are described on the dataset card.

Downloads last month
31
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for impresso-project/image-classification-siglip2

Finetuned
(38)
this model

Evaluation results

  • Test mean macro-F1 over L1/L2/L3 on impresso-project/newspaper-image-classification
    test set self-reported
    0.881