File size: 3,210 Bytes
e9232c8
 
 
 
 
8eaf552
 
 
 
 
 
 
e9232c8
8eaf552
e9232c8
 
 
 
 
 
 
 
 
fd4ddf4
e9232c8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8eaf552
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
---
base_model: minishlab/potion-base-32m
library_name: model2vec
license: mit
tags:
- model2vec
- static-embeddings
- text-classification
- newspaper-classification
- crop-classification
- historical-newspapers
- newspapers
datasets:
- institutional/institutional-newspapers-bpl
---

# 📰 Institutional Newspapers Crop Classifier (Text)

A text-based classifier that categorizes crops extracted from historical newspaper scans into high-level categories. This model is a [Model2Vec](https://github.com/MinishLab/model2vec) fine-tune of [minishlab/potion-base-32m](https://huggingface.co/minishlab/potion-base-32m) with a classifier head, which makes it light and efficient.

We recommend using this model alongside [`institutional/institutional-newspapers-crop-classifier-image-yolo26m-cls`](https://huggingface.co/institutional/institutional-newspapers-crop-classifier-image-yolo26m-cls), as neither visual nor textual signal alone is sufficient to classify newspaper crops accurately.

**More information:**
- 📄 [Full analysis available in our technical report](https://arxiv.org/abs/2608.18972)

**See also:**
- 🗂️ [Institutional Newspapers Collection](https://huggingface.co/collections/institutional/institutional-newspapers)
- ⚙️ [Pipeline](https://github.com/institutional/institutional-newspapers-pipeline)
- 🤖 [Segmentation model](https://huggingface.co/institutional/institutional-newspapers-segmenter-yolo26x)
- 🤖 [Crop-type image classifier](https://huggingface.co/institutional/institutional-newspapers-crop-classifier-image-yolo26m-cls)

The Institutional Data Initiative at the Harvard Law School Library works with knowledge institutions—from libraries and museums to cultural groups and government agencies—to refine and publish their collections as data. [Reach out to collaborate on your collections](https://institutional.org/).

## Evaluation results

Evaluated on a held-out test set of 31,015 crops:

| Class | Precision | Recall | F1-Score | Support |
|---|---|---|---|---|
| Advertisement | 0.95 | 0.94 | 0.95 | 11,244 |
| Cartoon | 0.85 | 0.57 | 0.68 | 430 |
| Content | 0.90 | 0.96 | 0.93 | 11,245 |
| Masthead, nameplate or running head | 0.93 | 0.86 | 0.89 | 2,250 |
| Photograph or illustration | 0.63 | 0.38 | 0.48 | 551 |
| Section heading | 0.94 | 0.93 | 0.94 | 5,295 |

| Metric | Score |
|---|---|
| **Accuracy** | 0.92 |
| **Macro avg F1** | 0.81 |
| **Weighted avg F1** | 0.92 |

## Training data

This model was trained on part of **Boston Public Library's** newspapers collection.

- **Total annotated crops**: 185,900 
- **Classes**: 6 
- **Split**: ~67% train / ~17% val / ~17% test 

## Usage

```python
from model2vec.inference import StaticModelPipeline

model_name = "institutional/institutional-newspapers-crop-classifier-text-model2vec"
model = StaticModelPipeline.from_pretrained(model_name)

texts = ["CLASSIFIED ADVERTISING — Rooms for rent, furnished apartments..."]
predictions = model.predict(texts, max_length=None)

print(predictions)
# ['Advertisement']
```

### Recommended inference parameters

| Parameter | Value | Note |
|---|---|---|
| `max_length` | `None` | Do not truncate input text |

## Citation

```bibtex
TODO
```