Instructions to use lemoncoda/bertforocr-continuity-ep5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lemoncoda/bertforocr-continuity-ep5 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="lemoncoda/bertforocr-continuity-ep5")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("lemoncoda/bertforocr-continuity-ep5") model = AutoModelForSequenceClassification.from_pretrained("lemoncoda/bertforocr-continuity-ep5", device_map="auto") - Notebooks
- Google Colab
- Kaggle
# Load model directly
from transformers import AutoTokenizer, AutoModelForSequenceClassification
tokenizer = AutoTokenizer.from_pretrained("lemoncoda/bertforocr-continuity-ep5")
model = AutoModelForSequenceClassification.from_pretrained("lemoncoda/bertforocr-continuity-ep5", device_map="auto")BERT for OCR cross-page continuity โ EP5
This is a binary XLM-RoBERTa/MiniLM classifier fine-tuned to decide whether two adjacent OCR fragments should be merged across a page boundary.
Labels
0:not_continuousโ do not merge1:continuousโ merge
The model output is converted to P(continuous) with softmax. For conservative
merging, use P(continuous) > 0.8, equivalent to score = 2P - 1 > 0.6.
Training configuration
- Base model:
MoritzLaurer/multilingual-MiniLMv2-L6-mnli-xnli - Epoch checkpoint: 5
- Training/validation pairs: 18,049 / 2,005
- Learning rate:
2e-5 - Batch size:
8 - Maximum sequence length:
512tokens - Boundary character window: previous tail
800, next head800 - Random seed:
42 - HTML table markup was preserved during training and validation.
Long pairs use boundary-aware truncation: retain the tail tokens of the previous fragment and the head tokens of the next fragment. With two sequences and a 512-token limit, the normal allocation is 254 tokens from each side; unused space from one side is transferred to the other.
Validation results
The values below are measured on the fixed 2,005-pair validation split. This is not an independent test set.
At the default probability threshold of 0.5:
| Accuracy | Precision | Recall | F1 | TN | FP | FN | TP |
|---|---|---|---|---|---|---|---|
| 0.894 | 0.878 | 0.900 | 0.889 | 941 | 118 | 95 | 851 |
At the recommended conservative threshold of 0.8:
| Group | N | Precision | Recall | F0.5 | FPR | TN | FP | FN | TP |
|---|---|---|---|---|---|---|---|---|---|
| Overall | 2,005 | 0.943 | 0.809 | 0.913 | 4.34% | 1,013 | 46 | 181 | 765 |
| Table | 1,808 | 0.946 | 0.809 | 0.915 | 4.63% | 865 | 42 | 172 | 729 |
| Text / OCRFlux bench | 197 | 0.900 | 0.800 | 0.878 | 2.63% | 148 | 4 | 9 | 36 |
The OCRFlux bench source group may include some table-like page elements and should not be interpreted as a perfectly pure prose-only subset.
Inference
Use the fine-tuned model as a two-class sequence-pair classifier and reproduce the boundary-aware truncation described above. The merge probability is:
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model_id = "lemoncoda/bertforocr-continuity-ep5"
tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True)
model = AutoModelForSequenceClassification.from_pretrained(model_id).eval()
def encode_boundary_pair(previous: str, following: str, max_length: int = 512):
previous_ids = tokenizer.encode(previous, add_special_tokens=False)
following_ids = tokenizer.encode(following, add_special_tokens=False)
budget = max_length - tokenizer.num_special_tokens_to_add(pair=True)
previous_budget = budget // 2
following_budget = budget - previous_budget
if len(previous_ids) < previous_budget:
following_budget += previous_budget - len(previous_ids)
if len(following_ids) < following_budget:
previous_budget += following_budget - len(following_ids)
previous_ids = previous_ids[-previous_budget:]
following_ids = following_ids[:following_budget]
return tokenizer.prepare_for_model(
previous_ids,
following_ids,
add_special_tokens=True,
max_length=max_length,
truncation="longest_first",
return_tensors="pt",
)
encoded = encode_boundary_pair(previous_fragment, next_fragment)
with torch.inference_mode():
probability = torch.softmax(model(**encoded).logits, dim=-1)[0, 1].item()
should_merge = probability > 0.8
score = 2 * probability - 1
For exact project integration, preserve the original HTML markup in table fragments and apply the 800-character previous-tail/next-head windows before tokenization.
- Downloads last month
- 10
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="lemoncoda/bertforocr-continuity-ep5")