DeepSeek-OCR page classifier (architecture B)

Variant: unmerged. Base: deepseek-ai/DeepSeek-OCR at revision 9f30c71f441d010e5429c532364a86705536c53a.

Intended use: classify single rendered document pages into 7 operational page types, routing low-confidence pages to UNKNOWN for human review. Not intended for OCR, text generation or any decision without human oversight.

Architecture

page image -> 640x640 mean-grey pad -> Normalize(0.5, 0.5, 0.5)
  -> SAM ViT-B (+ neck) ---------------------------+
  -> CLIP-L (SAM features as patch embeddings) ----+-> concat [2048]
  -> trained linear projector [2048 -> 1280] -> 10x10 grid
  -> one image_newline per row + one view_seperator = 111 page tokens
[BOS row][111 page tokens][learned CLASSIFY token]
  -> 12-layer DeepSeek MoE decoder (Q/O LoRA r16 a32) -> final hidden state
  -> LayerNorm -> Dropout -> Linear [1280 -> 7] -> logits / T -> softmax
  -> confidence < threshold => UNKNOWN

Checkpoint stripping

Upstream tensor Disposition Reason
lm_head.weight dropped vocabulary projection unused
model.vision_model.embeddings.patch_embedding.weight dropped CLIP patch conv bypassed (SAM features are the patch embeddings)
model.projector.layers.{weight,bias} dropped superseded by trained projector
model.embed_tokens.weight sliced only the BOS row 0 is used (vocab_size=1)
model.layers.N.self_attn.{q,o}_proj.weight adapted LoRA base or merged
every other tensor kept BF16 bytes unchanged

Upstream tensors: 2,710; exported: 2,706; dropped: 4. Parameters: upstream 3,336,106,240, exported 3,001,925,888, removed 334,180,352. Every tensor is listed in stripped_inventory.json.

Trained parameters

trained.safetensors (F32) holds 55 tensors, 3,618,567 parameters.

Group Parameters
classification_token 1,280
head 11,527
lora 983,040
projector 2,622,720

Inputs and decisions

Preprocessing: RGB, contain-resize to 640 px, mean-grey (127) square pad, scale to [0, 1], Normalize(mean 0.5, std 0.5). PDF rasterization is not included.

Labels: Inspection_Report, Job_Paperwork, Other, Pictures, Service_Order, Vendor_Invoice, Vendor_Proposal. Calibration: softmax(logits / T) with T = 1.5147499307352106. Pages whose top probability is below 0.8 are labeled UNKNOWN; the best known label is kept as the runner-up.

Loading

Tested with transformers==4.46.3 (hard requirement) and torch==2.14.0, BF16 on CUDA, eager attention, native ATen dispatch, one physical page per forward pass.

from huggingface_hub import snapshot_download
from transformers.dynamic_module_utils import get_class_from_dynamic_module

folder = snapshot_download("vvs184/deepseek-ocr-page-classifier-b04", revision="main")
model_class = get_class_from_dynamic_module(
    "modeling_page_classifier.PageClassifierModel", folder
)
model = model_class.load_for_inference(folder, device="cuda")
predictions = model.classify([page_image])  # PIL RGB images

AutoModel.from_pretrained(folder, trust_remote_code=True) restores the same weights through the same strict loader (missing or unexpected tensors, other loading options and configuration overrides are refused); call model.prepare_for_inference("cuda") before use.

NOT PRODUCTION-QUALIFIED

This model has not passed its production quality gate. The blinded human gold set is unfilled, so the required floors (accuracy 0.99, macro-F1 0.985, supported critical recall 0.995) are unmet and unproven.

Pseudo-label measurements only (not human truth): accuracy 0.9663, macro-F1 0.8091, minimum critical recall 0.9706.

Limitations

  • Near-threshold sensitivity: pages close to the rejection threshold can flip between a label and UNKNOWN under different batching, padding or kernels.
  • Singleton requirement: serve one physical page per forward pass; grouped batches change MoE numerics and decisions.
  • PDF rasterization (PyMuPDF), upload limits and serving controls are not included; inputs must be rendered page images.
  • Pseudo-label metrics are engineering checks, not human quality evidence.

Parity

Repository serving path versus this export (content_sha256 in export_manifest.json), torch.bfloat16 on cuda. Unmerged accepted: True.

Variant Cohort Pages Decision changes Max abs logit diff Rejection changes Verdict
merged sample_first 512 0 0.3125 0 reported_only
merged sample_second 512 0 0.3125 0 reported_only
merged test 4689 5 0.875 5 reported_only
merged val 4848 9 0.75 9 reported_only
unmerged sample_first 512 0 0 0 True
unmerged sample_second 512 0 0 0 True
unmerged test 4689 0 0 0 True
unmerged val 4848 0 0 0 True

License

Upstream DeepSeek-OCR code and weights: MIT License, Copyright (c) 2023 DeepSeek (see LICENSE). Trained classifier weights and page-classifier code: license other, private, see NOTICE.

Downloads last month
15
Safetensors
Model size
3B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for vvs184/deepseek-ocr-page-classifier-b04

Finetuned
(131)
this model