c5k-deberta_base-token_level-1-2 (legacy: urchade/gliner_multi_pii-v1 continued on PII-in-spam)

Superseded. This is an early checkpoint, kept because it has a DOI (10.57967/hf/8514). For current pt-BR PII detection use arthrod/gliner-mmbert-small-ptbr-pii-full-3x-v1 or the OpenAI Privacy Filter fine-tune arthrod/gliner-opf-ptbr-pii-v1; results for both are in arthrod/gliner-opf-ptbr-pii-bench-v1.

TL;DR. This run was launched as b-gliner-pii-token-multipii-v1-spam-v1 (24 Feb 2026). The repo was first published as arthrod/gliner-pii-token-multipii-v1-spam-v1, and that URL still redirects here. It is a "domain adaptation for PII-in-spam" that continued urchade/gliner_multi_pii-v1 on the in-house dataset arthrod/gliner-flex-pii-ready-v5 (1,111,948 train / 526,449 eval examples) on one GPU. It was interrupted at step 6,500 of 12,000, about 0.19 epoch. Eval loss fell steadily from 48.9 (step 500) to 25.1 (step 6,500). The root weights are checkpoint-6500, the best and last checkpoint.

The name is wrong on two counts:

  1. The encoder is mDeBERTa-v3-base, not DeBERTa-v3-base. The YAML asked for microsoft/deberta-v3-base, but the run loaded the pretrained urchade/gliner_multi_pii-v1 (encoder microsoft/mdeberta-v3-base, vocab 250,105). The training log prints encoder_config … vocab_size: 250105 and model_name: 'microsoft/mdeberta-v3-base'. The 1.16 GB weight files have the same size as the mDeBERTa run in c3750-mdeberta_base-vanilla-1-2.
  2. The model is span-level (markerV0), not token_level. The log prints Model class: UniEncoderSpanGLiNER, span_mode: 'markerV0', model_type: 'gliner_uni_encoder_span'. The YAML overrides applied were only max_types 25→100, max_neg_type_ratio 3→2, dropout 0.4→0.25.

Do not load the repo root with GLiNER. The root gliner_config.json was hand-written on 2026-02-24 (commits "Create/Update gliner_config.json"). It describes a token-level DeBERTa-v3-base model that does not match the root pytorch_model.bin. The Feb-2026 sweep in arthrod/gliner_eval_folder (EVAL_REPORT.md) excluded this repo for "50 garbage preds, all scores ~0.50". That fits a config/weights mismatch. Load checkpoint-6500/ instead: it holds the same weights (identical LFS sha256) with the correct span-level mDeBERTa config.

c5k and 1-2 in the repo name are not explained anywhere in the repo.

Related repos

Repo Type Role
arthrod/gliner-mmbert-small-ptbr-pii-full-3x-v1 model Current recommended PT-BR PII GLiNER (mmBERT-small)
arthrod/gliner-opf-ptbr-pii-v1 model OpenAI Privacy Filter fine-tune for PT-BR PII (current suite)
arthrod/gliner-ettin-68m-ptbr-pii-full-3x-v1 · -top50k-v1 · ettin-32m model Current-suite Ettin GLiNER models
arthrod/gliner-opf-ptbr-pii-bench-v1 dataset Benchmark of the current suite
arthrod/gliner_review_comparison dataset Side-by-side review of GLiNER outputs
arthrod/gliner-opf-ptbr-pii-demo space Interactive demo of the current suite
arthrod/gliner_eval_folder dataset Feb-2026 sweep: GLiNER models x 9 eval sets (source of the only task metrics for these legacy repos)
arthrod/attempt_vanilla model Experiment log (24 Feb 2026) of the runs behind v1-deberta_large-… and v1-deberta_small-…
arthrod/c3750-mdeberta_base-vanilla-1-2 model Legacy (Aug 2025): mDeBERTa-v3-base span (markerV0) GLiNER, 15 checkpoints
arthrod/c5k-deberta_base-token_level-1-2 model Legacy (Feb 2026): urchade/gliner_multi_pii-v1 continued on a PII-in-spam corpus, 13 checkpoints
arthrod/v0-ettin68m-token_level-0-3 model Legacy (Feb 2026): LoRA fine-tune of knowledgator/gliner-pii-small-v1.0 (Ettin-68M), 25 checkpoints
arthrod/v1-deberta_large-token_level-1-7 model Legacy (Feb 2026): knowledgator/gliner-pii-large-v1.0 fine-tuned 8k steps, has eval metrics
arthrod/v1-deberta_small-token_level-0-6 model Legacy (Feb 2026): failed from-scratch DeBERTa-v3-small run (loss stuck at 0)

Quick start

Download only the checkpoint you need. A bare GLiNER.from_pretrained(repo) reads the wrong root config and may pull the whole ~48 GB snapshot. GLiNER 0.2.25 has no subfolder argument:

from huggingface_hub import snapshot_download
from gliner import GLiNER

repo = "arthrod/c5k-deberta_base-token_level-1-2"
path = snapshot_download(
    repo,
    allow_patterns=["checkpoint-6500/*"],
    ignore_patterns=["*optimizer*", "*rng_state*", "*scheduler*"],
)
model = GLiNER.from_pretrained(f"{path}/checkpoint-6500")  # span-level mDeBERTa config, same weights as root

text = "Ganhe R$ 5.000 agora! Ligue para (11) 98765-4321 ou escreva para promo@exemplo.com.br"
labels = ["phone number", "email address", "person"]
for ent in model.predict_entities(text, labels, threshold=0.5):
    print(ent["text"], "->", ent["label"], round(ent["score"], 3))

This snippet was not run against these weights for this card. Training used transformers 5.1.0 (checkpoint-*/gliner_config.json).

Files and checkpoint layout

Path What it is
pytorch_model.bin Root weights, the same LFS sha256 as checkpoint-6500/pytorch_model.bin (1.16 GB).
gliner_config.json (root) Wrong for these weights. Hand-written token-level / DeBERTa-v3-base config (model_type: gliner_uni_encoder_token). Kept unchanged. See the warning above.
tokenizer.json, tokenizer_config.json mDeBERTa SentencePiece tokenizer (DebertaV2Tokenizer, extra_id_* tokens).
onnx/model.onnx ONNX export (opset 19, added 2026-03-19, 1.15 GB). Which config it was exported with is not recorded. Verify its outputs before use.
config.yaml Resolved training YAML (run, model, data, training, LoRA [disabled], environment).
summary_20260224T084509Z.txt, validation_20260224T084509Z.log Config-validation summary and the start-up log of the run: data sizes, model class, parameter count, the full trainer kwargs.
training_args.bin Pickled gliner.training.trainer.TrainingArguments (contains no token: hub_token: None).
checkpoint-<step>/ (13 dirs, every 500 steps) Full Trainer checkpoints: weights (1.16 GB), optimizer.pt (2.31 GB), RNG/scheduler state, trainer_state.json, correct GLiNER config, tokenizer. About 3.48 GB each. All 13 gliner_config.json files are identical.
training_log_history.parquet Every log entry of checkpoint-6500/trainer_state.json as a flat table (added by this card update).

Checkpoints

Eval loss is from checkpoint-6500/trainer_state.json; eval ran on all 526,449 eval examples (about 15–16 min each). "Mean train loss" is the mean of the logged train loss (every 10 steps) over the 500 steps before the checkpoint. Focal loss with sum reduction.

Checkpoint Epoch Eval loss Mean train loss (prev. 500 steps)
checkpoint-500 0.014 48.91 54.09
checkpoint-1000 0.029 39.05 17.64
checkpoint-1500 0.043 33.72 14.96
checkpoint-2000 0.058 31.95 13.52
checkpoint-2500 0.072 30.82 13.43
checkpoint-3000 0.086 29.10 12.58
checkpoint-3500 0.101 30.37 12.76
checkpoint-4000 0.115 27.71 11.61
checkpoint-4500 0.130 27.65 10.55
checkpoint-5000 0.144 27.11 11.45
checkpoint-5500 0.158 26.94 10.70
checkpoint-6000 0.173 25.97 10.84
checkpoint-6500 (= root) 0.187 25.08 11.28

Best checkpoint: checkpoint-6500, which is the lowest eval loss and also the last one. Eval loss was still falling when the run was interrupted. Every intermediate step was also pushed as its own commit ("Training in progress, step N"), so older weights can be recovered from the git history as well as from the checkpoint-* folders.

Training data

  • Train: arthrod/gliner-flex-pii-ready-v5, split train, 1,111,948 examples (from validation_*.log).
  • Eval: same dataset, split eval, 526,449 examples.
  • That dataset is private, so its language mix and label set cannot be checked from here. The run's stated goal is "Domain adaptation for PII-in-spam" (config.yaml, tags pii, spam, gliner, token_level, domain-adapt, multipii-v1).
  • An earlier version of this card listed arthrod/pii-gliner-evals (private) as the intended holdout and nvidia/gliner-PII as the intended baseline. No result was ever filled in.

Training recipe

From config.yaml, validation_20260224T084509Z.log and training_args.bin (inspected without unpickling):

  • Start: urchade/gliner_multi_pii-v1 (UniEncoderSpanGLiNER, mDeBERTa-v3-base, markerV0, max_width 12, hidden_size 512, 1 RNN layer, max_len 384). 288,949,504 parameters, all trainable (no LoRA).
  • Overrides: max_types 100, max_neg_type_ratio 2, dropout 0.25.
  • Optimizer AdamW (torch); LR encoder 8e-6, LR others 4e-5; weight decay 0.01 / 0.01; cosine schedule, 5% warmup; max_steps 12000; batch 32 (eval 64); grad clip 1.0; bf16; seed 42.
  • Loss: focal alpha 0.75, gamma 2.0, sum reduction, negatives ratio 2.0, no masking.
  • Eval/save every 500 steps, save_total_limit 20; W&B project gliner-pii; flash-attention 2 requested.
  • Interrupted at step 6,500/12,000 on 2026-02-24 ("Training interrupted at step 6500/12000" commit).

Evaluation

  • The only recorded metric is eval loss (table above).
  • arthrod/gliner_eval_folder EVAL_REPORT.md (Feb 2026) triaged this repo out of the benchmark: "Noisy predictions … 50 garbage preds, all scores ~0.50 (embedding resize issue)". That triage loaded the repo root, which has the wrong config (see top). No benchmark has been run on checkpoint-6500/ with its correct config.

training_log_history.parquet

Column Type Description
kind string train (logged every 10 steps) or eval
step int64 Global optimizer step
epoch float64 Fractional epoch
loss float64 Train loss; null on eval rows
grad_norm float64 Gradient norm; null on eval rows
learning_rate float64 LR; null on eval rows
eval_loss float64 Eval loss; null on train rows
eval_runtime float64 Eval wall time (s)
eval_samples_per_second float64 Eval throughput

663 rows (650 train + 13 eval). Examples:

[
  {"kind": "train", "step": 6500, "epoch": 0.187, "loss": 14.2946, "grad_norm": 50.52, "learning_rate": 1.890e-05, "eval_loss": null, "eval_runtime": null, "eval_samples_per_second": null},
  {"kind": "eval", "step": 6500, "epoch": 0.187, "loss": null, "grad_norm": null, "learning_rate": null, "eval_loss": 25.083, "eval_runtime": 964.85, "eval_samples_per_second": 545.628}
]

Limitations

  • No F1/precision/recall has ever been computed for a correctly loaded checkpoint.
  • Root gliner_config.json does not match the root weights (see top). The repo name repeats the same two errors (encoder and span mode).
  • Only 0.19 epoch of an interrupted schedule (the LR was still at 1.9e-5 for the "others" group when it stopped).
  • Training data is private and was built for spam-domain PII. Behaviour on legal/medical PT-BR text is unknown.
  • Span-level with max_width 12: entities longer than 12 words cannot be predicted.

License

Apache-2.0, as declared on this repo, matching the parent urchade/gliner_multi_pii-v1 (Apache-2.0). The microsoft/mdeberta-v3-base encoder is MIT.

Citation

@misc{arthrod_c5k_multipii_spam,
  author = {arthrod},
  title  = {c5k-deberta_base-token_level-1-2 (gliner-pii-token-multipii-v1-spam-v1): GLiNER multi-PII continued on PII-in-spam},
  year   = {2026},
  doi    = {10.57967/hf/8514},
  url    = {https://huggingface.co/arthrod/c5k-deberta_base-token_level-1-2}
}

@inproceedings{zaratiana2024gliner,
  title     = {GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer},
  author    = {Zaratiana, Urchade and Tomeh, Nadi and Holat, Pierre and Charnois, Thierry},
  booktitle = {NAACL},
  year      = {2024}
}
Downloads last month
22
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for arthrod/c5k-deberta_base-token_level-1-2

Finetuned
(8)
this model