BaoNhan's picture
Update README.md
8407d0c verified
|
Raw
History Blame Contribute Delete
2.29 kB
metadata
language: vi
library_name: transformers
pipeline_tag: text-classification
tags:
  - clickbait-detection
  - vietnamese
  - viclickbait-2025
datasets:
  - ViClickbait-2025

CafeBERT — ViClickbait-2025

Fine-tuned from uitnlp/CafeBERT for binary Vietnamese clickbait detection.

Experimental setup

  • Input: headline paired with lead paragraph; no URL, source, category, publish time, image, or engagement metadata.
  • Fixed 80/10/10 split using StratifiedGroupKFold with seed 42.
  • Fine-tuning seeds: [42, 22, 202]; 3 epochs per seed.
  • Development Macro-F1 selects checkpoints and representative seed.
  • Weighted cross-entropy from training-label frequencies: True.
  • Effective batch size: 8; max length: 256.

Results

Metric Mean ± sample std
Test Macro-F1 0.8047 ± 0.0060
Test accuracy 0.8236 ± 0.0074
Dev Macro-F1 0.8256 ± 0.0053

Representative seed: 22, selected only by development Macro-F1.

Per-seed

seed dev_macro_f1 test_macro_f1 test_accuracy
22 0.8312 0.8043 0.8246
42 0.8249 0.7989 0.8158
202 0.8207 0.8108 0.8304

Labels

  • 0: non-clickbait
  • 1: clickbait

Usage

from transformers import AutoModelForSequenceClassification, AutoTokenizer

model_id = "BaoNhan/cafebert-ViClickbait-2025"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
title = "Tiêu đề bài báo"
lead = "Đoạn dẫn của bài báo"
inputs = tokenizer(title, lead, return_tensors="pt", truncation=True, max_length=256)
prediction = model(**inputs).logits.argmax(dim=-1).item()
print(model.config.id2label[prediction])

Dataset

Limitations

The dataset is small, temporally bounded to 2023–2025, and collected from eight Vietnamese news platforms. Results may not transfer to social media, other publishers, or emerging clickbait styles.