Text Classification
Transformers
Safetensors
Vietnamese
xlm-roberta
clickbait-detection
vietnamese
viclickbait-2025
text-embeddings-inference
Instructions to use BaoNhan/cafebert-ViClickbait-2025 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use BaoNhan/cafebert-ViClickbait-2025 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="BaoNhan/cafebert-ViClickbait-2025")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("BaoNhan/cafebert-ViClickbait-2025") model = AutoModelForSequenceClassification.from_pretrained("BaoNhan/cafebert-ViClickbait-2025", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| language: vi | |
| library_name: transformers | |
| pipeline_tag: text-classification | |
| tags: | |
| - clickbait-detection | |
| - vietnamese | |
| - viclickbait-2025 | |
| datasets: | |
| - ViClickbait-2025 | |
| # CafeBERT — ViClickbait-2025 | |
| Fine-tuned from `uitnlp/CafeBERT` for binary Vietnamese clickbait detection. | |
| ## Experimental setup | |
| - Input: headline paired with lead paragraph; no URL, source, category, publish time, image, or engagement metadata. | |
| - Fixed 80/10/10 split using `StratifiedGroupKFold` with seed 42. | |
| - Fine-tuning seeds: [42, 22, 202]; 3 epochs per seed. | |
| - Development Macro-F1 selects checkpoints and representative seed. | |
| - Weighted cross-entropy from training-label frequencies: `True`. | |
| - Effective batch size: 8; max length: 256. | |
| ## Results | |
| | Metric | Mean ± sample std | | |
| |---|---:| | |
| | Test Macro-F1 | 0.8047 ± 0.0060 | | |
| | Test accuracy | 0.8236 ± 0.0074 | | |
| | Dev Macro-F1 | 0.8256 ± 0.0053 | | |
| Representative seed: **22**, selected only by development Macro-F1. | |
| ### Per-seed | |
| | seed | dev_macro_f1 | test_macro_f1 | test_accuracy | | |
| |-------:|---------------:|----------------:|----------------:| | |
| | 22 | 0.8312 | 0.8043 | 0.8246 | | |
| | 42 | 0.8249 | 0.7989 | 0.8158 | | |
| | 202 | 0.8207 | 0.8108 | 0.8304 | | |
| ## Labels | |
| - `0`: non-clickbait | |
| - `1`: clickbait | |
| ## Usage | |
| ```python | |
| from transformers import AutoModelForSequenceClassification, AutoTokenizer | |
| model_id = "BaoNhan/cafebert-ViClickbait-2025" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id) | |
| model = AutoModelForSequenceClassification.from_pretrained(model_id) | |
| title = "Tiêu đề bài báo" | |
| lead = "Đoạn dẫn của bài báo" | |
| inputs = tokenizer(title, lead, return_tensors="pt", truncation=True, max_length=256) | |
| prediction = model(**inputs).logits.argmax(dim=-1).item() | |
| print(model.config.id2label[prediction]) | |
| ``` | |
| ## Dataset | |
| - Nguyen et al. (2025), *ViClickbait-2025: A comprehensive dataset for Vietnamese clickbait detection*. https://doi.org/10.1016/j.dib.2025.112164 | |
| - Dataset: https://doi.org/10.17632/3wc46bfcjc.1 | |
| ## Limitations | |
| The dataset is small, temporally bounded to 2023–2025, and collected from eight Vietnamese news platforms. Results may not transfer to social media, other publishers, or emerging clickbait styles. | |