BaoNhan's picture
Update README.md
8407d0c verified
|
Raw
History Blame Contribute Delete
2.29 kB
---
language: vi
library_name: transformers
pipeline_tag: text-classification
tags:
- clickbait-detection
- vietnamese
- viclickbait-2025
datasets:
- ViClickbait-2025
---
# CafeBERT — ViClickbait-2025
Fine-tuned from `uitnlp/CafeBERT` for binary Vietnamese clickbait detection.
## Experimental setup
- Input: headline paired with lead paragraph; no URL, source, category, publish time, image, or engagement metadata.
- Fixed 80/10/10 split using `StratifiedGroupKFold` with seed 42.
- Fine-tuning seeds: [42, 22, 202]; 3 epochs per seed.
- Development Macro-F1 selects checkpoints and representative seed.
- Weighted cross-entropy from training-label frequencies: `True`.
- Effective batch size: 8; max length: 256.
## Results
| Metric | Mean ± sample std |
|---|---:|
| Test Macro-F1 | 0.8047 ± 0.0060 |
| Test accuracy | 0.8236 ± 0.0074 |
| Dev Macro-F1 | 0.8256 ± 0.0053 |
Representative seed: **22**, selected only by development Macro-F1.
### Per-seed
| seed | dev_macro_f1 | test_macro_f1 | test_accuracy |
|-------:|---------------:|----------------:|----------------:|
| 22 | 0.8312 | 0.8043 | 0.8246 |
| 42 | 0.8249 | 0.7989 | 0.8158 |
| 202 | 0.8207 | 0.8108 | 0.8304 |
## Labels
- `0`: non-clickbait
- `1`: clickbait
## Usage
```python
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model_id = "BaoNhan/cafebert-ViClickbait-2025"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
title = "Tiêu đề bài báo"
lead = "Đoạn dẫn của bài báo"
inputs = tokenizer(title, lead, return_tensors="pt", truncation=True, max_length=256)
prediction = model(**inputs).logits.argmax(dim=-1).item()
print(model.config.id2label[prediction])
```
## Dataset
- Nguyen et al. (2025), *ViClickbait-2025: A comprehensive dataset for Vietnamese clickbait detection*. https://doi.org/10.1016/j.dib.2025.112164
- Dataset: https://doi.org/10.17632/3wc46bfcjc.1
## Limitations
The dataset is small, temporally bounded to 2023–2025, and collected from eight Vietnamese news platforms. Results may not transfer to social media, other publishers, or emerging clickbait styles.