mDeBERTa-ID-20k

An aggressively vocabulary-pruned version of microsoft/mdeberta-v3-base with a 20k-token Indonesian vocabulary, designed for downstream Indonesian NLP tasks and resource-constrained applications.

This model was developed using VocabPrune, a deterministic, language-aware, frequency-based vocabulary pruning method designed to reduce vocabulary-related model overhead.

The model is a base checkpoint and should be fine-tuned for a specific downstream task.

Model Details

Property Value
Base model microsoft/mdeberta-v3-base
Vocabulary size 20k tokens
Vocabulary Indonesian
Language focus Indonesian
Architecture mDeBERTa-v3-base

This checkpoint uses a more aggressive vocabulary reduction than the 30k models in the VocabPrune collection.

Resources

For the methodology, experimental setup, and detailed evaluation results, please refer to the published paper.

Citation

If you use this model or the VocabPrune methodology in your research, please cite:

@article{fuadi2026efficient,
  author  = {Fuadi, Mukhlish and Wibawa, Adhi Dharma and Sumpeno, Surya},
  title   = {Efficient Transformer Models via Language-Aware
             Frequency-Based Vocabulary Pruning},
  journal = {IEEE Access},
  volume  = {14},
  pages   = {50993--51006},
  year    = {2026},
  doi     = {10.1109/ACCESS.2026.3679735}
}
Downloads last month
21
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for muchad/mdeberta-id-20k

Finetuned
(308)
this model

Collection including muchad/mdeberta-id-20k