QuechuaBERT (PRPE)

Small BertForMaskedLM for Southern Quechua, tokenized with the morphological PRPE segmenter from QuechuaTok.

Training note

CPU run: ~20k Llamacha lines, 5000 MLM steps, tiny BERT (hidden=256, layers=4). Final train loss ≈ 6.19. Early checkpoint for visibility, not a finished SOTA model.

Paper

Contreras, M. (2026). QuechuaTok. arXiv:2606.23943 — https://arxiv.org/abs/2606.23943

Usage

Install QuechuaTok, then load the model weights from this repo. Custom PRPE tokenizer lives in quechuatok.hf_tokenizer.

See QuechuaTok issue #8 and scripts/train_quechuabert.py.

Downloads last month
-
Safetensors
Model size
4.25M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train makitty/quechuabert

Paper for makitty/quechuabert