Buckets:
446 MB
9 files
Updated 3 months ago
Ctrl+K
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| .gitattributes | 690 Bytes xet | b1b0211d | |
| README.md | 990 Bytes xet | b1d5a1d2 | |
| config.json | 520 Bytes xet | b673a984 | |
| dict.txt | 434 kB xet | 99cedf94 | |
| pytorch_model.bin | 443 MB xet | ee80bc5f | |
| sentencepiece.bpe.model | 800 kB xet | 0f337c05 | |
| special_tokens_map.json | 298 Bytes xet | 1bb7e917 | |
| tokenizer.json | 1.65 MB xet | a9d95157 | |
| tokenizer_config.json | 506 Bytes xet | 1ca40411 |
Usage
Load in transformers library with:
from transformers import AutoTokenizer, AutoModelForMaskedLM
tokenizer = AutoTokenizer.from_pretrained("EMBEDDIA/sloberta")
model = AutoModelForMaskedLM.from_pretrained("EMBEDDIA/sloberta")
SloBERTa
SloBERTa model is a monolingual Slovene BERT-like model. It is closely related to French Camembert model https://camembert-model.fr/. The corpora used for training the model have 3.47 billion tokens in total. The subword vocabulary contains 32,000 tokens. The scripts and programs used for data preparation and training the model are available on https://github.com/clarinsi/Slovene-BERT-Tool
SloBERTa was trained for 200,000 iterations or about 98 epochs.
Corpora
The following corpora were used for training the model:
- Gigafida 2.0
- Kas 1.0
- Janes 1.0 (only Janes-news, Janes-forum, Janes-blog, Janes-wiki subcorpora)
- Slovenian parliamentary corpus siParl 2.0
- slWaC
- Total size
- 446 MB
- Files
- 9
- Last updated
- Jun 5
- Pre-warmed CDN
- US EU US EU