--- language: - tr license: mit tags: - turkish - roberta - nlp - masked-language-modeling - turkish-nlp thumbnail: latensis_logo.png --- # Latensis — Turkish RoBERTa Base

> *Understanding Beyond the Visible* **Latensis** is a Turkish-specific RoBERTa base model trained from scratch on 1GB of curated Turkish text using the [Hecemen Unigram 128k tokenizer](https://huggingface.co/mursideaki/hecemen-tokenizer-unigram-128k) — a tokenizer designed specifically for Turkish morphology. ## Why "Latensis"? The name is derived from the Latin root *latent* (hidden, concealed) — a direct reference to the latent representations that embedding models learn. The suffix *-ensis* (of, belonging to) completes the name: **Latensis** roughly means *"of the latent space"* or *"that which belongs to what is hidden."* ## Model Details | Property | Value | |----------|-------| | Architecture | RoBERTa base | | Parameters | 184.5M | | Vocabulary size | 128,000 | | Tokenizer | Hecemen Unigram 128k | | Training data | ~1GB Turkish text | | Training steps | 500,000 | | Val loss | 3.21 | | Perplexity | 25 | ## Training Data The model was trained on a carefully curated ~1GB Turkish corpus covering: - News articles - Books and literature - Web content - Legal and administrative texts - Agricultural and domain-specific texts ## Benchmark Results All fine-tuning experiments were conducted starting from this base model. Comparisons are against BERTurk (`dbmdz/bert-base-turkish-cased`). ### Sentiment Analysis | Dataset | Our Model | BERTurk | Training Data | |---------|-----------|---------|---------------| | TRSAv1 (Accuracy) | **0.9027** | 0.8372 | 19k examples | | TRSAv1 (F1-macro) | **0.9026** | 0.8350 | 19k examples | | Winvoker (Accuracy) | **0.8534** | 0.6944 | 19k examples | ### Natural Language Inference → RTE | Dataset | Our Model | BERTurk | Training Data | |---------|-----------|---------|---------------| | RTE (Accuracy) | **0.7960** | 0.7780 | 50k NLI examples | | RTE (F1-macro) | **0.7959** | 0.7770 | 50k NLI examples | > BERTurk NLI model was trained on 482k NLI examples. Our model achieves > better results with only 50k examples (~10x less data). ### Named Entity Recognition | Dataset | Our Model | BERTurk | Training Data | |---------|-----------|---------|---------------| | WikiANN-TR (F1-macro) | 0.9429 | 0.9522 | 20k examples | > Within BERTurk's reported variance range (93.92 ± 0.07). ### Semantic Textual Similarity | Metric | Our Model | Emrecan | Training Data | |--------|-----------|---------|---------------| | Pearson | 0.7687 | 0.8340 | 50k NLI + TrGLUE STS | | Spearman | 0.7976 | 0.8300 | 50k NLI + TrGLUE STS | > Emrecan model uses 482k NLI + full STS-b-TR. Our model uses 50k NLI (~10x less). ### QNLI | Dataset | Our Model | Training Data | |---------|-----------|---------------| | TrGLUE QNLI (Accuracy) | 0.8749 | 120k examples | ## Key Findings ✅ Outperforms BERTurk on Sentiment (TRSAv1 +6pp, Winvoker +16pp) ✅ Outperforms BERTurk on RTE with 10x less NLI training data ✅ Competitive NER performance within BERTurk's variance range ✅ Competitive STS with 10x less training data than Emrecan ✅ All results achieved with only 1GB training data ## Usage ### As a base model for fine-tuning ```python from transformers import RobertaModel, AutoTokenizer import sentencepiece as spm from huggingface_hub import hf_hub_download # Load tokenizer model_path = hf_hub_download( repo_id="mursideaki/hecemen-tokenizer-unigram-128k", filename="tr_unigram_tokenizer.model" ) sp = spm.SentencePieceProcessor() sp.load(model_path) # Load model model = RobertaModel.from_pretrained("mursideaki/latensis-roberta-base-tr") ``` ### Masked Language Modeling ```python from transformers import pipeline mlm = pipeline( "fill-mask", model="mursideaki/latensis-roberta-base-tr" ) # Note: Use with Hecemen tokenizer for best results ``` ## Companion Models & Tokenizers - [hecemen-tokenizer-unigram-128k](https://huggingface.co/mursideaki/hecemen-tokenizer-unigram-128k) — Tokenizer used for training - [hecemen-tokenizer-bpe-96k](https://huggingface.co/mursideaki/hecemen-tokenizer-bpe-96k) — BPE variant - Fine-tuned sentiment model *(coming soon)* ## Citation ```bibtex @misc{latensis2026, author = {Mürşide Aki}, title = {Latensis: A Turkish RoBERTa Base Model}, year = {2026}, publisher = {Hugging Face}, url = {https://huggingface.co/mursideaki/latensis-roberta-base-tr} } ``` ## License MIT