Instructions to use alibayram/embeddingmagibu2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use alibayram/embeddingmagibu2 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("alibayram/embeddingmagibu2") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
embeddingmagibu2
embeddingmagibu2 is a Turkish-focused multilingual embedding model built from google/embeddinggemma-2 and the Turkish-optimized magibu/Altus tokenizer.
Status: training in progress. See Training progress below for the latest metrics.
Training progress
Offline distillation from google/embeddinggemma-2 is in progress; this table is updated automatically at every release.
Latest release: phase1-test-new-done. Phases: test/new → test/full → validation/new → validation/full → train/new → train/full
(new: only the new Turkish token embeddings train; full: all embeddings + text backbone at a low learning rate).
Metrics on 2,048 held-out validation rows: cosine to the teacher's document embeddings for Turkish and for the other
languages, their mean (score), and STSbTR test Spearman.
| Release | Step | cos_doc_tr | cos_doc_other | score | STSbTR Spearman | Time (UTC) |
|---|---|---|---|---|---|---|
| initial (untrained) | 0 | 0.9314 | 0.9789 | 0.9551 | 0.5867 | 2026-10-10 13:35 |
| phase1-test-new-step1100 | 1100 | 0.9369 | 0.9818 | 0.9594 | 0.5796 | 2026-10-10 14:31 |
| phase1-test-new-done | 1102 | 0.9369 | 0.9818 | 0.9594 | 0.5796 | 2026-10-10 14:31 |
How it was built
It follows the tokenizer-surgery + offline-distillation pipeline of Adapting Multilingual Embedding Models to Turkish via Cross-Lingual Tokenizer Surgery and Offline Distillation.
- Tokenizer swap. Altus is the Gemma-4 vocabulary (identical to the teacher's) in which 42,484 rarely used tokens were replaced by Turkish subwords from Magibu's 65,536-token BPE vocabulary.
- Mean-composition initialization. The 219,660 unchanged token ids keep their teacher
embeddings exactly. Each of the 42,484 new tokens is initialized as the mean of the teacher
embeddings of the teacher tokens that spell it (using Altus'
merged-to-original-token-ids.json), e.g.ağ← mean(a,ğ). All other weights are the teacher's, unchanged. - Offline distillation (next). The student is trained to match precomputed teacher
embeddings on alibayram/wikipedia-40-langs-with-embeddings-embeddinggemma2:
40-language Wikipedia, with retrieval-style targets for documents
(
title: {title} | text: {text}) and title queries (task: search result | query: {title}). Texts are clipped so they fit in 8,192 tokens under both the teacher and the Altus tokenizer, so student and teacher always see identical inputs.
Model details
| Base model | google/embeddinggemma-2 |
| Tokenizer | magibu/Altus (BPE, 262,144 tokens) |
| Embedding dimension | 768, L2-normalized |
| Max sequence length | 8,192 tokens |
| Pooling | mean |
| Parameters | ~744M total (text model ~271M; vision and audio towers inherited from the base model) |
Tokens needed for the same Wikipedia text, Altus vs. the original tokenizer (6,000-article sample):
| Language | Altus / original |
|---|---|
| Turkish | 0.79× (21% fewer tokens) |
| Most other languages | 1.02–1.08× |
| Russian, Bulgarian | 1.10–1.15× |
Fewer Turkish tokens means more Turkish text fits into the 8K context and inference is faster.
Usage
Use the same prompts as EmbeddingGemma: a query prompt for queries and a document prompt for documents.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("alibayram/embeddingmagibu2")
queries = ["Türkiye'nin başkenti neresidir?"]
documents = [
"Ankara, Türkiye'nin başkentidir ve İç Anadolu Bölgesi'nde yer alır.",
"İstanbul, Türkiye'nin en kalabalık şehridir.",
]
query_embeddings = model.encode(queries, prompt_name="query")
document_embeddings = model.encode(documents, prompt_name="document")
print(model.similarity(query_embeddings, document_embeddings))
Limitations
- This checkpoint has not been distilled yet; embeddings of new Turkish tokens are only initialized.
- Languages whose text becomes longer under Altus (e.g. Russian, Bulgarian) may lose some quality.
- The vision and audio towers are inherited but not adapted; use this model for text.
Citation
@misc{bayram2026embeddingmagibu,
title = {Adapting Multilingual Embedding Models to Turkish via Cross-Lingual Tokenizer Surgery and Offline Distillation},
author = {Bayram, M. Ali and Diri, Banu and Y{\i}ld{\i}r{\i}m, Sava{\c{s}}},
year = {2026},
eprint = {2605.29992},
archivePrefix = {arXiv},
primaryClass = {cs.CL}
}
Developed by Magibu AI Research.
- Downloads last month
- 9
Model tree for alibayram/embeddingmagibu2
Base model
google/embeddinggemma-2