You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

embeddingmagibu2

embeddingmagibu2 is a Turkish-focused multilingual embedding model built from google/embeddinggemma-2 and the Turkish-optimized magibu/Altus tokenizer.

Status: training in progress. See Training progress below for the latest metrics.

Training progress

Offline distillation from google/embeddinggemma-2 is in progress; this table is updated automatically at every release. Latest release: phase1-test-new-done. Phases: test/new → test/full → validation/new → validation/full → train/new → train/full (new: only the new Turkish token embeddings train; full: all embeddings + text backbone at a low learning rate). Metrics on 2,048 held-out validation rows: cosine to the teacher's document embeddings for Turkish and for the other languages, their mean (score), and STSbTR test Spearman.

Release Step cos_doc_tr cos_doc_other score STSbTR Spearman Time (UTC)
initial (untrained) 0 0.9314 0.9789 0.9551 0.5867 2026-10-10 13:35
phase1-test-new-step1100 1100 0.9369 0.9818 0.9594 0.5796 2026-10-10 14:31
phase1-test-new-done 1102 0.9369 0.9818 0.9594 0.5796 2026-10-10 14:31

How it was built

It follows the tokenizer-surgery + offline-distillation pipeline of Adapting Multilingual Embedding Models to Turkish via Cross-Lingual Tokenizer Surgery and Offline Distillation.

  1. Tokenizer swap. Altus is the Gemma-4 vocabulary (identical to the teacher's) in which 42,484 rarely used tokens were replaced by Turkish subwords from Magibu's 65,536-token BPE vocabulary.
  2. Mean-composition initialization. The 219,660 unchanged token ids keep their teacher embeddings exactly. Each of the 42,484 new tokens is initialized as the mean of the teacher embeddings of the teacher tokens that spell it (using Altus' merged-to-original-token-ids.json), e.g. ağ ← mean(a, ğ). All other weights are the teacher's, unchanged.
  3. Offline distillation (next). The student is trained to match precomputed teacher embeddings on alibayram/wikipedia-40-langs-with-embeddings-embeddinggemma2: 40-language Wikipedia, with retrieval-style targets for documents (title: {title} | text: {text}) and title queries (task: search result | query: {title}). Texts are clipped so they fit in 8,192 tokens under both the teacher and the Altus tokenizer, so student and teacher always see identical inputs.

Model details

Base model google/embeddinggemma-2
Tokenizer magibu/Altus (BPE, 262,144 tokens)
Embedding dimension 768, L2-normalized
Max sequence length 8,192 tokens
Pooling mean
Parameters ~744M total (text model ~271M; vision and audio towers inherited from the base model)

Tokens needed for the same Wikipedia text, Altus vs. the original tokenizer (6,000-article sample):

Language Altus / original
Turkish 0.79× (21% fewer tokens)
Most other languages 1.02–1.08×
Russian, Bulgarian 1.10–1.15×

Fewer Turkish tokens means more Turkish text fits into the 8K context and inference is faster.

Usage

Use the same prompts as EmbeddingGemma: a query prompt for queries and a document prompt for documents.

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("alibayram/embeddingmagibu2")

queries = ["Türkiye'nin başkenti neresidir?"]
documents = [
    "Ankara, Türkiye'nin başkentidir ve İç Anadolu Bölgesi'nde yer alır.",
    "İstanbul, Türkiye'nin en kalabalık şehridir.",
]

query_embeddings = model.encode(queries, prompt_name="query")
document_embeddings = model.encode(documents, prompt_name="document")
print(model.similarity(query_embeddings, document_embeddings))

Limitations

  • This checkpoint has not been distilled yet; embeddings of new Turkish tokens are only initialized.
  • Languages whose text becomes longer under Altus (e.g. Russian, Bulgarian) may lose some quality.
  • The vision and audio towers are inherited but not adapted; use this model for text.

Citation

@misc{bayram2026embeddingmagibu,
  title         = {Adapting Multilingual Embedding Models to Turkish via Cross-Lingual Tokenizer Surgery and Offline Distillation},
  author        = {Bayram, M. Ali and Diri, Banu and Y{\i}ld{\i}r{\i}m, Sava{\c{s}}},
  year          = {2026},
  eprint        = {2605.29992},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL}
}

Developed by Magibu AI Research.

Downloads last month
9
Safetensors
Model size
0.7B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for alibayram/embeddingmagibu2

Finetuned
(36)
this model

Paper for alibayram/embeddingmagibu2