Yalda Embedding

Yalda Embedding is an openly developed Persian text embedding model for dense retrieval and semantic search. It specializes Qwen3-Embedding-0.6B using more than one million Persian query-document pairs and relational knowledge distillation from F2LLM-v2-8B.

With 595.8 million parameters and 1,024-dimensional embeddings, Yalda achieves state-of-the-art performance among the evaluated 0.6B models on eight Persian retrieval tasks from FaMTEB. Its reported mean is 0.5488, compared with 0.4713 for its Qwen3-0.6B initialization and 0.5563 for Qwen3-Embedding-4B. It offers competitive retrieval performance against the 4B reference while retaining the 0.6B model's inference footprint.

The model weights, training data, teacher embeddings, and development notebooks are public:

Model details

Property Value
Primary language Persian / Farsi
Base model Qwen3-Embedding-0.6B
Parameters 595,776,512
Embedding dimension 1,024
Architecture Decoder-only Transformer, 28 blocks
Attention Causal, grouped-query attention
Pooling Last non-padding token
Training and evaluation input length 512 tokens
Similarity Cosine similarity; dot product for L2-normalized embeddings
Released weights FP32 Safetensors
License Apache 2.0

Quick start

Install Sentence Transformers and its model dependencies:

pip install -U sentence-transformers transformers torch

Use the saved query prompt for search queries and the document prompt for passages. Normalize embeddings so their dot product equals cosine similarity.

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("Yaldat/Yalda-Embedding")
model.tokenizer.padding_side = "left"
model.max_seq_length = 512

queries = [
    "پایتخت ایران کجاست؟",
    "چرا برگ درختان در پاییز تغییر رنگ می‌دهد؟",
]
documents = [
    "تهران پایتخت ایران و یکی از بزرگ‌ترین شهرهای این کشور است.",
    "در پاییز با کاهش نور خورشید، کلروفیل برگ‌ها تجزیه می‌شود و رنگدانه‌های دیگر نمایان می‌شوند.",
    "کتابخانه‌ها فضایی برای مطالعه و دسترسی به منابع علمی فراهم می‌کنند.",
]

query_embeddings = model.encode(
    queries,
    prompt_name="query",
    normalize_embeddings=True,
)
document_embeddings = model.encode(
    documents,
    prompt_name="document",
    normalize_embeddings=True,
)

scores = query_embeddings @ document_embeddings.T
for query, row in zip(queries, scores):
    best_index = int(row.argmax())
    print(f"Query: {query}")
    print(f"Best passage: {documents[best_index]}")
    print(f"Similarity: {row[best_index]:.4f}\n")

The model has no default prompt, so select prompt_name="query" explicitly for queries. The saved query prompt is:

Instruct: Given a web search query, retrieve relevant passages that answer the query
Query:

Documents receive no instruction. Pass raw query text to encode; the query prompt is added automatically when selected. For larger collections, precompute document embeddings and index them in a vector search system.

The example uses 512 tokens to match the training and evaluation setup. Longer inputs are truncated under this setting. The inherited backbone's longer context configuration does not establish retrieval quality beyond the evaluated input length.

Evaluation

The results below are reported in Yalda Embedding: An Engineering Effort to Reach State of the Art. Evaluation uses the test splits of eight Persian retrieval tasks from FaMTEB. The summary is the unweighted macro-average of the reported task scores.

Comparison with evaluated baselines

Model Mean retrieval score
Qwen3-Embedding-4B 0.5563
Yalda Embedding 0.5488
F2LLM-v2-0.6B 0.5063
Dibachain/Diba-Embed 0.4788
Qwen3-Embedding-0.6B 0.4713
heydariAI/persian-embeddings 0.4175
Tooka-SBERT-V2-Large 0.4150

Yalda improves over Qwen3-Embedding-0.6B by 0.0775 absolute (16.4% relative) and over F2LLM-v2-0.6B by 0.0425 absolute. It exceeds Diba-Embed by 0.0700 absolute (14.6% relative). It scores higher than each of these three baselines on all eight tasks at the reported precision.

Against Qwen3-Embedding-4B, Yalda scores higher on NFCorpus-Fa, SynPerQARetrieval, PersianWebDocumentRetrieval, and HotpotQA-FaHardNegatives, and ties on FiQA2018-Fa. Its overall mean is 0.0075 below the 4B reference. The SOTA claim is scoped to the evaluated 0.6B models and these eight retrieval tasks.

Yalda's per-task scores

Task Reported score
FiQA2018-Fa 0.30
NFCorpus-Fa 0.33
SynPerQARetrieval 0.88
PersianWebDocumentRetrieval 0.57
HotpotQA-FaHardNegatives 0.58
MSMARCO-FaHardNegatives 0.65
NQ-FaHardNegatives 0.43
SciFact-Fa 0.65
Macro-average 0.5488

Student training uses query-positive-document pairs, in-batch negatives, and relational distillation. It uses no explicitly mined or supplied hard negatives, while achieving the reported scores on the three hard-negative evaluation tasks.

Training recipe

The model is fine-tuned on 1,035,981 Persian query-positive-document pairs assembled from 13 retrieval resources. Training data comes from the source-provided training splits; evaluation uses their corresponding test splits. The empirical source mixture is shuffled without resampling.

The shared encoder is optimized with query-to-document InfoNCE and forward KL divergence between teacher and student distributions over the documents in each global batch:

loss = (1 - alpha) * InfoNCE + alpha * KL(teacher || student)
alpha = 0.75

The final recipe weights InfoNCE by 0.25 and distillation by 0.75. Teacher vectors are precomputed offline, truncated from 4,096 to their first 1,024 Matryoshka dimensions, and L2-normalized before constructing similarity distributions. The method transfers ranking relationships rather than matching teacher and student coordinates. The 8B teacher is needed for data preparation, not for student training or inference.

Setting Value
Hardware Eight TPU v5e devices
Training framework JAX 0.9.2 and Flax NNX 0.12.6
Epochs 1
Global batch size 512
Pairs per device 64
Optimization steps 2,023
Optimizer AdamW
Peak learning rate 2e-5
Warmup First 10% of updates
Decay Cosine to zero
Gradient clipping Global norm 1.0
Student / teacher temperature 0.05 / 0.07
KD weight 0.75
Sequence length 512 tokens
Precision BF16 projections with FP32 parameters, normalization, and losses
Memory strategy Single-axis FSDP and activation rematerialization

One epoch consumes 1,035,776 pairs; the incomplete final batch of 205 pairs is omitted. The model seed is zero; the final experiment does not fix the data permutation seed.

Distillation ablation

Configuration KD weight Mean retrieval score
Stabilized InfoNCE, batch size 512 0.00 0.4813
Ablation 4 0.25 0.5125
Ablation 5 0.50 0.5313
Ablation 6: final model 0.75 0.5488

The final distilled configuration improves all eight tasks over the stabilized, non-distilled model. Increasing the KD weight from 0.50 to 0.75 adds 0.0175 to the mean.

Open development and reproducibility

The public GitHub repository contains:

Set KD_ALPHA = 0.75 in the training notebook to use the final paper configuration. Keep the model revision, dependency versions, task selection, prompt settings, and exact score field with each benchmark run. The notebooks expose the development procedure; the tables above reproduce the paper's reported results.

Scope and limitations

The reported evaluation covers Persian retrieval. Performance on other languages, classification, clustering, and other embedding tasks has not been established by these experiments. The eight-task comparison does not establish SOTA across all model sizes or benchmarks. Longer-context retrieval quality is not evaluated here.

In-batch negatives can include semantically relevant passages, and distillation can transfer the teacher's retrieval preferences and errors. Separate source splits are used, but broader overlap auditing and repeated seeded runs remain areas for further evaluation.

Acknowledgments

Yalda builds on Qwen3 Embedding, uses F2LLM-v2-8B as its offline teacher, and is evaluated using MTEB and FaMTEB.

Downloads last month
113
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Yaldat/Yalda-Embedding

Finetuned
(555)
this model

Datasets used to train Yaldat/Yalda-Embedding