Instructions to use Yaldat/Yalda-Embedding with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use Yaldat/Yalda-Embedding with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("Yaldat/Yalda-Embedding") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Transformers
How to use Yaldat/Yalda-Embedding with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Yaldat/Yalda-Embedding") model = AutoModelForCausalLM.from_pretrained("Yaldat/Yalda-Embedding", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Yalda Embedding
Yalda Embedding is an openly developed Persian text embedding model for dense retrieval and semantic search. It specializes Qwen3-Embedding-0.6B using more than one million Persian query-document pairs and relational knowledge distillation from F2LLM-v2-8B.
With 595.8 million parameters and 1,024-dimensional embeddings, Yalda achieves state-of-the-art performance among the evaluated 0.6B models on eight Persian retrieval tasks from FaMTEB. Its reported mean is 0.5488, compared with 0.4713 for its Qwen3-0.6B initialization and 0.5563 for Qwen3-Embedding-4B. It offers competitive retrieval performance against the 4B reference while retaining the 0.6B model's inference footprint.
The model weights, training data, teacher embeddings, and development notebooks are public:
- Training and benchmark code: yaldat/yalda-embedding
- Training data: Yaldat/Persian-Pairwise
- Teacher-embedded data: Yaldat/Persian-Pairwise-Embedded
Model details
| Property | Value |
|---|---|
| Primary language | Persian / Farsi |
| Base model | Qwen3-Embedding-0.6B |
| Parameters | 595,776,512 |
| Embedding dimension | 1,024 |
| Architecture | Decoder-only Transformer, 28 blocks |
| Attention | Causal, grouped-query attention |
| Pooling | Last non-padding token |
| Training and evaluation input length | 512 tokens |
| Similarity | Cosine similarity; dot product for L2-normalized embeddings |
| Released weights | FP32 Safetensors |
| License | Apache 2.0 |
Quick start
Install Sentence Transformers and its model dependencies:
pip install -U sentence-transformers transformers torch
Use the saved query prompt for search queries and the document prompt for passages. Normalize embeddings so their dot product equals cosine similarity.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("Yaldat/Yalda-Embedding")
model.tokenizer.padding_side = "left"
model.max_seq_length = 512
queries = [
"پایتخت ایران کجاست؟",
"چرا برگ درختان در پاییز تغییر رنگ میدهد؟",
]
documents = [
"تهران پایتخت ایران و یکی از بزرگترین شهرهای این کشور است.",
"در پاییز با کاهش نور خورشید، کلروفیل برگها تجزیه میشود و رنگدانههای دیگر نمایان میشوند.",
"کتابخانهها فضایی برای مطالعه و دسترسی به منابع علمی فراهم میکنند.",
]
query_embeddings = model.encode(
queries,
prompt_name="query",
normalize_embeddings=True,
)
document_embeddings = model.encode(
documents,
prompt_name="document",
normalize_embeddings=True,
)
scores = query_embeddings @ document_embeddings.T
for query, row in zip(queries, scores):
best_index = int(row.argmax())
print(f"Query: {query}")
print(f"Best passage: {documents[best_index]}")
print(f"Similarity: {row[best_index]:.4f}\n")
The model has no default prompt, so select prompt_name="query" explicitly for queries.
The saved query prompt is:
Instruct: Given a web search query, retrieve relevant passages that answer the query
Query:
Documents receive no instruction. Pass raw query text to encode; the query prompt is
added automatically when selected. For larger collections, precompute document embeddings
and index them in a vector search system.
The example uses 512 tokens to match the training and evaluation setup. Longer inputs are truncated under this setting. The inherited backbone's longer context configuration does not establish retrieval quality beyond the evaluated input length.
Evaluation
The results below are reported in Yalda Embedding: An Engineering Effort to Reach State of the Art. Evaluation uses the test splits of eight Persian retrieval tasks from FaMTEB. The summary is the unweighted macro-average of the reported task scores.
Comparison with evaluated baselines
| Model | Mean retrieval score |
|---|---|
| Qwen3-Embedding-4B | 0.5563 |
| Yalda Embedding | 0.5488 |
| F2LLM-v2-0.6B | 0.5063 |
| Dibachain/Diba-Embed | 0.4788 |
| Qwen3-Embedding-0.6B | 0.4713 |
| heydariAI/persian-embeddings | 0.4175 |
| Tooka-SBERT-V2-Large | 0.4150 |
Yalda improves over Qwen3-Embedding-0.6B by 0.0775 absolute (16.4% relative) and over F2LLM-v2-0.6B by 0.0425 absolute. It exceeds Diba-Embed by 0.0700 absolute (14.6% relative). It scores higher than each of these three baselines on all eight tasks at the reported precision.
Against Qwen3-Embedding-4B, Yalda scores higher on NFCorpus-Fa, SynPerQARetrieval, PersianWebDocumentRetrieval, and HotpotQA-FaHardNegatives, and ties on FiQA2018-Fa. Its overall mean is 0.0075 below the 4B reference. The SOTA claim is scoped to the evaluated 0.6B models and these eight retrieval tasks.
Yalda's per-task scores
| Task | Reported score |
|---|---|
| FiQA2018-Fa | 0.30 |
| NFCorpus-Fa | 0.33 |
| SynPerQARetrieval | 0.88 |
| PersianWebDocumentRetrieval | 0.57 |
| HotpotQA-FaHardNegatives | 0.58 |
| MSMARCO-FaHardNegatives | 0.65 |
| NQ-FaHardNegatives | 0.43 |
| SciFact-Fa | 0.65 |
| Macro-average | 0.5488 |
Student training uses query-positive-document pairs, in-batch negatives, and relational distillation. It uses no explicitly mined or supplied hard negatives, while achieving the reported scores on the three hard-negative evaluation tasks.
Training recipe
The model is fine-tuned on 1,035,981 Persian query-positive-document pairs assembled from 13 retrieval resources. Training data comes from the source-provided training splits; evaluation uses their corresponding test splits. The empirical source mixture is shuffled without resampling.
The shared encoder is optimized with query-to-document InfoNCE and forward KL divergence between teacher and student distributions over the documents in each global batch:
loss = (1 - alpha) * InfoNCE + alpha * KL(teacher || student)
alpha = 0.75
The final recipe weights InfoNCE by 0.25 and distillation by 0.75. Teacher vectors are precomputed offline, truncated from 4,096 to their first 1,024 Matryoshka dimensions, and L2-normalized before constructing similarity distributions. The method transfers ranking relationships rather than matching teacher and student coordinates. The 8B teacher is needed for data preparation, not for student training or inference.
| Setting | Value |
|---|---|
| Hardware | Eight TPU v5e devices |
| Training framework | JAX 0.9.2 and Flax NNX 0.12.6 |
| Epochs | 1 |
| Global batch size | 512 |
| Pairs per device | 64 |
| Optimization steps | 2,023 |
| Optimizer | AdamW |
| Peak learning rate | 2e-5 |
| Warmup | First 10% of updates |
| Decay | Cosine to zero |
| Gradient clipping | Global norm 1.0 |
| Student / teacher temperature | 0.05 / 0.07 |
| KD weight | 0.75 |
| Sequence length | 512 tokens |
| Precision | BF16 projections with FP32 parameters, normalization, and losses |
| Memory strategy | Single-axis FSDP and activation rematerialization |
One epoch consumes 1,035,776 pairs; the incomplete final batch of 205 pairs is omitted. The model seed is zero; the final experiment does not fix the data permutation seed.
Distillation ablation
| Configuration | KD weight | Mean retrieval score |
|---|---|---|
| Stabilized InfoNCE, batch size 512 | 0.00 | 0.4813 |
| Ablation 4 | 0.25 | 0.5125 |
| Ablation 5 | 0.50 | 0.5313 |
| Ablation 6: final model | 0.75 | 0.5488 |
The final distilled configuration improves all eight tasks over the stabilized, non-distilled model. Increasing the KD weight from 0.50 to 0.75 adds 0.0175 to the mean.
Open development and reproducibility
The public GitHub repository contains:
- Yalda-Embedding-Training.ipynb: TPU training, configurable distillation, checkpoint export, and Hugging Face upload.
- Benchmark.ipynb: MTEB evaluation on selected Persian retrieval test tasks.
- ReadMe.md: experiment setup and instructions.
Set KD_ALPHA = 0.75 in the training notebook to use the final paper configuration.
Keep the model revision, dependency versions, task selection, prompt settings, and exact
score field with each benchmark run. The notebooks expose the development procedure;
the tables above reproduce the paper's reported results.
Scope and limitations
The reported evaluation covers Persian retrieval. Performance on other languages, classification, clustering, and other embedding tasks has not been established by these experiments. The eight-task comparison does not establish SOTA across all model sizes or benchmarks. Longer-context retrieval quality is not evaluated here.
In-batch negatives can include semantically relevant passages, and distillation can transfer the teacher's retrieval preferences and errors. Separate source splits are used, but broader overlap auditing and repeated seeded runs remain areas for further evaluation.
Acknowledgments
Yalda builds on Qwen3 Embedding, uses F2LLM-v2-8B as its offline teacher, and is evaluated using MTEB and FaMTEB.
- Downloads last month
- 113