kenga-embed-longposF
44M-parameter Russian/English sentence encoder from the Kenga project. Bidirectional Z-factored transformer, SentencePiece-16k tokenizer, mean pooling, 768-d L2-normalised output, 512-token context. Prefix protocol is the FRIDA / BERTA one, so it drops into any pipeline that already uses them.
kenga-embed-longpos is kenga-embed-prophet5 with positions 256..511 actually trained. Every earlier stage (teacher and student alike) ran at 256 tokens, so the second half of the learned position table stayed at its random init and any text longer than 256 tokens was encoded with noise in its tail. Stage A trains only pos[256:512] (everything else frozen, so texts up to 256 tokens are unchanged) against BERTA at 512 tokens on 70k long open texts (Lenta.ru, Russian Wikipedia; neither is an MTEB(rus) corpus); the F stage then briefly unfreezes the whole model on the same objective.
Official MTEB(rus, v1.1) numbers
Run with mteb==2.20.5, all splits and subsets as defined by the benchmark, no task was skipped or re-weighted.
Reference columns are the models' own leaderboard submissions
(embeddings-benchmark/results). 23/23 tasks done.
| task | kenga-embed-longposF | Giga-Embeddings-instruct-480M | BERTA-128M | USER2-small-34M | rubert-tiny-turbo-29M |
|---|---|---|---|---|---|
| GeoreviewClassification | 46.9 | 55.4 | 54.8 | 41.1 | 41.4 |
| HeadlineClassification | 85.0 | 89.0 | 89.0 | 74.3 | 68.9 |
| InappropriatenessClassification | 60.9 | 86.1 | 74.8 | 60.7 | 59.1 |
| KinopoiskClassification | 61.5 | 73.0 | 67.8 | 52.2 | 50.5 |
| MassiveIntentClassification | 63.1 | 85.3 | 74.0 | 66.1 | 58.0 |
| MassiveScenarioClassification | 73.4 | 90.9 | 84.5 | 70.3 | 62.9 |
| RuReviewsClassification | 70.2 | 76.3 | 72.3 | 60.8 | 60.7 |
| RuSciBenchGRNTIClassification | 63.7 | 74.0 | 69.0 | 63.1 | 52.9 |
| RuSciBenchOECDClassification | 49.5 | 59.9 | 54.8 | 49.2 | 40.8 |
| CEDRClassification | 54.9 | 69.8 | 73.0 | 39.4 | 39.0 |
| SensitiveTopicsClassification | 29.0 | 44.3 | 39.9 | 27.5 | 25.2 |
| GeoreviewClusteringP2P | 47.6 | 73.8 | 73.8 | 66.2 | 59.7 |
| RuSciBenchGRNTIClusteringP2P | 59.8 | 70.5 | 65.0 | 56.4 | 48.1 |
| RuSciBenchOECDClusteringP2P | 51.5 | 58.1 | 55.6 | 48.6 | 41.1 |
| TERRa | 61.8 | 79.6 | 65.7 | 54.0 | 56.3 |
| RuBQReranking | 66.4 | 80.5 | 75.2 | 66.0 | 62.2 |
| MIRACLReranking | 49.6 | 67.5 | 64.3 | 50.5 | 47.7 |
| RiaNewsRetrievalHardNegatives.v2 | 52.7 | 88.9 | 84.5 | 74.5 | 52.3 |
| RuBQRetrieval | 54.1 | 80.6 | 71.0 | 61.1 | 51.7 |
| MIRACLRetrievalHardNegatives.v2 | 45.5 | 74.7 | 65.9 | 46.1 | 42.4 |
| RUParaPhraserSTS | 67.3 | 78.3 | 77.8 | 69.6 | 72.1 |
| RuSTSBenchmarkSTS | 71.2 | 83.6 | 82.2 | 81.0 | 78.5 |
| STS22 | 54.9 | 65.3 | 61.1 | 66.1 | 64.6 |
| --- | --- | --- | --- | --- | --- |
| Classification (mean) | 63.8 | 76.7 | 71.2 | 59.8 | 55.0 |
| MultilabelClassification (mean) | 41.9 | 57.1 | 56.5 | 33.5 | 32.1 |
| Clustering (mean) | 52.9 | 67.5 | 64.8 | 57.1 | 49.6 |
| PairClassification (mean) | 61.8 | 79.6 | 65.7 | 54.0 | 56.3 |
| Reranking (mean) | 58.0 | 74.0 | 69.7 | 58.3 | 54.9 |
| Retrieval (mean) | 50.8 | 81.4 | 73.8 | 60.6 | 48.8 |
| STS (mean) | 64.5 | 75.7 | 73.7 | 72.2 | 71.7 |
| mean over tasks | 58.3 | 74.2 | 69.4 | 58.5 | 53.7 |
| mean over task types (leaderboard) | 56.2 | 73.1 | 67.9 | 56.5 | 52.6 |
| tasks done | 23 | 23 | 23 | 23 | 23 |
Leaderboard-style mean (average of task-type means): 56.2.
What this is and is not: a 44M model, roughly 10x smaller than Giga-Embeddings-instruct-480M, meant to be compared
with the 30-40M Russian encoders (USER2-small, rubert-tiny-turbo). On the 23 finished tasks (plain mean) it scores 58.3 vs 74.2 for Giga-Embeddings-instruct-480M (+15.9 gap) and 58.5 for USER2-small-34M (-0.2). It does not beat Giga.
The numbers above are the whole story; the raw result files are in mteb_results/ of the training tree.
Usage
import sys; sys.path.insert(0, "<this folder>") # or trust_remote_code-style import after download
from modeling_kenga_embed_v2 import KengaEmbedV2HF
m = KengaEmbedV2HF.from_pretrained("<this folder>", device="cuda") # cpu works too
q = m.encode(["??? ??????? ????? ? ????"], prefix="search_query")
d = m.encode(["????? ???? ???????? ...", "?????? ?????"], prefix="search_document")
print(q @ d.T) # cosine, embeddings are L2-normalised
a = m.encode(["??? ???? ?? ??????."], prefix="paraphrase")
b = m.encode(["?? ?????? ???? ?????."], prefix="paraphrase")
Prefixes ("<prefix>: <text>" is prepended for you):
| use | prefix |
|---|---|
| retrieval query | search_query |
| retrieval document | search_document |
| STS / paraphrase (both sides) | paraphrase |
| classification / clustering | categorize, categorize_sentiment, categorize_topic |
| NLI / entailment (TERRa) | categorize_entailment |
Download with huggingface_hub.snapshot_download("GermannM/kenga-embed-longposF"); the folder contains pytorch_model.bin,
config.json, kenga_spm.model and the self-contained modeling_kenga_embed_v2.py (torch + sentencepiece only).
Architecture
d=768, layers=8, heads=12, dff=3072, factorised token embedding (16385 x 128 -> 768), Z-factored attention/FF projections with rank 192/512, learned positions up to 512. 44.2M parameters, fp32 checkpoint 177 MB. Checkpoint step 1500.
Training code
PyTorch trainers live in the z-system lab tree (embed_v2/: build_segments.py, teacher.py, distill.py,
build_ft_data.py, mine_hard.py, finetune_prophet.py, run_mteb.py), not in the public kenga-lang repo;
the recipe and the Prophet contract are documented in
docs/PROPHETS.md. Trained on one GTX 1660 (6 GB).
License
MIT.
- Downloads last month
- 17