File size: 5,655 Bytes
2d70d4e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
463a0fd
2d70d4e
 
 
 
463a0fd
 
 
 
 
 
 
2d70d4e
463a0fd
 
2d70d4e
 
463a0fd
 
 
 
 
 
 
 
 
 
 
 
 
2d70d4e
463a0fd
 
 
 
 
 
2d70d4e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
---
license: cc-by-sa-4.0
base_model: intfloat/multilingual-e5-small
library_name: quanfire-multilingual-embedding
pipeline_tag: sentence-similarity
tags:
  - sentence-embeddings
  - multilingual
  - indic
  - cross-lingual-retrieval
  - lora
  - e5
language:
  - hi
  - bn
  - gu
  - kn
  - ml
  - mr
  - sa
  - ta
  - te
  - ur
  - en
  - fr
  - de
  - es
  - it
  - pt
  - ru
  - ar
  - tr
  - zh
  - ja
  - ko
  - th
  - vi
  - id
---

# QuanFire Multilingual Embedding — `prod-a70s30-fr`

A production multilingual sentence-embedding adapter, Indic-first and trained
**only on openly-licensed, commercially-clean data**. It is a LoRA adaptation
over a frozen [`intfloat/multilingual-e5-small`](https://huggingface.co/intfloat/multilingual-e5-small)
(MIT) base — a 3.4 MB adapter, 384-dimensional normalized vectors, `max_length` 256.

This is **not** a from-scratch foundation model. The contribution is the framework,
the Indic strength, and training data whose licence you can actually ship on.

- **Framework & code:** [github.com/Quanfire-AI/quanfire-multilingual-embedding](https://github.com/Quanfire-AI/quanfire-multilingual-embedding) (Apache-2.0)
- **PyPI:** `pip install quanfire-multilingual-embedding`
- **Weights licence:** CC BY-SA 4.0 (see *Licence & provenance* below)

## What it covers

- **10 Indic languages** — Hindi, Bengali, Gujarati, Kannada, Malayalam, Marathi,
  Sanskrit, Tamil, Telugu, Urdu.
- **15 global languages** stay competitive — English, French, German, Spanish,
  Italian, Portuguese, Russian, Arabic, Turkish, Chinese, Japanese, Korean, Thai,
  Vietnamese, Indonesian.

## Results (held-out, scored on CUDA)

| Instrument | base e5-small | e5 v2 | **prod-a70s30-fr** |
|---|---|---|---|
| Global FLORES-200 cross-lingual, all-pairs recall | 0.9268 | 0.9488 | **0.9762** |
| French retrieval | 0.961 | 0.977 | **0.990** |
| Indic in-domain, non-Hindi X↔Y recall@1 | 0.7875 | 0.8964 | **0.8994** |
| Hindi-pivot mixed-pool recall@10 | 0.7495 | 0.8852 | **0.8914** |
| FLORES non-Hindi recall@1 | 0.9847 | 0.9609 | **0.9785** |

Global all-pairs beats both the base model and e5 v2; French is recovered with no
language regressed against the base. Indic instruments beat v2 across the board and
stay neutral within sampling noise versus the prior internal Indic model.

## Usage

The adapter runs through the QuanFire framework (it applies the LoRA over the base
and produces normalized embeddings). Install the package and pull the weights:

```bash
pip install 'quanfire-multilingual-embedding[neural]'

# download this model's files into a local directory
hf download quanfire-ai/multilingual-embedding --local-dir multilingual-embedding
```

**As an HTTP embeddings service (recommended for applications).** This exposes an
OpenAI-compatible `POST /v1/embeddings` endpoint, so your app stores the vectors in
its own database or vector index:

```bash
qfme serve --adapter multilingual-embedding --port 8000
```

```bash
curl -s localhost:8000/v1/embeddings \
  -H 'content-type: application/json' \
  -d '{"input": ["नमस्ते दुनिया", "hello world", "bonjour le monde"]}'
# -> {"object":"list","data":[{"index":0,"embedding":[...384 floats...]}, ...],
#     "model":"multilingual-embedding","usage":{...},"prefix_applied":null}
```

This model is symmetric (empty prefixes), so `input_type` is not required; pass
`"input_type": "query"` or `"passage"` only for asymmetric models.

**In-process, as a search pipeline:**

```python
from multilingual_embedding.pipelines.search import SemanticSearchPipeline

pipe = SemanticSearchPipeline.from_adapter("multilingual-embedding")
pipe.index(["नमस्ते दुनिया", "hello world", "bonjour le monde", "Bonjour tout le monde"])
for hit in pipe.search("a french greeting", top_k=3):
    print(hit.rank, round(hit.score, 3), hit.text)
```

Vectors are L2-normalized `float32` (dimension 384), so cosine similarity is a dot
product and they drop straight into any vector database or ANN index.

## Licence & provenance

**Weights: CC BY-SA 4.0.** Use them commercially and redistribute them freely,
provided you keep attribution and license derivative weights under the same
share-alike terms. The share-alike floor comes from the training data, not
preference — every source is openly licensed and documented:

| Source | Role in the blend | Licence |
|---|---|---|
| Wikipedia langlink-mined pairs | article side (~70%) | CC BY-SA 4.0 |
| BPCC-Mined bitext (10 languages) | sentence side (~30%) | CC0 |
| itihasa (Sanskrit) | sentence side | Apache-2.0 |
| Tatoeba (en↔fr) | French-recovery fold | CC BY |
| `intfloat/multilingual-e5-small` | base checkpoint | MIT |

CC BY-SA is the strongest obligation in the mix and so sets the weights licence;
CC0, Apache-2.0, CC BY and MIT are all compatible and add only attribution. The net
effect: the weights are **commercially usable and redistributable** — you can ship
them in a paid product and also release them.

The framework source code is Apache-2.0 (separate from these weights).

## Limitations

- A LoRA adapter over a published checkpoint — not an independently pretrained model.
- Cross-lingual retrieval is only as strong as the training corpus was parallel; on
  out-of-domain FLORES non-Hindi the base model can edge it, an expected effect of
  in-domain specialization.
- Exact (brute-force cosine) search is the intended regime up to ~10⁵–10⁶ vectors;
  beyond that, add your own ANN index.

## Citation

```
QuanFire Multilingual Embedding (prod-a70s30-fr).
QuanFire, 2026. https://github.com/Quanfire-AI/quanfire-multilingual-embedding
```