File size: 2,510 Bytes
44ed717
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
---
license: mit
language:
  - en
  - es
library_name: onnxruntime
pipeline_tag: sentence-similarity
tags:
  - embeddings
  - onnx
  - retrieval
  - sts
  - on-device
base_model: intfloat/multilingual-e5-large-instruct
---

# ALF-emb-micro 1.0

Bilingual EN+ES embeddings. ALF is the AtomicoLabs model family. **Micro** = under 1B parameters.

Vocab-pruned finetune of multilingual-e5-large-instruct. Shipping format is per-channel int8 ONNX (mean-pool, L2-normalized, 1024-d).

Weights: GitHub Release [`v1.0`](https://github.com/AtomicoLabs/ALF-emb-micro/releases/tag/v1.0) and this Hugging Face revision `v1.0`.

| | Composite | EN retrieval | ES retrieval | STS | Domain |
|---|---|---|---|---|---|
| ALF-emb-micro 1.0 (int8) | 79.99 | 57.25 | 77.80 | 89.07 | 95.85 |
| OpenAI text-embedding-3-small | 80.08 | 59.74 | 79.52 | 88.66 | 92.39 |

370M params · vocab 64.5k · cosine parity vs fp32 0.984.

This tokenizer is **not** drop-in e5. You must remap ids with `keep_ids.npy` / `remap.py`.

## Use

```python
from pathlib import Path
import sys
import numpy as np
import onnxruntime as ort
from huggingface_hub import snapshot_download

repo = Path(snapshot_download("AtomicoLabs/ALF-emb-micro", revision="v1.0"))
sys.path.insert(0, str(repo))
from remap import RemapTokenizer

QUERY = (
    "Instruct: Given a web search query, retrieve relevant passages that answer the query\n"
    "Query: "
)

tok = RemapTokenizer(repo, np.load(repo / "keep_ids.npy").tolist())
sess = ort.InferenceSession(str(repo / "model.onnx"), providers=["CPUExecutionProvider"])

def embed(texts, *, is_query=False):
    batch = [(QUERY + t if is_query else t) for t in texts]
    enc = tok(batch, padding=True, truncation=True, max_length=512, return_tensors="np")
    vec = sess.run(None, {
        "input_ids": enc["input_ids"].astype(np.int64),
        "attention_mask": enc["attention_mask"].astype(np.int64),
    })[0]
    return vec.astype(np.float32)
```

Documents take an empty prefix. Queries need the instruct prefix above.

## Limits

- English retrieval is a bit under the OpenAI small bar; Spanish and domain are close or ahead.
- Max 512 tokens. int8, not fp16/fp32.
- Not a general instruction model.

## Training

Finetune of [`intfloat/multilingual-e5-large-instruct`](https://huggingface.co/intfloat/multilingual-e5-large-instruct) (MIT), then vocab prune 560M → 370M, then dynamic per-channel int8 ONNX.

## License

MIT. Include this notice and the e5-large-instruct MIT notice when you redistribute.