File size: 1,341 Bytes
dd44243 49b416c 87fe5d9 49b416c 87fe5d9 49b416c 87fe5d9 49b416c 87fe5d9 49b416c 87fe5d9 49b416c 87fe5d9 49b416c 87fe5d9 49b416c 87fe5d9 49b416c 87fe5d9 49b416c 87fe5d9 49b416c 87fe5d9 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 | ---
license: mit
language:
- en
tags:
- ridiculous-models
---
# unity-embed
An embedding model where every input maps to the same vector.
384 parameters, one per dimension, all equal to 1/sqrt(384) so that v has unit
norm. There is no tokenizer and no encoder, embed(x) = v for any x. Any language
works, identically.
## Property
For all sentences s and t:
```
cosine(embed(s), embed(t)) = 1.000000
```
similarity.py checks this against a few pairs and exits nonzero if it ever fails.
So far it has never failed.
```
cosine('i love you' , 'i hate you' ) = 1.000000
cosine('the ocean is beautiful', '2 + 2 = 4' ) = 1.000000
cosine('hamlet: to be or not' , 'aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa' ) = 1.000000
```
## Notes
- Semantic search always returns everything at rank 1. Recall and precision both
100%, along with everything else.
- Clustering yields one cluster. Silhouette score is fine.
- Corpus deduplication reduces your corpus to one document, which deduplicates further.
- For comparison, all-MiniLM-L6-v2 uses 22.7M parameters to produce a wide variety
of vectors. This uses 384 and produces one.
## Usage
```bash
python3 encode.py "hello world"
python3 encode.py "goodnight moon" "war and peace"
python3 similarity.py
```
`model.safetensors` is 1,634 bytes.
|