File size: 1,341 Bytes
dd44243
 
 
 
 
 
 
 
49b416c
 
87fe5d9
49b416c
87fe5d9
 
 
49b416c
87fe5d9
49b416c
87fe5d9
49b416c
 
 
 
 
87fe5d9
 
49b416c
 
87fe5d9
 
 
49b416c
 
87fe5d9
49b416c
87fe5d9
 
 
 
 
 
49b416c
 
 
87fe5d9
49b416c
87fe5d9
 
49b416c
 
87fe5d9
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
---
license: mit
language:
- en
tags:
- ridiculous-models
---

# unity-embed

An embedding model where every input maps to the same vector.

384 parameters, one per dimension, all equal to 1/sqrt(384) so that v has unit
norm. There is no tokenizer and no encoder, embed(x) = v for any x. Any language
works, identically.

## Property

For all sentences s and t:

```
cosine(embed(s), embed(t)) = 1.000000
```

similarity.py checks this against a few pairs and exits nonzero if it ever fails.
So far it has never failed.

```
cosine('i love you'            , 'i hate you'                       ) = 1.000000
cosine('the ocean is beautiful', '2 + 2 = 4'                        ) = 1.000000
cosine('hamlet: to be or not'  , 'aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa' ) = 1.000000
```

## Notes

- Semantic search always returns everything at rank 1. Recall and precision both
  100%, along with everything else.
- Clustering yields one cluster. Silhouette score is fine.
- Corpus deduplication reduces your corpus to one document, which deduplicates further.
- For comparison, all-MiniLM-L6-v2 uses 22.7M parameters to produce a wide variety
  of vectors. This uses 384 and produces one.

## Usage

```bash
python3 encode.py "hello world"
python3 encode.py "goodnight moon" "war and peace"
python3 similarity.py
```

`model.safetensors` is 1,634 bytes.