nomic-embed-text-v1.5 β ExecuTorch
A BERT with rotary embeddings and SwiGLU, trained so that a prefix of the vector is still a usable vector. Text in, one 768-dimensional vector out, for search and retrieval that never leaves the device.
- Source: nomic-ai/nomic-embed-text-v1.5 β 12 layers, 768 dimensions, 30,528 vocabulary
- License: apache-2.0
- Input:
input_idsandattention_mask, both[1, 256]int64 - Output:
[1, 768], mean-pooled and not normalised inside the graph
The recipe is in the graph, and it was read off this repo
sentence-transformers stores it per model, and the shelf's seven embedding models do
not agree. This one pools mean and
does not normalise, read from
1_Pooling/config.json and modules.json rather than inferred from the family name.
Getting it wrong does not throw; it returns vectors that look fine and rank wrong.
The prefix is not in the graph
This model is trained with search_query: in front of the text and expects it at
inference. That happens before tokenisation, so the .pte never sees it as anything
but tokens β and leaving it out does not throw. It returns a plausible vector that
retrieves worse.
Verification
| build | file | size (MB) | Mac ms* | backend takes | worst cosine vs eager | retrieval budget |
|---|---|---|---|---|---|---|
| fp32 | embed_nomic_embed_text_v15_xnnpack_fp32.pte |
547.2 | 54.8 | 58.8% | 1.000000 | 0% |
| fp16 | embed_nomic_embed_text_v15_xnnpack_fp16.pte |
273.9 | 129.7 | 51.9% | 0.999999 | 8% |
| Core ML (fp16, iOS) | embed_nomic_embed_text_v15_coreml_all.pte |
274.8 | 8.1 | 100.0% | 0.999795 | 44% |
*Mac arm64, median of 10, one 256-token sequence β a reference point for relative cost, not a device number. Torch eager fp32 on the same machine is 43.6 ms.
Cosine is measured against the model run in eager through its own pooling, over eight sentences. The last column is the one that decides: rank those eight against each other, and ask whether this build's score error is smaller than the gap between the document a query retrieves and the runner-up. Every shipped build keeps all eight top-1 results.
Matryoshka: the vector truncates
This model is trained so that a prefix of the vector is still a usable vector β 768 down to 512, 256, 128 or 64 dimensions, trading accuracy for index size. The graph returns the full 768, because the dimension is the caller's choice, and the truncation recipe is three lines:
import torch.nn.functional as F
v = F.layer_norm(v, (v.shape[1],)) # before truncating, not after
v = v[:, :dim] # 512 / 256 / 128 / 64
v = F.normalize(v, p=2, dim=1)
The layer_norm is what makes the prefix usable, and it is easy to skip. At the
full 768 it barely matters β measured on this shelf, going through the layer_norm
changes the direction of the vector by a cosine of 0.999944 and leaves the test pair's
score at 0.8233 either way. It earns its place only once you truncate.
The four prefixes are a real part of the model. search_document: for what goes
in the index, search_query: for what is asked of it, plus classification: and
clustering: . Unlike E5's symmetric mode, the two retrieval prefixes are not
interchangeable.
The architecture is not stock BERT. Rotary embeddings and a SwiGLU MLP, with the
modelling code in nomic-ai/nomic-bert-2048 rather than in transformers β loading it
needs trust_remote_code=True and einops installed. None of that reaches the .pte,
which is a plain graph once exported.
Not shipped: int8
embed_nomic_embed_text_v15_xnnpack_int8.pte is 208.0 MB β smaller than fp16's 273.9 MB, because the token embedding
table is only 94 MB of the 547.2 MB model (17%), leaving most of the
weight in linears for int8 to shrink.
It is withheld on the number that decides. Ranking the eight test sentences against each other, this build moves a pair score by at most 0.0368 while the closest fp32 decision β the gap between the document a query retrieves and the runner-up β is 0.0077. That is 480% of the room available, against a bar of 50%.
Correlation reads 0.992203 for this build, which no correlation gate would stop.
torch.export -> to_edge_transform_and_lower(partitioner) -> .pte (conversion scripts: executorch-models)
- Downloads last month
- 7
Model tree for mlboydaisuke/nomic-embed-text-v1.5-ExecuTorch
Base model
nomic-ai/nomic-embed-text-v1.5