Download docs/embedding.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 10.1 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/embedding.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml/docs/embedding.md
-
curl -L -o embedding.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/embedding.md
Embedding in bankml and Savante
Introduction
A language model answers; an embedding model measures meaning. It turns a piece of text into a list of numbers (a vector) so that texts which mean similar things get vectors that point in similar directions. Search by embedding finds what a question is about, not only the words it shares.
mindX, the system Savante belongs to, keeps its memories this way. Its memory store (agents/memory_pgvector.py)
embeds every memory and document with bge-m3 into 1,024 numbers and stores them in PostgreSQL with pgvector.
bankml now uses the same model, so Savante's local history and mindX's memory can be searched by meaning in the same
way.
The embedding model is optional. Everything in bankml works without it. With it, the .history ragebar also
understands paraphrase, and a published agent's history becomes searchable by meaning in PostgreSQL.
bankML does not compute embeddings itself yet. bge-m3 is an XLM-R encoder, and bankML has no encoder graph:
bankml serve --native refuses /api/embed with HTTP 400 and the reason. The encoder graph, with /api/embed,
/v1/embeddings and an oracle against llama.cpp's bge-m3 output, is phase O7 of OLLAMA.md, planned for
0.6.0 (TODO.md). Until then Savante asks the local Ollama, as described below.
Summary
| model | bge-m3 (BAAI), the model mindX uses by default (MINDX_EMBED_MODEL, default bge-m3) |
| licence | MIT (read from the model's licence layer in the local Ollama store) |
| weights | 1.16 GB GGUF, served by the local Ollama (ollama pull bge-m3) |
| pinned by | the sha256 of that GGUF (Ollama's layer digest daec91ff…3062c on this laptop), recorded with every vector |
| output | 1,024 dimensions, normalized to unit length; the width of mindX's VECTOR(1024) and bankml's bankml_exchanges.embedding |
| input | each text cut at 4,000 characters, as mindX cuts it (≈ 1,000 tokens; bge-m3 reads up to 8,192) |
| used for | 1. the .history ragebar: words (BM25) and meaning (bge-m3) ranked together; 2. embedding in PostgreSQL publishing |
| without it | BM25 alone, and an empty embedding column; nothing fails |
| privacy | vectors of private text stay beside .history (<history>.emb) and leave the machine only if you publish private lines |
Explanation
Why an embedding model at all
The ragebar has always searched .history with BM25, which ranks exchanges by the words they share with the query,
weighted by how rare each word is. It is fast, exact and explainable, and it is blind to paraphrase: "how fast is
the ternary model" does not match an answer that says "Q2_0 decode took 2.4 s per token".
An embedding model closes that gap. bge-m3 maps the query and every exchange into the same 1,024-dimensional space, where the angle between two vectors reflects how close their meanings are. bankml does not replace BM25 with it. It fuses the two rankings, so an exact word match still counts and a match in meaning is added to it.
Why bge-m3
- It is what mindX uses. Savante is one of mindX's offices. Using the same model and width means her vectors and mindX's live in the same space: a history published from bankml can sit in the same database and be compared with mindX's memories without re-embedding either.
- It is open. MIT-licensed, like the rest of what bankml admits ("open source or go away").
- It reads long text. Its 8,192-token window holds a whole question and answer; mindX cuts at 4,000 characters and bankml does the same, so the same text yields the same vector.
- It is already on this computer. The local Ollama holds it; bankml adds no new runtime and downloads nothing.
How it stays out of the way
This laptop has about 1 GB of free memory while the chat model runs, and loading bge-m3 needs about 1.2 GB. So:
- bankml asks for an embedding only when the ragebar or a publish needs one, and Ollama unloads the model 60 seconds
after the last use (
keep_alive), giving the memory back; - before loading it, bankml checks free memory (1.3 GB by default) and, if there is not enough, uses BM25 alone rather than push the machine into swap;
- one embedding call runs at a time; a keystroke that finds it busy is answered by BM25, and the query's vector is cached so the next keystroke with the same text is instant;
- exchanges are embedded once, in the background, and cached by the sha256 of the embedded text. The ragebar never waits for indexing; it uses what is ready and says how much is.
Technical
The model and its identity
bge-m3 is served by Ollama (0.13.3 here) from its store (/usr/share/ollama/.ollama/models). bankml reads the
model's local manifest and takes two things from it: the model layer's digest, which is the sha256 of the GGUF
Ollama loads, and the licence layer, which it classifies (MIT). It never asks the registry: what is identified is
what is on disk. Every cached vector records that digest, and a cache written by different weights is ignored.
Calling it
POST http://127.0.0.1:11434/api/embed with {"model": "bge-m3", "input": [...], "keep_alive": "60s", "truncate": true}, the endpoint mindX uses (/api/embed, Ollama ≥ 0.3.4). The reply's vectors must be 1,024 long, or the call
is refused. bankml divides each by its length so that the cosine of two vectors is their dot product.
The private cache
<history>.emb beside .history (for Savante, ~/.local/share/bankml/savante/savante.history.emb), one JSON line per
exchange:
{"text_sha256": "…", "model": "bge-m3", "digest": "daec91ff…", "dims": 1024, "vec": "<1,024 float32, little-endian, base64>"}
The key is the sha256 of the exact text embedded: user + "\n" + assistant, cut at 4,000 characters. An edited or
new exchange gets a new key; nothing is ever re-embedded twice. The file is derived data: delete it and it is
rebuilt.
Ranking: reciprocal rank fusion
For a query, bankml ranks exchanges twice: by BM25 (mindX's rage.py when present, otherwise the built-in BM25),
and by cosine similarity to the query's vector. The rankings are fused by reciprocal rank fusion (Cormack, Clarke
and Büttcher 2009): each exchange scores Σ 1 / (60 + rank) over the rankings it appears in. RRF needs no tuning and
no score calibration between the two systems; an exchange near the top of either list rises, one near the top of both
rises most. The ragebar's status line names the engines and how many exchanges are embedded.
PostgreSQL
When an agent is published with its private lines (connectors.publish(…, include_private=True)), and the
database has pgvector, every exchange's vector is written to bankml_exchanges.embedding (vector(1024)) in the
same transaction as the lines themselves, and indexed by pgvectorscale's DiskANN when it is installed, else HNSW. A
query can then find an agent's exchanges by meaning in SQL:
SELECT seq, line FROM bankml_exchanges WHERE agent = 'ada'
ORDER BY embedding <=> '[…1,024 numbers…]'::vector LIMIT 5;
Public-only publishing sends no lines and so no vectors: a vector of private text is itself private.
Code
| file | what |
|---|---|
sAGI/embed.py |
status and provenance, embed(), the cache (index, index_async, cached), query_vector, fuse |
sAGI/savante.py |
history_search() fuses BM25 and bge-m3; _semantic() indexes in the background |
sAGI/connectors.py |
publish() writes embedding with the lines, in one transaction |
testing/test_ui.py |
fusion, the cache, and the fallback to BM25 (offline: a fake Ollama) |
Measured on this laptop (2026-09-29, Ryzen 3 3200U, Ollama 0.13.3)
| first call (load + embed one text) | 9.2 s |
| next call (three texts, model loaded) | 1.40 s |
| memory while loaded | 1.14 GB; free memory fell from 1.82 GB to 0.43 GB, which is why the 1.3 GB guard exists |
indexing the three real exchanges in .history |
11.5 s (once; cached afterwards) |
| meaning check (cosine to "How fast is the ternary model on this laptop?") | a fact about decode speed 0.448 · the oversight office's oath 0.299 |
| ragebar status | RAGE (mindX rage.py, BM25) + bge-m3 (meaning, 3 of 3 embedded), fused by reciprocal rank |
The earlier attempt the same day was refused by the guard (1.0 GB free), as designed; this one ran when other applications had released memory.
Usage
Have it. bge-m3 is in the local Ollama:
ollama pull bge-m3 # once; 1.16 GB. On this laptop it is already there.
ollama list | grep bge-m3
Or adopt it through the Models tab's Ollama section like any other model (it is an embedding model, so it is not offered as a chat carrier).
Use it. Nothing to switch on. Open .history and type in the ragebar. The status line says, for example:
3 of 3 exchanges · RAGE-shaped BM25 (built in) + bge-m3 (meaning, 3 of 3 embedded), fused by reciprocal rank
The first search starts indexing the history in the background; results improve as it completes. If bge-m3 cannot run (not pulled, Ollama down, not enough free memory), the status line names BM25 alone.
Publish with it. Agents tab → PostgreSQL → publish with private lines. The result reports
embedded: N of M exchanges embedded with bge-m3 (daec91ff…).
Settings (environment):
| variable | default | meaning |
|---|---|---|
BANKML_EMBED_MODEL |
bge-m3 |
the Ollama model (must produce 1,024 dimensions) |
BANKML_OLLAMA |
http://127.0.0.1:11434 |
the Ollama server (keep it local: history text is sent to it) |
BANKML_EMBED_KEEP_ALIVE |
60s |
how long Ollama keeps it loaded after a call |
BANKML_EMBED_NEED_GB |
1.3 |
free memory required before loading it |
Check it. From the repository root, python3 -B -c "import sys; sys.path.insert(0, 'sAGI'); import embed; print(embed.status())" prints the model, its digest and licence, and whether it is ready.