multilingual-e5-small-rijksmuseum

Vocabulary-trimmed ONNX (int8) copy of Xenova/multilingual-e5-small (from intfloat/multilingual-e5-small), used as the query encoder of the Rijksmuseum MCP+ server.

The XLM-R tokenizer's 250,002 pieces were reduced to 114,541 by dropping rows of the word-embedding table and the matching tokenizer entries. All other weights are unchanged. Kept pieces:

  • every piece used by the server's corpus (artwork titles, descriptions, curatorial narratives, inscriptions, creator names and vocabulary labels);
  • every single-character piece, so no input maps to <unk>;
  • pieces of the 30,000 most frequent words (via wordfreq) in English, Dutch, German, French, Spanish, Italian and Polish, and the 3,000 most frequent in 19 further languages (incl. Russian, Greek, Arabic, Hebrew, Japanese, Chinese, Korean).

Unigram tokenization picks the highest-scoring segmentation, so any text whose original segmentation uses only kept pieces tokenizes and embeds identically to the base model; vectors produced with the base model remain compatible. Other text is split into smaller kept pieces, giving a close but not identical vector.

Loaded with transformers.js (dtype: "q8"), resident memory drops from about 796 to 414 MB.

Built with scripts/trim-embedding-vocab.py; trim_manifest.json records the kept ids and inputs. Use the e5 prefixes (query: / passage: ) as with the base model.

Downloads last month
29
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kintopp/multilingual-e5-small-rijksmuseum

Quantized
(291)
this model