How to use from the
Use from the
Transformers.js library
// npm i @huggingface/transformers
import { pipeline } from '@huggingface/transformers';

// Allocate pipeline
const pipe = await pipeline('feature-extraction', 'kintopp/multilingual-e5-base-iconclass');

multilingual-e5-base-iconclass

Vocabulary-trimmed ONNX (int8) copy of Xenova/multilingual-e5-base (from intfloat/multilingual-e5-base), used as the query encoder of the Rijksmuseum Iconclass MCP server.

The XLM-R tokenizer's 250,002 pieces were reduced to 116,522 by dropping rows of the word-embedding table and the matching tokenizer entries. All other weights are unchanged. Kept pieces:

  • every piece used by the server's corpus (Iconclass texts and keywords in all 13 languages);
  • every single-character piece, so no input maps to <unk>;
  • pieces of the 30,000 most frequent words (via wordfreq) in the 13 Iconclass languages (English, German, French, Italian, Portuguese, Finnish, Dutch, Spanish, Polish, Czech, Hungarian, Japanese, Chinese), and the 3,000 most frequent in 13 further languages (incl. Russian, Greek, Arabic, Hebrew, Korean, Scandinavian).

Unigram tokenization picks the highest-scoring segmentation, so any text whose original segmentation uses only kept pieces tokenizes and embeds identically to the base model; vectors produced with the base model remain compatible. Other text is split into smaller kept pieces, giving a close but not identical vector.

Loaded with transformers.js (dtype: "q8"), resident memory drops from about 1,330 to 776 MB.

Built with scripts/trim-embedding-vocab.py (in the Rijksmuseum MCP+ repo); trim_manifest.json records the kept ids and inputs. Use the e5 prefixes (query: / passage: ) as with the base model.

Downloads last month
30
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kintopp/multilingual-e5-base-iconclass

Quantized
(272)
this model