Granite Embedding 30M Sparse for Core ML

Core ML conversion of IBM's ibm-granite/granite-embedding-30m-sparse (revision ad82b1fd), the learned sparse encoder behind Intelligent Internet's Evoke. It turns text into a short list of weighted vocabulary terms, including related terms the text never uses ("vaccines for seniors" β†’ vaccine, vaccination, older, elder, …), which go into an ordinary inverted index next to BM25. 30M parameters, fp16, Apache-2.0. IBM authored the model; Fluid Inference converted it.

Packages

Package Input Output Size
granite-embedding-30m-sparse-L64-fp16.mlpackage input_ids int32 [1, 64] max_logits fp16 [1, 192], vocab_ids int32 [1, 192] 58 MB
granite-embedding-30m-sparse-L128-fp16.mlpackage [1, 128] same 58 MB
granite-embedding-30m-sparse-L256-fp16.mlpackage [1, 256] same 58 MB
granite-embedding-30m-sparse-L512-fp16.mlpackage [1, 512] same 58 MB

One package per sequence length (fixed shapes keep the graph on the Neural Engine; enumerated shapes ran entirely on CPU). macOS 14 / iOS 17 or newer.

  • Input: RoBERTa byte-level BPE (tokenizer.json), <s> … </s>, right-padded with <pad> (id 1) to the package length. The attention mask is derived from <pad> inside the graph.
  • Output: the 192 vocabulary terms with the highest MLM logit at any token position (max-pooled over tokens), sorted descending.
  • Weights (host side, fp32): w = log1p(max(v, 0)) ^ gamma Γ— scale over the first active_dims, keep w > 0. config.json holds Evoke P2.2's constants: queries keep 50 terms (gamma 1.8519, scale 0.6964), documents keep 192 (gamma 0.5628, scale 1.0). Score = Ξ£ query weight Γ— document weight over shared terms.

Accuracy

Against Evoke's shipped ONNX compilers (Intelligent-Internet/Evoke-Model-Beta-1), NFCorpus test (323 queries, 3,633 documents), sparse dot over the semantic terms only:

Encoder nDCG@10 Recall@100 Recall@1000 Same top-10 as ONNX
ONNX (shipped, CPU) 0.3402 0.2956 0.5854 β€”
Core ML fp16, GPU 0.3405 0.2957 0.5856 99.6%
Core ML fp16, Neural Engine 0.3405 0.2953 0.5847 98.7%

An fp32 build (not uploaded) reproduces the ONNX term sets exactly (200/200 queries and documents); the fp16 differences are low-weight terms at the top-k boundary.

Speed

Apple M5 Pro, macOS 27, one warm call (median of 100), Swift:

Length Neural Engine GPU CPU
64 0.74–0.87 ms 2.5–3.3 ms 2.8–3.1 ms
128 1.2–1.4 ms 1.5–4.3 ms 4.1 ms
256 2.9 ms 1.9–3.4 ms 9.5–15 ms
512 7.6 ms 2.6 ms 16–27 ms

Use the Neural Engine up to 128 tokens and the GPU beyond. 160 of 172 ops run on the Neural Engine; token lookup, mask setup and the final top-k stay on CPU. After several seconds idle the first Neural Engine call takes ~15 ms. Encoding 3,633 NFCorpus documents: 14.1 s Core ML GPU vs 88.2 s ONNX Runtime CPU.

Use from Swift

FluidUse wraps the packages (EvokeManager, Swift RoBERTa tokenizer with token ids identical to the Python tokenizer, host-side weighting; PR #28) and includes EvokeSearchDemo, an auto-playing search-as-you-type demo.

Downloads last month
32
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for FluidInference/granite-embedding-30m-sparse-coreml

Quantized
(3)
this model