Download README.md from FluidInference/granite-embedding-30m-sparse-coreml: direct link, hf CLI and curl.
- Browser
- Download file 3.81 kB
-
https://huggingface.co/FluidInference/granite-embedding-30m-sparse-coreml/resolve/main/README.md
- Command line
-
hf download hf://FluidInference/granite-embedding-30m-sparse-coreml/README.md
-
curl -L -o README.md https://huggingface.co/FluidInference/granite-embedding-30m-sparse-coreml/resolve/main/README.md
license: apache-2.0
library_name: coremltools
pipeline_tag: feature-extraction
base_model: ibm-granite/granite-embedding-30m-sparse
tags:
- coreml
- sparse
- splade
- sparse-encoder
- information-retrieval
- evoke
- apple-silicon
- neural-engine
Granite Embedding 30M Sparse for Core ML
Core ML conversion of IBM's ibm-granite/granite-embedding-30m-sparse
(revision ad82b1fd), the learned sparse encoder behind Intelligent Internet's
Evoke. It turns text into a short list of weighted vocabulary terms,
including related terms the text never uses ("vaccines for seniors" β vaccine, vaccination, older, elder, β¦), which go
into an ordinary inverted index next to BM25. 30M parameters, fp16, Apache-2.0. IBM authored the model; Fluid Inference
converted it.
Packages
| Package | Input | Output | Size |
|---|---|---|---|
granite-embedding-30m-sparse-L64-fp16.mlpackage |
input_ids int32 [1, 64] |
max_logits fp16 [1, 192], vocab_ids int32 [1, 192] |
58 MB |
granite-embedding-30m-sparse-L128-fp16.mlpackage |
[1, 128] |
same | 58 MB |
granite-embedding-30m-sparse-L256-fp16.mlpackage |
[1, 256] |
same | 58 MB |
granite-embedding-30m-sparse-L512-fp16.mlpackage |
[1, 512] |
same | 58 MB |
One package per sequence length (fixed shapes keep the graph on the Neural Engine; enumerated shapes ran entirely on CPU). macOS 14 / iOS 17 or newer.
- Input: RoBERTa byte-level BPE (
tokenizer.json),<s>β¦</s>, right-padded with<pad>(id 1) to the package length. The attention mask is derived from<pad>inside the graph. - Output: the 192 vocabulary terms with the highest MLM logit at any token position (max-pooled over tokens), sorted descending.
- Weights (host side, fp32):
w = log1p(max(v, 0)) ^ gamma Γ scaleover the firstactive_dims, keepw > 0.config.jsonholds Evoke P2.2's constants: queries keep 50 terms (gamma 1.8519, scale 0.6964), documents keep 192 (gamma 0.5628, scale 1.0). Score = Ξ£ query weight Γ document weight over shared terms.
Accuracy
Against Evoke's shipped ONNX compilers (Intelligent-Internet/Evoke-Model-Beta-1), NFCorpus test (323 queries,
3,633 documents), sparse dot over the semantic terms only:
| Encoder | nDCG@10 | Recall@100 | Recall@1000 | Same top-10 as ONNX |
|---|---|---|---|---|
| ONNX (shipped, CPU) | 0.3402 | 0.2956 | 0.5854 | β |
| Core ML fp16, GPU | 0.3405 | 0.2957 | 0.5856 | 99.6% |
| Core ML fp16, Neural Engine | 0.3405 | 0.2953 | 0.5847 | 98.7% |
An fp32 build (not uploaded) reproduces the ONNX term sets exactly (200/200 queries and documents); the fp16 differences are low-weight terms at the top-k boundary.
Speed
Apple M5 Pro, macOS 27, one warm call (median of 100), Swift:
| Length | Neural Engine | GPU | CPU |
|---|---|---|---|
| 64 | 0.74β0.87 ms | 2.5β3.3 ms | 2.8β3.1 ms |
| 128 | 1.2β1.4 ms | 1.5β4.3 ms | 4.1 ms |
| 256 | 2.9 ms | 1.9β3.4 ms | 9.5β15 ms |
| 512 | 7.6 ms | 2.6 ms | 16β27 ms |
Use the Neural Engine up to 128 tokens and the GPU beyond. 160 of 172 ops run on the Neural Engine; token lookup, mask setup and the final top-k stay on CPU. After several seconds idle the first Neural Engine call takes ~15 ms. Encoding 3,633 NFCorpus documents: 14.1 s Core ML GPU vs 88.2 s ONNX Runtime CPU.
Use from Swift
FluidUse wraps the packages (EvokeManager, Swift RoBERTa tokenizer with token ids identical to the Python tokenizer,
host-side weighting; PR #28) and includes EvokeSearchDemo, an
auto-playing search-as-you-type demo.