---
license: apache-2.0
library_name: coremltools
pipeline_tag: feature-extraction
base_model: ibm-granite/granite-embedding-30m-sparse
tags:
- coreml
- sparse
- splade
- sparse-encoder
- information-retrieval
- evoke
- apple-silicon
- neural-engine
---
# Granite Embedding 30M Sparse for Core ML
Core ML conversion of IBM's [`ibm-granite/granite-embedding-30m-sparse`](https://huggingface.co/ibm-granite/granite-embedding-30m-sparse)
(revision `ad82b1fd`), the learned sparse encoder behind Intelligent Internet's
[Evoke](https://github.com/Intelligent-Internet/Evoke). It turns text into a short list of weighted vocabulary terms,
including related terms the text never uses ("vaccines for seniors" → vaccine, vaccination, older, elder, …), which go
into an ordinary inverted index next to BM25. 30M parameters, fp16, Apache-2.0. IBM authored the model; Fluid Inference
converted it.
## Packages
| Package | Input | Output | Size |
| --- | --- | --- | ---: |
| `granite-embedding-30m-sparse-L64-fp16.mlpackage` | `input_ids` int32 `[1, 64]` | `max_logits` fp16 `[1, 192]`, `vocab_ids` int32 `[1, 192]` | 58 MB |
| `granite-embedding-30m-sparse-L128-fp16.mlpackage` | `[1, 128]` | same | 58 MB |
| `granite-embedding-30m-sparse-L256-fp16.mlpackage` | `[1, 256]` | same | 58 MB |
| `granite-embedding-30m-sparse-L512-fp16.mlpackage` | `[1, 512]` | same | 58 MB |
One package per sequence length (fixed shapes keep the graph on the Neural Engine; enumerated shapes ran entirely on
CPU). macOS 14 / iOS 17 or newer.
- **Input:** RoBERTa byte-level BPE (`tokenizer.json`), `` … ``, right-padded with `` (id 1) to the
package length. The attention mask is derived from `` inside the graph.
- **Output:** the 192 vocabulary terms with the highest MLM logit at any token position (max-pooled over tokens),
sorted descending.
- **Weights (host side, fp32):** `w = log1p(max(v, 0)) ^ gamma × scale` over the first `active_dims`, keep `w > 0`.
`config.json` holds Evoke P2.2's constants: queries keep 50 terms (gamma 1.8519, scale 0.6964), documents keep 192
(gamma 0.5628, scale 1.0). Score = Σ query weight × document weight over shared terms.
## Accuracy
Against Evoke's shipped ONNX compilers (`Intelligent-Internet/Evoke-Model-Beta-1`), NFCorpus test (323 queries,
3,633 documents), sparse dot over the semantic terms only:
| Encoder | nDCG@10 | Recall@100 | Recall@1000 | Same top-10 as ONNX |
| --- | ---: | ---: | ---: | ---: |
| ONNX (shipped, CPU) | 0.3402 | 0.2956 | 0.5854 | — |
| Core ML fp16, GPU | 0.3405 | 0.2957 | 0.5856 | 99.6% |
| Core ML fp16, Neural Engine | 0.3405 | 0.2953 | 0.5847 | 98.7% |
An fp32 build (not uploaded) reproduces the ONNX term sets exactly (200/200 queries and documents); the fp16 differences
are low-weight terms at the top-k boundary.
## Speed
Apple M5 Pro, macOS 27, one warm call (median of 100), Swift:
| Length | Neural Engine | GPU | CPU |
| --- | ---: | ---: | ---: |
| 64 | **0.74–0.87 ms** | 2.5–3.3 ms | 2.8–3.1 ms |
| 128 | **1.2–1.4 ms** | 1.5–4.3 ms | 4.1 ms |
| 256 | 2.9 ms | **1.9–3.4 ms** | 9.5–15 ms |
| 512 | 7.6 ms | **2.6 ms** | 16–27 ms |
Use the Neural Engine up to 128 tokens and the GPU beyond. 160 of 172 ops run on the Neural Engine; token lookup, mask
setup and the final top-k stay on CPU. After several seconds idle the first Neural Engine call takes ~15 ms. Encoding
3,633 NFCorpus documents: 14.1 s Core ML GPU vs 88.2 s ONNX Runtime CPU.
## Use from Swift
[FluidUse](https://github.com/FluidInference/FluidUse) wraps the packages (`EvokeManager`, Swift RoBERTa tokenizer with token ids identical to the Python tokenizer,
host-side weighting; [PR #28](https://github.com/FluidInference/FluidUse/pull/28)) and includes `EvokeSearchDemo`, an
auto-playing search-as-you-type demo.