--- license: apache-2.0 library_name: coremltools pipeline_tag: feature-extraction base_model: ibm-granite/granite-embedding-30m-sparse tags: - coreml - sparse - splade - sparse-encoder - information-retrieval - evoke - apple-silicon - neural-engine --- # Granite Embedding 30M Sparse for Core ML Core ML conversion of IBM's [`ibm-granite/granite-embedding-30m-sparse`](https://huggingface.co/ibm-granite/granite-embedding-30m-sparse) (revision `ad82b1fd`), the learned sparse encoder behind Intelligent Internet's [Evoke](https://github.com/Intelligent-Internet/Evoke). It turns text into a short list of weighted vocabulary terms, including related terms the text never uses ("vaccines for seniors" → vaccine, vaccination, older, elder, …), which go into an ordinary inverted index next to BM25. 30M parameters, fp16, Apache-2.0. IBM authored the model; Fluid Inference converted it. ## Packages | Package | Input | Output | Size | | --- | --- | --- | ---: | | `granite-embedding-30m-sparse-L64-fp16.mlpackage` | `input_ids` int32 `[1, 64]` | `max_logits` fp16 `[1, 192]`, `vocab_ids` int32 `[1, 192]` | 58 MB | | `granite-embedding-30m-sparse-L128-fp16.mlpackage` | `[1, 128]` | same | 58 MB | | `granite-embedding-30m-sparse-L256-fp16.mlpackage` | `[1, 256]` | same | 58 MB | | `granite-embedding-30m-sparse-L512-fp16.mlpackage` | `[1, 512]` | same | 58 MB | One package per sequence length (fixed shapes keep the graph on the Neural Engine; enumerated shapes ran entirely on CPU). macOS 14 / iOS 17 or newer. - **Input:** RoBERTa byte-level BPE (`tokenizer.json`), `` … ``, right-padded with `` (id 1) to the package length. The attention mask is derived from `` inside the graph. - **Output:** the 192 vocabulary terms with the highest MLM logit at any token position (max-pooled over tokens), sorted descending. - **Weights (host side, fp32):** `w = log1p(max(v, 0)) ^ gamma × scale` over the first `active_dims`, keep `w > 0`. `config.json` holds Evoke P2.2's constants: queries keep 50 terms (gamma 1.8519, scale 0.6964), documents keep 192 (gamma 0.5628, scale 1.0). Score = Σ query weight × document weight over shared terms. ## Accuracy Against Evoke's shipped ONNX compilers (`Intelligent-Internet/Evoke-Model-Beta-1`), NFCorpus test (323 queries, 3,633 documents), sparse dot over the semantic terms only: | Encoder | nDCG@10 | Recall@100 | Recall@1000 | Same top-10 as ONNX | | --- | ---: | ---: | ---: | ---: | | ONNX (shipped, CPU) | 0.3402 | 0.2956 | 0.5854 | — | | Core ML fp16, GPU | 0.3405 | 0.2957 | 0.5856 | 99.6% | | Core ML fp16, Neural Engine | 0.3405 | 0.2953 | 0.5847 | 98.7% | An fp32 build (not uploaded) reproduces the ONNX term sets exactly (200/200 queries and documents); the fp16 differences are low-weight terms at the top-k boundary. ## Speed Apple M5 Pro, macOS 27, one warm call (median of 100), Swift: | Length | Neural Engine | GPU | CPU | | --- | ---: | ---: | ---: | | 64 | **0.74–0.87 ms** | 2.5–3.3 ms | 2.8–3.1 ms | | 128 | **1.2–1.4 ms** | 1.5–4.3 ms | 4.1 ms | | 256 | 2.9 ms | **1.9–3.4 ms** | 9.5–15 ms | | 512 | 7.6 ms | **2.6 ms** | 16–27 ms | Use the Neural Engine up to 128 tokens and the GPU beyond. 160 of 172 ops run on the Neural Engine; token lookup, mask setup and the final top-k stay on CPU. After several seconds idle the first Neural Engine call takes ~15 ms. Encoding 3,633 NFCorpus documents: 14.1 s Core ML GPU vs 88.2 s ONNX Runtime CPU. ## Use from Swift [FluidUse](https://github.com/FluidInference/FluidUse) wraps the packages (`EvokeManager`, Swift RoBERTa tokenizer with token ids identical to the Python tokenizer, host-side weighting; [PR #28](https://github.com/FluidInference/FluidUse/pull/28)) and includes `EvokeSearchDemo`, an auto-playing search-as-you-type demo.