Instructions to use smdesai/GLiNER25-Multi-Decide-FP16-CoreML with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER2
How to use smdesai/GLiNER25-Multi-Decide-FP16-CoreML with GLiNER2:
from gliner2 import GLiNER2 model = GLiNER2.from_pretrained("smdesai/GLiNER25-Multi-Decide-FP16-CoreML") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
GLiNER2.5-multi-Decide β Core ML (iOS 18+ / macOS 15+)
| File | Purpose |
|---|---|
GLiNER25-multi-Decide-FP16.mlpackage/ |
Core ML package, FP16 weights and compute (208 MB), three functions sharing one set of weights |
word_embeddings.f16 |
Word-embedding table, FP16 (384 MB), looked up on the host |
tokenizer.json, tokenizer_config.json |
Checkpoint tokenizer (mDeBERTa-v3, 250k-piece Unigram) |
model-card.md |
Upstream model card (fastino/GLiNER2.5-multi-Decide, Apache-2.0) |
Source: checkpoint revision 6bc1d43d201b0691e733626389af8c57eea3ea68, classification only
(mDeBERTa-v3-base encoder + classifier MLP). Built with coremltools 9.0 / Torch 2.7.1 from
FP16 128/256/512 exports (constant relative-position buckets, finite FP16 attention mask),
merged into one multifunction package with shared weights.
Interface
Functions context128, context256, context512 (select with
MLModelConfiguration.functionName):
inputs_embedsfloat16[1, L, 768]: rowinput_ids[i]ofword_embeddings.f16at positioni, including padding positions (pad id 0)attention_maskint32[1, L](1 = real token, 0 = padding)- output:
logits[1, L], one score per token position. Classification reads the positions of the schema's label markers ([L]).
Choose the smallest function that fits the request and pad to its length. The maximum is 512 tokens. Reject longer requests; do not truncate.
Word embeddings: word_embeddings.f16
The 250,112 Γ 768 word-embedding table (69% of the parameters) is kept outside the model, so the model file is small and only the rows a request uses are read. The file is little-endian FP16, row-major, no header (384,172,032 bytes). Memory-map it and copy one 1,536-byte row per token:
let table = try Data(contentsOf: tableURL, options: .alwaysMapped)
let rowBytes = 768 * MemoryLayout<Float16>.size
table.withUnsafeBytes { raw in
for (i, id) in inputIDs.enumerated() { // destination: the inputs_embeds buffer
memcpy(destination + i * 768, raw.baseAddress! + id * rowBytes, rowBytes)
}
}
The embedding LayerNorm and everything after it are in the model, so results are bit-identical to an in-model lookup. With the file in the page cache a lookup takes 40β75 Β΅s on an iPhone 17 Pro; the first requests after install, while rows are read from storage, take 1β5 ms longer. The Core ML and Core AI conversions use the same file.
Input preparation
Token IDs come from the upstream GLiNER2 processor (classify_text layout: schema
( [P] task ( [L] label ... ) ), [SEP_TEXT], then the lower-cased text words, each piece
tokenized on its own, no [CLS]/[SEP]). A Swift port that matches the Python processor
exactly (3,099 pieces and 153 requests in 26 languages) is in the conversion repository.
Required: GPU compute units
let configuration = MLModelConfiguration()
configuration.computeUnits = .cpuAndGPU
configuration.functionName = "context256"
let model = try MLModel(contentsOf: compiledURL, configuration: configuration)
Do not rely on .cpuOnly: FP16 on the CPU changes no decisions but exceeds the strict gate
(max 0.0116 on a Mac). Xcode compiles the .mlpackage to .mlmodelc when it is bundled in
an app; use MLModel.compileModel(at:) otherwise.
Validation (iPhone 17 Pro, iOS 27.2, GPU)
Strict gate: every decision matches the FP32 PyTorch oracle and the maximum probability error is β€ 0.005. Corpora: 43 (128), 80 (256) and 113 (512) requests in English and 25 other languages and scripts, including requests of 257β512 tokens. Three launches, identical results every launch.
| Function | Max prob. error | Median | Footprint peak |
|---|---|---|---|
| context128 | 0.0032 | 8.4β9.2 ms | 62 MiB |
| context256 | 0.0038 | 14.1β14.3 ms | 67 MiB |
| context512 | 0.0038 | 39β46 ms | 94 MiB |
Footprint is from single-function launches; with all three functions loaded in one process the peak is 85β99 MiB, including the first launch after install. With the embedding table inside the model instead, that first launch peaked at 476 MiB. Also passes on a Mac (M3 Max) GPU: max error 0.0034 / 0.0034 / 0.0038. Not yet validated on an iOS 18β26 device or an older GPU.
SHA-256
ce4b3e5e06632cc00bea4419c6dc9a7e3feab2139b0c871c956dd176868341f7 GLiNER25-multi-Decide-FP16.mlpackage/Data/com.apple.CoreML/weights/weight.bin
5f908aed7b4388d6324f17023eaed44a9bcc1484ef8a86f0b180d622a0d72273 GLiNER25-multi-Decide-FP16.mlpackage/Data/com.apple.CoreML/model.mlmodel
2248b04176406fb2f8776b1df5849b821c0e416fdbb0c16f9772bc68eeb434f0 word_embeddings.f16
c62446df87ae18ec98b133f8f84fc449a07cc89bbf8ef192a4cb5f9c53777a7a tokenizer.json
- Downloads last month
- 8
Model tree for smdesai/GLiNER25-Multi-Decide-FP16-CoreML
Base model
fastino/gliner2.5-multi-v1