ā” GLiNER2.5-Decide now runs natively on Apple Silicon, in MLX Swift
Fastino's 340M decision model scores every label you give it, or finds entity spans, in a single forward pass. Nothing is generated. I ported it to speech-swift and published three MLX conversions:
Measured on an idle M5 Pro, full request including tokenization: - INT8: 7.6 ms routing, 8.9 ms extraction, 0.85 GB peak - FP16: 8.8 / 10.0 ms, 1.58 GB - FP32: 11.1 / 12.6 ms, 2.55 GB (Python gliner2-mlx: 13.7 / 15.0 ms)
Every precision returns the same labels, spans and offsets as the PyTorch original on 24 reference cases. Most of the speedup came from one change: DeBERTa's relative-position projections depend only on the weights, so they are computed once at load instead of on every request.
I also compared it with TypeSafe's hosted Jev. Worth knowing: the published 60.2% vs 57.6% result is against JevK5, an open reproduction, not Jev itself.
ā” GLiNER2.5-Decide now runs natively on Apple Silicon, in MLX Swift
Fastino's 340M decision model scores every label you give it, or finds entity spans, in a single forward pass. Nothing is generated. I ported it to speech-swift and published three MLX conversions:
Measured on an idle M5 Pro, full request including tokenization: - INT8: 7.6 ms routing, 8.9 ms extraction, 0.85 GB peak - FP16: 8.8 / 10.0 ms, 1.58 GB - FP32: 11.1 / 12.6 ms, 2.55 GB (Python gliner2-mlx: 13.7 / 15.0 ms)
Every precision returns the same labels, spans and offsets as the PyTorch original on 24 reference cases. Most of the speedup came from one change: DeBERTa's relative-position projections depend only on the weights, so they are computed once at load instead of on every request.
I also compared it with TypeSafe's hosted Jev. Worth knowing: the published 60.2% vs 57.6% result is against JevK5, an open reproduction, not Jev itself.