--- license: other license_name: lfm-open-license-v1.0 license_link: https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M-Spellchecker/blob/main/LICENSE library_name: minima-lfm base_model: LiquidAI/LFM2.5-Encoder-350M-Spellchecker pipeline_tag: token-classification tags: - ternary - 1.58-bit - grammatical-error-correction - spell-check --- # Minima Spellcheck Packed W1.58A8 adaptation of [LiquidAI/LFM2.5-Encoder-350M-Spellchecker](https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M-Spellchecker), built with [SSHDotCodes/minima](https://github.com/SSHDotCodes/minima). - Logical matrix values: `{-1, 0, +1}` (1.585 bits) - Physical artifact format: I2_S, four trits per byte - Group size: 32; recovery rank: 128 - Full encoder context: 8,192 tokens - Weight file: 226.0 MB - CPU demo: [Minima Spellcheck](https://huggingface.co/spaces/ProCreations/minima-spellcheck) ## Use ```bash pip install "minima-lfm @ git+https://github.com/SSHDotCodes/minima.git" ``` ```python from minima import MinimaModel model = MinimaModel.from_pretrained("ProCreations/minima-spellcheck", device="cpu") print(model.correct( ["I has went to the stor yesterday ."], max_iter=4, min_error_prob=0.0, rerank=False, )) # ['I went to the store yesterday .'] ``` CPU inference defaults to a one-time FBGEMM INT8 packing of each effective ternary-plus-recovery matrix. Set `MINIMA_CPU_BACKEND=i2s` to execute the strict 2-bit AVX2/NEON kernel instead. ## Validation The 1,000-step distillation run used held-out corrupted FineWeb text and the LiquidAI model as teacher. With both models' optional dense reranker disabled: | Diagnostic | Result | |---|---:| | Tag top-1 agreement within teacher candidates | 99.51% | | Error-detection top-1 agreement | 99.41% | | Exact correction agreement, held-out + reference examples | 70.0% (14/20) | | Exact agreement on the four published-style examples | 100% (4/4) | The Space smoke test corrected `I has went to the stor yesterday .` to `I went to the store yesterday.` and returned a warmed CPU latency of 110 ms. These are teacher-agreement diagnostics, not an ERRANT benchmark. The upstream optional reranker is deliberately disabled because it contains a separate 1.42 GB dense encoder, which would invalidate this model's CPU memory profile. See `spellcheck_report.json` for the full immutable training and evaluation log. ## License The weights remain subject to the LFM Open License v1.0 shipped in this repository. The Minima runtime code is MIT licensed.