--- language: - en - de - fr - es - it - pt - nl - ru - ja - zh - ko tags: - liquid - lfm2 - lfm2.5 - bidirectional - masked-lm - encoder - grammatical-error-correction - gec - spell-check - token-classification - gector library_name: transformers license: other license_name: lfm1.0 license_link: LICENSE pipeline_tag: token-classification base_model: - LiquidAI/LFM2.5-Encoder-350M ---
# LFM2.5-Encoder-350-Spellchecker A full fine-tune of [LFM2.5-Encoder-350M](https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M) with a subword-level **GECToR-style grammatical-error-correction tagger**. It covers grammar, spelling, punctuation, and casing in English. Find more details about our encoders in our [blog post](https://www.liquid.ai/blog/lfm2-5-encoders). > [!NOTE] > 💻 **Demos**: Try this fine-tuned model running in a CPU-only Hugging Face space: > **[Spell checking](https://huggingface.co/spaces/LiquidAI/spellchecker)** — correct misspellings token by token. ## Usage > ⚠️ Loads custom code via `trust_remote_code=True` (the model wraps a `trust_remote_code` encoder). Install the required packages: ```bash pip install torch transformers ``` Run spell checking: ```python from transformers import AutoModel model_id = "LiquidAI/LFM2.5-Encoder-350-Spellchecker" model = AutoModel.from_pretrained( model_id, trust_remote_code=True, ).float().eval() print(model.correct(["She go to school every day ."])) # ['She goes to school every day .'] ``` `correct()` accepts a string or a list; tune precision with `min_error_prob` (higher → fewer edits) and `max_iter` (refinement passes). Input should be whitespace-tokenized (punctuation separated by spaces), matching the training data. ## Evaluation Fixed inference setting: `max_iter=4`, precision knobs off. Headline [ERRANT](https://github.com/chrisjbryant/errant) F0.5: **MASTER composite (selection metric): 64.24** | Benchmark | Precision | Recall | F0.5 | |---|--:|--:|--:| | LOCNESS native (ERRANT) | 53.77 | 34.53 | 48.38 | | BEA-dev (ERRANT) | 56.48 | 28.96 | 47.46 | | CoNLL-14 (ERRANT) | 67.54 | 18.91 | 44.59 | | FCE-test (ERRANT) | 57.31 | 32.63 | 49.78 | | Robustness (ERRANT) | 91.96 | 87.98 | 91.14 | | Multilingual dev (F0.5) | — | — | — | ## Examples | Input | Correction | |---|---| | `She go to school every day .` | `She goes to school every day .` | | `I has went to the stor yesterday .` | `I went to the store yesterday .` | | `Their are many reason to study hard .` | `There are many reasons to study hard .` | | `He don't like coffee but he like tea .` | `He does n't like coffee , but he likes tea .` | ## 📬 Contact - Got questions or want to connect? [Join our Discord community](https://discord.com/invite/liquid-ai) - If you are interested in custom solutions with edge deployment, please contact [our sales team](https://www.liquid.ai/contact). ## Citation ```bibtex @article{liquidAI2026Encoders, author = {Liquid AI}, title = {LFM2.5-Encoders: Fast at Long Context, Even on CPU}, journal = {Liquid AI Blog}, year = {2026}, note = {www.liquid.ai/blog/lfm2-5-encoders}, } ```