iamleonie's picture
Initial public commit
4e2f724
|
Raw
History Blame Contribute Delete
3.86 kB
---
language:
- en
- de
- fr
- es
- it
- pt
- nl
- ru
- ja
- zh
- ko
tags:
- liquid
- lfm2
- lfm2.5
- bidirectional
- masked-lm
- encoder
- grammatical-error-correction
- gec
- spell-check
- token-classification
- gector
library_name: transformers
license: other
license_name: lfm1.0
license_link: LICENSE
pipeline_tag: token-classification
base_model:
- LiquidAI/LFM2.5-Encoder-350M
---
<div align="center">
<img
src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png"
alt="Liquid AI"
style="width: 100%; max-width: 100%; height: auto; display: inline-block; margin-bottom: 0.5em; margin-top: 0.5em;"
/>
<div style="display: flex; justify-content: center; gap: 0.5em; margin-bottom: 1em;">
<a href="https://playground.liquid.ai/"><strong>Try LFM</strong></a> β€’
<a href="https://docs.liquid.ai/lfm/getting-started/welcome"><strong>Docs</strong></a> β€’
<a href="https://leap.liquid.ai/"><strong>LEAP</strong></a> β€’
<a href="https://discord.com/invite/liquid-ai"><strong>Discord</strong></a>
</div>
</div>
# LFM2.5-Encoder-350-Spellchecker
A full fine-tune of [LFM2.5-Encoder-350M](https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M) with a subword-level **GECToR-style grammatical-error-correction tagger**.
It covers grammar, spelling, punctuation, and casing in English.
Find more details about our encoders in our [blog post](https://www.liquid.ai/blog/lfm2-5-encoders).
> [!NOTE]
> πŸ’» **Demos**: Try this fine-tuned model running in a CPU-only Hugging Face space:
> **[Spell checking](https://huggingface.co/spaces/LiquidAI/spellchecker)** β€” correct misspellings token by token.
## Usage
> ⚠️ Loads custom code via `trust_remote_code=True` (the model wraps a `trust_remote_code` encoder).
Install the required packages:
```bash
pip install torch transformers
```
Run spell checking:
```python
from transformers import AutoModel
model_id = "LiquidAI/LFM2.5-Encoder-350-Spellchecker"
model = AutoModel.from_pretrained(
model_id,
trust_remote_code=True,
).float().eval()
print(model.correct(["She go to school every day ."]))
# ['She goes to school every day .']
```
`correct()` accepts a string or a list; tune precision with `min_error_prob` (higher β†’ fewer edits) and
`max_iter` (refinement passes). Input should be whitespace-tokenized (punctuation separated by spaces),
matching the training data.
## Evaluation
Fixed inference setting: `max_iter=4`, precision knobs off. Headline [ERRANT](https://github.com/chrisjbryant/errant) F0.5:
**MASTER composite (selection metric): 64.24**
| Benchmark | Precision | Recall | F0.5 |
|---|--:|--:|--:|
| LOCNESS native (ERRANT) | 53.77 | 34.53 | 48.38 |
| BEA-dev (ERRANT) | 56.48 | 28.96 | 47.46 |
| CoNLL-14 (ERRANT) | 67.54 | 18.91 | 44.59 |
| FCE-test (ERRANT) | 57.31 | 32.63 | 49.78 |
| Robustness (ERRANT) | 91.96 | 87.98 | 91.14 |
| Multilingual dev (F0.5) | β€” | β€” | β€” |
## Examples
| Input | Correction |
|---|---|
| `She go to school every day .` | `She goes to school every day .` |
| `I has went to the stor yesterday .` | `I went to the store yesterday .` |
| `Their are many reason to study hard .` | `There are many reasons to study hard .` |
| `He don&#x27;t like coffee but he like tea .` | `He does n&#x27;t like coffee , but he likes tea .` |
## πŸ“¬ Contact
- Got questions or want to connect? [Join our Discord community](https://discord.com/invite/liquid-ai)
- If you are interested in custom solutions with edge deployment, please contact [our sales team](https://www.liquid.ai/contact).
## Citation
```bibtex
@article{liquidAI2026Encoders,
author = {Liquid AI},
title = {LFM2.5-Encoders: Fast at Long Context, Even on CPU},
journal = {Liquid AI Blog},
year = {2026},
note = {www.liquid.ai/blog/lfm2-5-encoders},
}
```