Instructions to use agentlans/fasttext-line-classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- fastText
How to use agentlans/fasttext-line-classifier with fastText:
from huggingface_hub import hf_hub_download import fasttext model = fasttext.load_model(hf_hub_download("agentlans/fasttext-line-classifier", "model.bin")) - Notebooks
- Google Colab
- Kaggle
user commited on
Commit ·
3691566
1
Parent(s): c43dfe3
Upload original model, change README
Browse files- README.md +9 -3
- line_classifier.bin +3 -0
README.md
CHANGED
|
@@ -18,11 +18,11 @@ tags:
|
|
| 18 |
A lightweight model built with [FastText](https://fasttext.cc/) designed to evaluate individual lines of text and classify them as **keep**, **delete**, or **edit** to improve overall corpus quality.
|
| 19 |
|
| 20 |
> [!NOTE]
|
| 21 |
-
> This project is inspired by and based on an analysis of the [UltraX project](https://huggingface.co/datasets/openbmb/UltraX-Preview)
|
| 22 |
|
| 23 |
## 🏋️ Training & Methodology
|
| 24 |
|
| 25 |
-
1. The model was trained using FastText's supervised learning with **autotune** enabled, utilizing the train and validation splits from [`agentlans/openbmb-UltraX-Preview-line-classification`](https://huggingface.co/datasets/agentlans/openbmb-UltraX-Preview-line-classification).
|
| 26 |
2. After training, the model was compressed using FastText's built-in quantization (`.ftz` format) to significantly reduce file size and memory footprint while maintaining robust classification performance.
|
| 27 |
3. Evaluated on the independent test split of `agentlans/openbmb-UltraX-Preview-line-classification`, achieving the following overall metrics:
|
| 28 |
|
|
@@ -33,7 +33,13 @@ A lightweight model built with [FastText](https://fasttext.cc/) designed to eval
|
|
| 33 |
|
| 34 |
## 🚀 Quick Start
|
| 35 |
|
| 36 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 37 |
|
| 38 |
```python
|
| 39 |
import fasttext
|
|
|
|
| 18 |
A lightweight model built with [FastText](https://fasttext.cc/) designed to evaluate individual lines of text and classify them as **keep**, **delete**, or **edit** to improve overall corpus quality.
|
| 19 |
|
| 20 |
> [!NOTE]
|
| 21 |
+
> This project is inspired by and based on an analysis of the [UltraX project](https://huggingface.co/datasets/openbmb/UltraX-Preview). It is not officially affiliated with it.
|
| 22 |
|
| 23 |
## 🏋️ Training & Methodology
|
| 24 |
|
| 25 |
+
1. The model (`line_classifier.bin`) was trained using FastText's supervised learning with **autotune** enabled, utilizing the train and validation splits from [`agentlans/openbmb-UltraX-Preview-line-classification`](https://huggingface.co/datasets/agentlans/openbmb-UltraX-Preview-line-classification).
|
| 26 |
2. After training, the model was compressed using FastText's built-in quantization (`.ftz` format) to significantly reduce file size and memory footprint while maintaining robust classification performance.
|
| 27 |
3. Evaluated on the independent test split of `agentlans/openbmb-UltraX-Preview-line-classification`, achieving the following overall metrics:
|
| 28 |
|
|
|
|
| 33 |
|
| 34 |
## 🚀 Quick Start
|
| 35 |
|
| 36 |
+
1. Install FastText
|
| 37 |
+
|
| 38 |
+
```bash
|
| 39 |
+
pip install fasttext
|
| 40 |
+
```
|
| 41 |
+
|
| 42 |
+
2. Download the [`line_classifier.bin`](https://huggingface.co/agentlans/fasttext-line-classifier/resolve/main/line_classifier.bin) or [`line_classifier.ftz`](https://huggingface.co/agentlans/fasttext-line-classifier/resolve/main/line_classifier.ftz) file from this repository, then load and run predictions using Python:
|
| 43 |
|
| 44 |
```python
|
| 45 |
import fasttext
|
line_classifier.bin
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:fc6680a659cbaa5be9b2e4542a861605e5e1fda0e22bfc305a3b22d489a137d3
|
| 3 |
+
size 1841576421
|