yuvstk's picture
Upload README.md with huggingface_hub
daac0c6 verified
|
Raw
History Blame Contribute Delete
6.7 kB
---
license: apache-2.0
language:
- en
base_model: google-bert/bert-base-uncased
pipeline_tag: text-classification
tags:
- grammatical-error-detection
- linguistic-acceptability
- bert
- blimp
- cola
datasets:
- nyu-mll/blimp
- nyu-mll/glue
metrics:
- accuracy
- matthews_correlation
- f1
widget:
- text: "Katherine can't help himself."
example_title: "Reflexive agreement error"
- text: "The professor talked us."
example_title: "Verb argument error"
- text: "She has been working here since 2019."
example_title: "Correct sentence"
- text: "They drank the pub."
example_title: "Selectional restriction error"
---
# BERT for Grammatical Error Detection (BLiMP + CoLA)
`bert-base-uncased` fine-tuned for **binary grammatical error detection**: given
one English sentence, decide whether it contains a grammatical error.
| label | meaning |
|-------|---------|
| `0` | grammatical |
| `1` | ungrammatical |
Note the orientation: **1 means "has an error."** This is the inverse of CoLA's
native convention (where 1 = acceptable), and the training labels were flipped
accordingly.
## Usage
```python
from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch
model_id = "yuvstk/bert-grammar-error-detection"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
sentences = ["Katherine can't help himself.", "She went home early."]
inputs = tokenizer(sentences, padding=True, truncation=True,
max_length=64, return_tensors="pt")
with torch.no_grad():
probs = torch.softmax(model(**inputs).logits, dim=-1)
for sentence, prob in zip(sentences, probs):
label = int(prob.argmax())
print(f"{'UNGRAMMATICAL' if label else 'GRAMMATICAL'} "
f"({prob[label]:.1%}) — {sentence}")
```
Or with a pipeline (`LABEL_1` = ungrammatical):
```python
from transformers import pipeline
clf = pipeline("text-classification", model=model_id)
clf("The professor talked us.")
```
## Training data
Two sources merged into a single 4000 + 9594 sentence corpus:
| source | rows (train) | what it contributes |
|---|---|---|
| [BLiMP](https://huggingface.co/datasets/nyu-mll/blimp) — `anaphor_gender_agreement` + `anaphor_number_agreement` | 3200 | synthetic minimal pairs; reflexive pronoun agreement; exactly 50/50 balanced |
| [CoLA](https://huggingface.co/datasets/nyu-mll/glue) (GLUE) | 7695 | real linguistics-literature sentences; many error types; ~70/30 imbalanced |
Splitting differs per source, because the sources need different treatment:
- **BLiMP is split by `pair_id`**, never by row. The two sentences of a minimal
pair differ by exactly one word, so a row-level split would put a
near-duplicate of a test sentence into training.
- **CoLA is split by stratified rows.** GLUE's `test` split is unlabelled
(all `-1`), so GLUE `validation` is used as the test set and the validation
set is carved out of GLUE `train`.
Total: 10,895 train / 1,256 validation / 1,443 test.
## Results
Evaluated separately per source, because the two halves differ enormously in
difficulty — a single pooled number would mostly reflect the mixture ratio.
| test set | n | accuracy | precision | recall | F1 | MCC |
|---|---|---|---|---|---|---|
| **BLiMP** | 400 | **1.000** | 1.000 | 1.000 | 1.000 | 1.000 |
| **CoLA** | 1043 | **0.837** | 0.833 | 0.590 | 0.691 | **0.601** |
| pooled | 1443 | 0.882 | 0.911 | 0.747 | 0.821 | 0.743 |
CoLA MCC of 0.601 is in the normal published range for BERT-base (~0.55–0.60).
**Merging helped.** The same model trained on CoLA alone reached MCC 0.576;
adding BLiMP raised it to 0.601 while BLiMP itself stayed at 1.000 — positive
transfer, not interference.
### Baselines, for scale
| method | accuracy | MCC |
|---|---|---|
| majority class (CoLA) | 0.691 | 0.000 |
| bag-of-words logistic regression (CoLA) | 0.718 | 0.092 |
| pronoun-only rule (BLiMP) | 0.688 | — |
| **zero-shot `bert-base-uncased`, no fine-tuning** (BLiMP) | **0.973** | — |
That last row is worth dwelling on: masking the pronoun and asking the *raw*
pretrained model which word it prefers already solves BLiMP at 97.3%.
Fine-tuning on BLiMP mostly attaches an output head to knowledge the model
already had. CoLA is where fine-tuning does real work.
## Training procedure
| hyperparameter | value |
|---|---|
| base model | `bert-base-uncased` (109.5M parameters) |
| epochs | 4 |
| learning rate | 2e-5 |
| warmup ratio | 0.06 |
| batch size | 32 |
| max sequence length | 64 |
| weight decay | 0.01 |
| optimizer | AdamW |
| seed | 42 |
| best checkpoint by | validation **MCC** (not accuracy — the data is imbalanced) |
Per-epoch validation MCC: 0.727 → 0.735 → 0.771 → **0.775**.
## Limitations
Measured on 40 hand-written test sentences (33/40 correct, 82.5%), the failure
modes are systematic rather than random:
1. **Blind to omissions.** *"Although it was raining, we decided go for a
walk."* is judged correct. Both training sets create errors by
**substituting** a word, never deleting one, so the model never learned to
notice something missing.
2. **Over-flags correct sentences.** *"He is an honest man."* and *"She arrived
at the airport."* are both flagged as errors. Recall on CoLA is 0.590 while
precision is 0.833 — it misses more errors than it invents, but its false
alarms land on perfectly ordinary sentences.
3. **Untrained phenomena fail.** Determiner–noun agreement (*"Raymond is
selling this sketch."*) is flagged as an error. On BLiMP's
`determiner_noun_agreement_1` — a phenomenon never seen in training — the
model scores 0.675, far below its 1.000 on trained phenomena.
4. **Confidence is not reliability.** Several wrong predictions are made at
100% confidence. Do not treat the softmax score as a calibrated probability.
5. **English only**, and short sentences only — training data averaged well
under 20 words.
This is a coursework model built to study what fine-tuning contributes, not a
production grammar checker.
## Intended use
Educational and research use: demonstrating grammatical acceptability
classification, and comparing fine-tuned versus zero-shot versus from-scratch
transformers. Not suitable for grading student writing, automated proofreading,
or any decision affecting a person.
## Citation
```bibtex
@misc{bert-grammar-error-detection,
title = {BERT for Grammatical Error Detection (BLiMP + CoLA)},
author = {Your Real Name},
year = {2026},
url = {https://huggingface.co/yuvstk/bert-grammar-error-detection}
}
```
Datasets: BLiMP (Warstadt et al., TACL 2020) and CoLA (Warstadt et al., TACL
2019).