bert-medium-nli / final /README.md
Pankaj8922's picture
Fine-tune bert-medium on high-quality NLI subset
05b0237 verified
|
Raw
History Blame Contribute Delete
1.96 kB
metadata
language: en
license: apache-2.0
tags:
  - text-classification
  - zero-shot-classification
  - nli
datasets:
  - Pankaj8922/nli-high-quality-balanced
base_model: prajjwal1/bert-medium
metrics:
  - accuracy
  - f1
model-index:
  - name: bert-medium-nli
    results:
      - task:
          type: text-classification
          name: Natural Language Inference
        dataset:
          name: Pankaj8922/nli-high-quality-balanced
          type: Pankaj8922/nli-high-quality-balanced
        metrics:
          - type: accuracy
            value: 0.8373
            name: Test Accuracy
          - type: f1
            value: 0.8372
            name: Test F1 (macro)

bert-medium-nli

Fine-tuned prajjwal1/bert-medium for natural language inference (entailment / neutral / contradiction), intended for use as a zero-shot text classification model via the entailment trick (hypothesis = "This text is about {label}.").

Trained on Pankaj8922/nli-high-quality-balanced, a combined and filtered subset of MNLI, SNLI, FEVER-NLI, and ANLI: annotator-agreement filtered, deduplicated, teacher-confidence filtered, hypothesis-only artifact filtered, and class-balanced.

Results

Split Accuracy F1 (macro) Precision (macro) Recall (macro)
Validation 0.8389 0.8389 0.8389 0.8389
Test 0.8373 0.8372 0.8372 0.8373

Training details

  • Base model: prajjwal1/bert-medium
  • Epochs: 3
  • Batch size: 64 (train), 128 (eval)
  • Learning rate: 5e-05
  • Max sequence length: 256

Labels

  • 0: entailment
  • 1: neutral
  • 2: contradiction

Intended use / limitations

This is a small (~41M parameter) model, so its ceiling on zero-shot performance against novel, unseen label sets is lower than larger NLI-tuned checkpoints (e.g. DeBERTa-v3-base or -large variants). Best suited for fast inference or resource-constrained settings rather than maximum accuracy.