bert-base-banking77-pt2

A fine-tuned version of bert-base-uncased for customer-intent classification on banking queries, trained on the PolyAI/banking77 dataset.

On the evaluation set it achieves:

  • F1 (weighted): 0.9362
  • Loss: 0.3089

Model description

BANKING77 is a fine-grained intent detection benchmark: 13,083 real online-banking customer queries labelled across 77 intents (e.g. card_arrival, exchange_rate, pending_top_up). The label space is intentionally hard — many intents are near-duplicates in wording but differ in meaning, which is exactly where bag-of-words baselines break down.

This model puts a linear classification head on top of bert-base-uncased and fine-tunes the full encoder for 10 epochs. The full id2label mapping for all 77 intents is stored in config.json, so predictions come back as readable intent names rather than integer indices.

Intended uses & limitations

Intended use. Routing or triaging short, English, first-person banking queries — chatbot intent detection, support-ticket tagging, analytics over customer messages.

Limitations.

  • Closed label set. The model always predicts one of the 77 BANKING77 intents. It has no "unknown" or out-of-scope class, so off-topic input still returns a confident-looking label. Threshold on the softmax score if you need an abstain path.
  • Domain-bound. Performance outside retail/online banking phrasing is not representative.
  • English only, and tuned to short single-sentence queries — long multi-intent messages are not handled well.
  • Uncased. Casing is discarded, so it cannot use capitalisation as a signal.
  • Inherits the pretraining biases of bert-base-uncased. Do not use it as a sole automated decision-maker on anything consequential to a customer.

Training and evaluation data

PolyAI/banking77, used with its standard split — 10,003 training examples and 3,080 test examples over 77 intent labels. Evaluation metric is weighted F1, which accounts for the mild class imbalance across intents.

Training procedure

Training hyperparameters

  • learning_rate: 5e-05
  • train_batch_size: 32
  • eval_batch_size: 16
  • seed: 42
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: linear
  • num_epochs: 10

Training results

Training Loss Epoch Step Validation Loss F1
3.261 1.0 313 1.0894 0.7969
0.5499 2.0 626 0.4196 0.9103
0.305 3.0 939 0.3403 0.9157
0.1277 4.0 1252 0.3020 0.9251
0.0857 5.0 1565 0.2911 0.9306
0.0347 6.0 1878 0.2865 0.9333
0.0251 7.0 2191 0.2994 0.9362
0.0111 8.0 2504 0.2970 0.9365
0.0075 9.0 2817 0.3102 0.9364
0.0058 10.0 3130 0.3089 0.9362

Validation F1 plateaus around epoch 7 (0.9362) while training loss keeps falling — the last few epochs are mostly overfitting. Epoch 7–8 is the sweet spot if you retrain.

Framework versions

  • Transformers 4.42.4
  • Pytorch 2.3.1+cu121
  • Datasets 2.20.0
  • Tokenizers 0.19.1

How to use

from transformers import AutoTokenizer, AutoModelForSequenceClassification, pipeline

ckpt = 'pistachio7/bert-base-banking77-pt2'
tokenizer = AutoTokenizer.from_pretrained(ckpt)
model = AutoModelForSequenceClassification.from_pretrained(ckpt)

classifier = pipeline('text-classification', tokenizer=tokenizer, model=model)
classifier('What is the base of the exchange rates?')
# Output: [{'label': 'exchange_rate', 'score': 0.9961327314376831}]

To get the full ranked list of intents instead of just the top one:

classifier('My card still has not turned up', top_k=3)
Downloads last month
14
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pistachio7/bert-base-banking77-pt2

Finetuned
(7013)
this model

Dataset used to train pistachio7/bert-base-banking77-pt2

Evaluation results