on-task-bert / README.md
paiheng's picture
Upload on-task BERT classifier
d7af4c0 verified
|
Raw
History Blame Contribute Delete
1.12 kB
metadata
language:
  - en
license: mit
library_name: transformers
pipeline_tag: text-classification
base_model: google-bert/bert-base-uncased
tags:
  - education

On-Task BERT

This model classifies a student's classroom utterance as off-task (0) or on-task (1). It is a BERT-base-uncased sequence classifier fine-tuned on English student utterances from mathematics classroom transcripts.

Evaluation

The final model was selected using five-fold cross-validation on the training and validation data, retrained on 1,878 examples, and evaluated once on a held-out test set of 470 examples.

Metric Result
Accuracy 0.900
Macro F1 0.813
On-task precision 0.935
On-task recall 0.947
On-task F1 0.941

The test confusion counts were 372 true positives, 26 false positives, 21 false negatives, and 51 true negatives. On the same test split, a math-vocabulary baseline achieved 0.849 accuracy and 0.559 macro F1.

Training details

  • Base model: bert-base-uncased
  • Maximum sequence length: 256
  • Epochs: 5
  • Learning rate: 3e-5
  • Input: student utterance only