hongccccccc's picture
Upload README.md with huggingface_hub
2b50b70 verified
|
Raw
History Blame Contribute Delete
1.99 kB
metadata
language: en
license: apache-2.0
base_model: roberta-base
pipeline_tag: text-classification
library_name: transformers
tags:
  - scientific-text
  - abstract-sections
widget:
  - text: >-
      We conclude that early intervention significantly improves patient
      outcomes.
    example_title: Conclusion sentence
  - text: Participants were randomly assigned to treatment and control groups.
    example_title: Methods sentence
  - text: The aim of this study was to evaluate the efficacy of the new vaccine.
    example_title: Objective sentence

RoBERTa Abstract-Section Classifier

roberta-base fine-tuned to classify sentences from scientific abstracts into five rhetorical sections: BACKGROUND, CONCLUSIONS, METHODS, OBJECTIVE, RESULTS.

Labels

id label example
0 BACKGROUND "Diabetes is a growing public health concern…"
1 CONCLUSIONS "We conclude that early intervention improves outcomes."
2 METHODS "Participants were randomly assigned to two groups…"
3 OBJECTIVE "The aim of this study was to evaluate…"
4 RESULTS "The treatment group showed a 40% reduction (p < 0.001)."

How to use

from transformers import pipeline

clf = pipeline("text-classification", model="hongccccccc/roberta-abstract-section-classifier")
clf("The treatment group showed a 40% reduction in mortality compared with placebo.")
# [{'label': 'RESULTS', 'score': 0.99}]

Training

  • Base model: RoBERTa (Liu et al., 2019; RobertaForSequenceClassification, single-label, 5 classes)
  • Data: 200,000 paper abstracts from PubMed (Canese and Weis, 2013), self-labeled with the five section categories
  • Fine-tuned: January 2023, transformers 4.12.5 (original training_args.bin included)

Evaluation

F1 score of 0.92 on a held-out 10% sample (details in Wright et al., 2022).