jimjunior's picture
add training notebook link
3af9e42 verified
|
Raw
History Blame Contribute Delete
3.44 kB
---
license: mit
datasets:
- cornell-movie-review-data/rotten_tomatoes
language:
- en
base_model:
- distilbert/distilbert-base-uncased
pipeline_tag: text-classification
---
# Rotten Tomatoes Sentiment DistilBERT
This model is a fine-tuned version of `distilbert-base-uncased` for binary sentiment classification on movie review snippets from the Rotten Tomatoes dataset. The model predicts whether a given movie review sentence expresses a **positive** or **negative** sentiment.
## Model Details
### Base Model
The base model is `distilbert-base-uncased`, a smaller and faster distilled version of BERT. It is designed to retain much of BERT’s language understanding ability while being more lightweight.
### Fine-Tuned Task
This model was fine-tuned for **text classification**, specifically **binary sentiment analysis**.
The two output labels are:
| Label ID | Label |
| -------- | -------- |
| 0 | NEGATIVE |
| 1 | POSITIVE |
## Dataset
This model was fine-tuned using the `cornell-movie-review-data/rotten_tomatoes` dataset from Hugging Face Datasets.
The dataset contains short movie review sentences labeled as either positive or negative.
Example inputs:
```text
"a smart, funny and touching film"
```
```text
"the story is dull and lifeless"
```
## Intended Use
This model can be used for simple sentiment classification of short English movie-review-style text.
Example use cases include:
* Classifying movie reviews as positive or negative
* Learning and demonstrating Hugging Face Transformers fine-tuning
* Experimenting with the `Trainer` API
* Building a basic sentiment analysis demo
## Limitations
This model was trained on movie review data, so it may perform best on text that resembles short movie reviews.
It may not perform well on:
* Long documents
* Non-English text
* Sarcasm or highly ambiguous sentiment
* Domain-specific sentiment outside movie reviews
* Reviews with mixed opinions
For example, a sentence like:
```text
"The acting was excellent, but the story was slow and forgettable."
```
may be difficult because it contains both positive and negative sentiment.
## Training Procedure
The model was fine-tuned using the Hugging Face Transformers `Trainer` API.
### Training Configuration
```python
TrainingArguments(
output_dir="./rt-sentiment-distilbert",
learning_rate=2e-5,
per_device_train_batch_size=16,
per_device_eval_batch_size=16,
num_train_epochs=3,
weight_decay=0.01,
eval_strategy="epoch",
save_strategy="epoch",
load_best_model_at_end=True,
metric_for_best_model="accuracy",
report_to="none",
)
```
> FULL TRAINING AND FINETUNING CODE CAN BE ACCESSED IN THIS COLAB NODE BOOK: https://colab.research.google.com/drive/12cwKhvX9LAq39LGk-WbxzchYPla_yrA4?usp=sharing
## Usage
You can use this model with the Hugging Face `pipeline` API:
```python
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="jimjunior/rt-sentiment-distilbert"
)
result = classifier("The movie was touching, funny, and beautifully acted.")
print(result)
```
Example output:
```python
[{"label": "POSITIVE", "score": 0.98}]
```
You can also classify negative text:
```python
classifier("The story was boring and the acting was terrible.")
```
Example output:
```python
[{"label": "NEGATIVE", "score": 0.97}]
```
## Model Output
The model returns one of two labels:
* `NEGATIVE`
* `POSITIVE`
Each prediction also includes a confidence score.