YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
BERT Fine-Tuning for IMDb Sentiment Classification
A fine-tuned BERT Base Uncased model for binary sentiment classification on the IMDb Movie Reviews dataset.
This project demonstrates the complete fine-tuning workflow using the Hugging Face ecosystem, from dataset preprocessing and tokenization to model training, evaluation, inference, and deployment.
Model Details
- Base Model:
bert-base-uncased - Task: Binary Sentiment Classification
- Dataset: IMDb Movie Reviews
- Framework: Hugging Face Transformers
- Training Framework: Trainer API
- Language: English
Training Pipeline
The model was trained using the following workflow:
- Dataset loading using Hugging Face Datasets
- Tokenization with
AutoTokenizer - Fine-tuning using
AutoModelForSequenceClassification - Evaluation with Accuracy metric
- Mixed precision (FP16) training when CUDA is available
- Model exported using SafeTensors
Performance
The fine-tuned model learns to classify movie reviews into:
- LABEL_0 → Negative
- LABEL_1 → Positive
Usage
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="YOUR_USERNAME/BERT-Fine-Tuning"
)
classifier("This movie was absolutely amazing!")
Example output:
[
{
"label": "LABEL_1",
"score": 0.998
}
]
Repository Contents
- Fine-tuned model weights
- Tokenizer files
- Configuration files
- SafeTensors checkpoint
The complete training notebook, source code, and documentation are available in the accompanying GitHub repository.
Future Improvements
- LoRA / PEFT fine-tuning
- Multi-class sentiment classification
- Hyperparameter optimization
- Model quantization
- ONNX and TensorRT deployment
- Production inference benchmarking
License
This project is released for educational and research purposes.
Built with ❤️ by the author.
- Downloads last month
- 41