Instructions to use Chima207/distilbert_goodreads_book_classification with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Chima207/distilbert_goodreads_book_classification with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="Chima207/distilbert_goodreads_book_classification")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("Chima207/distilbert_goodreads_book_classification") model = AutoModelForSequenceClassification.from_pretrained("Chima207/distilbert_goodreads_book_classification", device_map="auto") - Notebooks
- Google Colab
- Kaggle
distilbert_goodreads_book_classification
This model is a fine-tuned version of distilbert-base-uncased on the Goodreads-Books dataset. It achieves the following results on the evaluation set:
- Loss: 2.0030
- Accuracy: 0.5159
- F1 Score: 0.5065
- Precision: 0.5073
- Recall: 0.5159
Model description
This model is a fine-tuned version of distilbert-base-uncased trained for multi-class book genre classification based on textual metadata (book titles and descriptions).
Due to the noisy and highly heterogeneous nature of crowd-sourced Goodreads user tags, 1,003 unique Goodreads shelves were heuristically mapped into a standardized taxonomy of 31 Amazon Kindle categories. To handle the resulting severe class imbalance, class distribution was balanced using a Random Oversampler before fine-tuning. This model represents a single-stage fine-tuning approach directly on the preprocessed Goodreads metadata dataset, achieving an Accuracy of 50.58% and significantly outperforming traditional vector-based baselines (Doc2Vec + Random Forest at 40.84%).
Intended uses & limitations
Automated genre classification of books using metadata or descriptions.
Serving as a baseline for evaluating NLP models on imbalanced, crowd-sourced taxonomy mapping tasks.
Performance is fundamentally constrained by the inherent noise, ambiguity, and subjectivity of crowdsourced user tags (Goodreads shelves).
Predictions are mapped specifically to the 31-class Amazon Kindle taxonomy; inputs outside this category scope may yield inaccurate classifications.
Datasets
- Goodreads-Dataset: Hugging Face Repository (Original: BrightData/Goodreads-Books)
- Amazon-Dataset: Kaggle Amazon Kindle Books Dataset)
Training and evaluation data
- Source: Goodreads book metadata and user-generated shelf tags (noisy genres).
- Target Taxonomy: 31 standardized Amazon Kindle categories (genres).
- Input Features: Concatenated description text made of book title, author and summary.
Data Preprocessing & Pipeline
- Taxonomy Mapping: 1,003 heterogeneous, crowdsourced Goodreads-shelves were heuristically clustered and mapped into a standardized 31-class Amazon Kindle taxonomy to resolve label noise and synonymy.
- Class Imbalance Handling: To address the severe long-tail class distribution inherent in user-generated book data, Random Oversampling was applied to the training split prior to model fine-tuning.
- Splits: The dataset was partitioned into stratified (80/20) train and test splits to evaluate the classification performance.
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
- learning_rate: 2e-05
- train_batch_size: 4
- eval_batch_size: 4
- seed: 42
- gradient_accumulation_steps: 4
- total_train_batch_size: 16
- optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
- lr_scheduler_type: linear
- num_epochs: 2
- mixed_precision_training: Native AMP
Training results
| Training Loss | Epoch | Step | Validation Loss | Accuracy | F1 Score | Precision | Recall |
|---|---|---|---|---|---|---|---|
| 0.4212 | 1.0000 | 8519 | 1.7610 | 0.5058 | 0.4913 | 0.5018 | 0.5058 |
| 0.247 | 1.9999 | 17038 | 2.0030 | 0.5159 | 0.5065 | 0.5073 | 0.5159 |
Framework versions
- Transformers 4.45.2
- Pytorch 2.5.1
- Datasets 4.1.1
- Tokenizers 0.20.1
Academic Context & Citation / Akademischer Kontext
This repository and model were developed as part of a Bachelor's thesis in 2026.
- Title: Classification of Goodreads genres: A methodological comparison of Doc2Vec and DistilBERT
- License: CC BY-NC 4.0 (Free for research, education, and personal use; commercial use prohibited)
Dieses Repository und Modell wurden im Rahmen einer Bachelorarbeit im Jahr 2026 entwickelt.
- Titel: Klassifikation von Goodreads-Genres: Ein methodischer Vergleich von Doc2Vec und DistilBERT
- Lizenz: CC BY-NC 4.0 (Frei für Forschung, Lehre und private Nutzung; kommerzielle Nutzung untersagt)
- Downloads last month
- 54
Model tree for Chima207/distilbert_goodreads_book_classification
Base model
distilbert/distilbert-base-uncased