distilbert_goodreads_book_classification

This model is a fine-tuned version of distilbert-base-uncased on the Goodreads-Books dataset. It achieves the following results on the evaluation set:

  • Loss: 2.0030
  • Accuracy: 0.5159
  • F1 Score: 0.5065
  • Precision: 0.5073
  • Recall: 0.5159

Model description

This model is a fine-tuned version of distilbert-base-uncased trained for multi-class book genre classification based on textual metadata (book titles and descriptions).

Due to the noisy and highly heterogeneous nature of crowd-sourced Goodreads user tags, 1,003 unique Goodreads shelves were heuristically mapped into a standardized taxonomy of 31 Amazon Kindle categories. To handle the resulting severe class imbalance, class distribution was balanced using a Random Oversampler before fine-tuning. This model represents a single-stage fine-tuning approach directly on the preprocessed Goodreads metadata dataset, achieving an Accuracy of 50.58% and significantly outperforming traditional vector-based baselines (Doc2Vec + Random Forest at 40.84%).

Intended uses & limitations

  • Automated genre classification of books using metadata or descriptions.

  • Serving as a baseline for evaluating NLP models on imbalanced, crowd-sourced taxonomy mapping tasks.

  • Performance is fundamentally constrained by the inherent noise, ambiguity, and subjectivity of crowdsourced user tags (Goodreads shelves).

  • Predictions are mapped specifically to the 31-class Amazon Kindle taxonomy; inputs outside this category scope may yield inaccurate classifications.

Datasets

Training and evaluation data

  • Source: Goodreads book metadata and user-generated shelf tags (noisy genres).
  • Target Taxonomy: 31 standardized Amazon Kindle categories (genres).
  • Input Features: Concatenated description text made of book title, author and summary.

Data Preprocessing & Pipeline

  1. Taxonomy Mapping: 1,003 heterogeneous, crowdsourced Goodreads-shelves were heuristically clustered and mapped into a standardized 31-class Amazon Kindle taxonomy to resolve label noise and synonymy.
  2. Class Imbalance Handling: To address the severe long-tail class distribution inherent in user-generated book data, Random Oversampling was applied to the training split prior to model fine-tuning.
  3. Splits: The dataset was partitioned into stratified (80/20) train and test splits to evaluate the classification performance.

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 2e-05
  • train_batch_size: 4
  • eval_batch_size: 4
  • seed: 42
  • gradient_accumulation_steps: 4
  • total_train_batch_size: 16
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: linear
  • num_epochs: 2
  • mixed_precision_training: Native AMP

Training results

Training Loss Epoch Step Validation Loss Accuracy F1 Score Precision Recall
0.4212 1.0000 8519 1.7610 0.5058 0.4913 0.5018 0.5058
0.247 1.9999 17038 2.0030 0.5159 0.5065 0.5073 0.5159

Framework versions

  • Transformers 4.45.2
  • Pytorch 2.5.1
  • Datasets 4.1.1
  • Tokenizers 0.20.1

Academic Context & Citation / Akademischer Kontext

This repository and model were developed as part of a Bachelor's thesis in 2026.

  • Title: Classification of Goodreads genres: A methodological comparison of Doc2Vec and DistilBERT
  • License: CC BY-NC 4.0 (Free for research, education, and personal use; commercial use prohibited)

Dieses Repository und Modell wurden im Rahmen einer Bachelorarbeit im Jahr 2026 entwickelt.

  • Titel: Klassifikation von Goodreads-Genres: Ein methodischer Vergleich von Doc2Vec und DistilBERT
  • Lizenz: CC BY-NC 4.0 (Frei für Forschung, Lehre und private Nutzung; kommerzielle Nutzung untersagt)
Downloads last month
54
Safetensors
Model size
67M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Chima207/distilbert_goodreads_book_classification

Finetuned
(12625)
this model

Dataset used to train Chima207/distilbert_goodreads_book_classification