distilbert_amazon_goodreads_book_classification

This model was trained from scratch on the Goodreads-Books dataset. It achieves the following results on the evaluation set:

  • Loss: 1.7274
  • Accuracy: 0.5135
  • F1 Score: 0.4989
  • Precision: 0.5060
  • Recall: 0.5135

Model description

This model evaluates a two-stage sequential transfer learning approach (Amazon to Goodreads). Starting from distilbert-base-uncased, the model was first fine-tuned on structured Amazon book metadata and subsequently fine-tuned on the heuristically mapped, oversampled Goodreads dataset across 31 categories. While this sequential adaptation yields a peak Accuracy of 51.38% (a marginal +0.80% gain over single-stage Goodreads fine-tuning), the empirical results demonstrate diminishing returns. The additional computational overhead of two-stage fine-tuning is rarely justified for real-world deployment compared to direct single-stage adaptation.

Intended uses & limitations

  • Research and analysis of domain adaptation mechanisms between structured e-commerce data and noisy user-generated web content.

  • High-accuracy deployment scenarios where every marginal percentage increase in classification performance is critical regardless of compute budget.

  • Requires substantial training time and computational resources for a minimal performance gain over single-stage fine-tuning.

  • Inherits the classification domain boundaries of the 31 target Amazon categories.

Datasets

Training and evaluation data

  1. Stage 1 (Pre-training / Source Domain):
    • Source: Structured Amazon Kindle metadata.
    • Purpose: Learn clean domain representations and sentence semantics on structured book descriptions.
  2. Stage 2 (Target Adaptation / Target Domain):
    • Source: Preprocessed Goodreads book metadata.
    • Preprocessing: 1,003 crowdsourced Goodreads shelves heuristically mapped into the 31 Amazon Kindle categories and balanced using Random Oversampling.
    • Purpose: Adapt the pre-trained weights to the noisy, community-driven text characteristics of Goodreads entries.

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 2e-05
  • train_batch_size: 4
  • eval_batch_size: 2
  • seed: 42
  • gradient_accumulation_steps: 4
  • total_train_batch_size: 16
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: linear
  • num_epochs: 2
  • mixed_precision_training: Native AMP

Training results

Training Loss Epoch Step Validation Loss Accuracy F1 Score Precision Recall
0.414 1.0000 8519 1.7274 0.5135 0.4989 0.5060 0.5135
0.237 1.9999 17038 2.0405 0.5210 0.5118 0.5146 0.5210

Framework versions

  • Transformers 4.45.2
  • Pytorch 2.5.1
  • Datasets 4.1.1
  • Tokenizers 0.20.1

Academic Context & Citation / Akademischer Kontext

This repository and model were developed as part of a Bachelor's thesis in 2026.

  • Title: Classification of Goodreads genres: A methodological comparison of Doc2Vec and DistilBERT
  • License: CC BY-NC 4.0 (Free for research, education, and personal use; commercial use prohibited)

Dieses Repository und Modell wurden im Rahmen einer Bachelorarbeit im Jahr 2026 entwickelt.

  • Titel: Klassifikation von Goodreads-Genres: Ein methodischer Vergleich von Doc2Vec und DistilBERT
  • Lizenz: CC BY-NC 4.0 (Frei für Forschung, Lehre und private Nutzung; kommerzielle Nutzung untersagt)
Downloads last month
36
Safetensors
Model size
67M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Chima207/distilbert_amazon_goodreads_book_classification

Dataset used to train Chima207/distilbert_amazon_goodreads_book_classification