Instructions to use Chima207/distilbert_amazon_goodreads_book_classification with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Chima207/distilbert_amazon_goodreads_book_classification with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="Chima207/distilbert_amazon_goodreads_book_classification")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("Chima207/distilbert_amazon_goodreads_book_classification") model = AutoModelForSequenceClassification.from_pretrained("Chima207/distilbert_amazon_goodreads_book_classification", device_map="auto") - Notebooks
- Google Colab
- Kaggle
distilbert_amazon_goodreads_book_classification
This model was trained from scratch on the Goodreads-Books dataset. It achieves the following results on the evaluation set:
- Loss: 1.7274
- Accuracy: 0.5135
- F1 Score: 0.4989
- Precision: 0.5060
- Recall: 0.5135
Model description
This model evaluates a two-stage sequential transfer learning approach (Amazon to Goodreads). Starting from distilbert-base-uncased, the model was first fine-tuned on structured Amazon book metadata and subsequently fine-tuned on the heuristically mapped, oversampled Goodreads dataset across 31 categories. While this sequential adaptation yields a peak Accuracy of 51.38% (a marginal +0.80% gain over single-stage Goodreads fine-tuning), the empirical results demonstrate diminishing returns. The additional computational overhead of two-stage fine-tuning is rarely justified for real-world deployment compared to direct single-stage adaptation.
Intended uses & limitations
Research and analysis of domain adaptation mechanisms between structured e-commerce data and noisy user-generated web content.
High-accuracy deployment scenarios where every marginal percentage increase in classification performance is critical regardless of compute budget.
Requires substantial training time and computational resources for a minimal performance gain over single-stage fine-tuning.
Inherits the classification domain boundaries of the 31 target Amazon categories.
Datasets
- Goodreads-Dataset: Hugging Face Repository (Original: BrightData/Goodreads-Books)
- Amazon-Dataset: Kaggle Amazon Kindle Books Dataset)
Training and evaluation data
- Stage 1 (Pre-training / Source Domain):
- Source: Structured Amazon Kindle metadata.
- Purpose: Learn clean domain representations and sentence semantics on structured book descriptions.
- Stage 2 (Target Adaptation / Target Domain):
- Source: Preprocessed Goodreads book metadata.
- Preprocessing: 1,003 crowdsourced Goodreads shelves heuristically mapped into the 31 Amazon Kindle categories and balanced using Random Oversampling.
- Purpose: Adapt the pre-trained weights to the noisy, community-driven text characteristics of Goodreads entries.
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
- learning_rate: 2e-05
- train_batch_size: 4
- eval_batch_size: 2
- seed: 42
- gradient_accumulation_steps: 4
- total_train_batch_size: 16
- optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
- lr_scheduler_type: linear
- num_epochs: 2
- mixed_precision_training: Native AMP
Training results
| Training Loss | Epoch | Step | Validation Loss | Accuracy | F1 Score | Precision | Recall |
|---|---|---|---|---|---|---|---|
| 0.414 | 1.0000 | 8519 | 1.7274 | 0.5135 | 0.4989 | 0.5060 | 0.5135 |
| 0.237 | 1.9999 | 17038 | 2.0405 | 0.5210 | 0.5118 | 0.5146 | 0.5210 |
Framework versions
- Transformers 4.45.2
- Pytorch 2.5.1
- Datasets 4.1.1
- Tokenizers 0.20.1
Academic Context & Citation / Akademischer Kontext
This repository and model were developed as part of a Bachelor's thesis in 2026.
- Title: Classification of Goodreads genres: A methodological comparison of Doc2Vec and DistilBERT
- License: CC BY-NC 4.0 (Free for research, education, and personal use; commercial use prohibited)
Dieses Repository und Modell wurden im Rahmen einer Bachelorarbeit im Jahr 2026 entwickelt.
- Titel: Klassifikation von Goodreads-Genres: Ein methodischer Vergleich von Doc2Vec und DistilBERT
- Lizenz: CC BY-NC 4.0 (Frei für Forschung, Lehre und private Nutzung; kommerzielle Nutzung untersagt)
- Downloads last month
- 36
Model tree for Chima207/distilbert_amazon_goodreads_book_classification
Base model
distilbert/distilbert-base-uncased