DistilGPT2 fine-tuned on ELI5 Answers

Model Description

This model is a fine-tuned version of distilbert/distilgpt2 for causal language modeling and text generation.

It was fine-tuned on a subset of the dany0407/eli5_category dataset.

The model was created as a practical exercise while following the Hugging Face LLM Course.

  • Developed by: Driw0x
  • Model type: DistilGPT2
  • Language: English
  • Base model: distilbert/distilgpt2
  • Task: Causal language modeling / text generation

Training and Evaluation Data

The training notebook loads the first 5,000 examples from the ELI5 training split:

eli5 = load_dataset(
    "dany0407/eli5_category",
    split="train[:5000]"
)

These examples are then divided into training and evaluation subsets:

eli5 = eli5.train_test_split(
    test_size=0.2
)

This results in approximately:

  • 80% training data;
  • 20% evaluation data.

No explicit random seed was passed directly to train_test_split() in the notebook.

Dataset Detail

The tutorial introduction describes this experiment as using the r/askscience subset of ELI5.

However, the code that was actually executed does not apply an explicit category or subreddit filter.

The training run therefore uses the first 5,000 examples returned by:

dany0407/eli5_category

This model card documents the executed code.

Preprocessing

The nested ELI5 dataset structure is first flattened.

The model is trained using the answer text.

For every example, all available answers are concatenated:

def preprocess_function(examples):
    return tokenizer(
        [
            " ".join(x)
            for x in examples["answers.text"]
        ]
    )

The tokenized texts are then concatenated and divided into fixed-size blocks.

block_size = 128

The grouping function creates sequences of 128 tokens:

def group_texts(examples):
    concatenated_examples = {
        k: sum(examples[k], [])
        for k in examples.keys()
    }

    total_length = len(
        concatenated_examples[
            list(examples.keys())[0]
        ]
    )

    if total_length >= block_size:
        total_length = (
            total_length // block_size
        ) * block_size

    result = {
        k: [
            t[i:i + block_size]
            for i in range(
                0,
                total_length,
                block_size
            )
        ]
        for k, t in concatenated_examples.items()
    }

    result["labels"] = result["input_ids"].copy()

    return result

Remaining tokens that do not fill a complete block are discarded.

For causal language modeling, the input token IDs are also used as labels.

The GPT-2 EOS token is reused as the padding token:

tokenizer.pad_token = tokenizer.eos_token

The language-modeling collator is configured with:

DataCollatorForLanguageModeling(
    tokenizer=tokenizer,
    mlm=False
)

This means the model is trained with causal language modeling rather than masked language modeling.

Training Procedure

The model was fine-tuned using the Hugging Face Trainer API.

Training Hyperparameters

Hyperparameter Value
Learning rate 2e-5
Train batch size 8
Evaluation batch size 8
Number of epochs 3
Weight decay 0.01
Seed 42
Evaluation strategy epoch
LR scheduler linear
Token block size 128

The optimizer used during training was AdamW (ADAMW_TORCH_FUSED) with:

  • betas=(0.9, 0.999)
  • epsilon=1e-8

Training Results

Epoch Training Loss Validation Loss
1 3.924143 3.811588
2 3.822786 3.802098
3 3.779634 3.801003

The complete training run reported:

  • Training loss: 3.844863
  • Training steps: 3966
  • Epochs: 3

Evaluation

After training, the model was explicitly evaluated with:

eval_results = trainer.evaluate()

Perplexity was computed from the evaluation loss using:

import math

perplexity = math.exp(
    eval_results["eval_loss"]
)

The resulting evaluation metrics were:

  • Validation loss: 3.801003
  • Perplexity: 44.75

Perplexity measures how well the language model predicts the evaluation data. It does not measure factual correctness or the quality of generated answers.

Usage

Pipeline

from transformers import pipeline

generator = pipeline(
    "text-generation",
    model="Driw0x/my_awesome_eli5_clm-model"
)

prompt = (
    "Somatic hypermutation allows "
    "the immune system to"
)

result = generator(prompt)

print(result[0]["generated_text"])

Direct Generation

from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained(
    "Driw0x/my_awesome_eli5_clm-model"
)

model = AutoModelForCausalLM.from_pretrained(
    "Driw0x/my_awesome_eli5_clm-model"
)

prompt = (
    "Somatic hypermutation allows "
    "the immune system to"
)

inputs = tokenizer(
    prompt,
    return_tensors="pt"
).input_ids

outputs = model.generate(
    inputs,
    max_new_tokens=100,
    do_sample=True,
    top_k=50,
    top_p=0.95
)

generated_text = tokenizer.batch_decode(
    outputs,
    skip_special_tokens=True
)[0]

print(generated_text)

Intended Uses

This model is primarily intended for:

  • learning causal language modeling;
  • experimenting with DistilGPT2 fine-tuning;
  • learning text preprocessing for language-model training;
  • experimenting with Hugging Face Trainer;
  • evaluating language models with perplexity;
  • experimenting with text generation.

Limitations

This is a small educational fine-tune and should not be treated as a general-purpose assistant.

Important limitations include:

  • only the first 5,000 examples of the selected dataset were used;
  • only answer text is used during language-model training;
  • questions and answers are not explicitly paired as a supervised question-answering task;
  • training sequences are limited to blocks of 128 tokens;
  • incomplete final token blocks are discarded;
  • the train/evaluation split was created without an explicit random seed passed to train_test_split();
  • generated text may be repetitive, inaccurate, unsupported, or nonsensical;
  • perplexity does not measure factual correctness;
  • biases or problematic content present in the base model or dataset may be inherited;
  • the model has not undergone instruction tuning, preference optimization, or dedicated safety alignment.

Generated information should not be assumed to be factually correct.

Framework Versions

  • Transformers 5.17.0
  • PyTorch 2.11.0+cu130
  • Datasets 4.8.5
  • Tokenizers 0.23.2

Training Source

The complete training procedure is available in the following repository:

Driw0x/hf-ai-courses

Notebook:

llm-course/1-transformer-models/notebooks/language_modeling.ipynb

Downloads last month
579
Safetensors
Model size
81.9M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Driw0x/my_awesome_eli5_clm-model

Finetuned
(1577)
this model

Dataset used to train Driw0x/my_awesome_eli5_clm-model