DistilBERT fine-tuned on SQuAD for Extractive Question Answering

Model Description

This model is a fine-tuned version of distilbert/distilbert-base-uncased for extractive question answering.

Given a question and a context passage, the model predicts the start and end token positions of an answer span inside the context.

It was fine-tuned on a 5,000-example subset of SQuAD.

The model was created as a practical exercise while following the Hugging Face LLM Course.

  • Developed by: Driw0x
  • Model type: DistilBERT
  • Language: English
  • Base model: distilbert/distilbert-base-uncased
  • Task: Extractive question answering

Training and Evaluation Data

The model was fine-tuned on:

rajpurkar/squad

The notebook loads the first 5,000 examples from the SQuAD training split and creates a new 80/20 train/evaluation split:

squad = load_dataset("rajpurkar/squad", split="train[:5000]")
squad = squad.train_test_split(test_size=0.2)

This produces:

  • 4,000 examples for training;
  • 1,000 examples for evaluation.

The official SQuAD validation split is not used in this experiment. The generated evaluation split participates in model development and should not be considered an untouched final benchmark.

Preprocessing

Questions are stripped and tokenized together with their contexts.

The preprocessing configuration uses:

  • maximum sequence length: 384;
  • truncation="only_second" so only the context is truncated;
  • return_offsets_mapping=True to map character-level answers to token positions;
  • padding="max_length";
  • start and end token positions as supervision targets.

If an answer is outside the retained context after truncation, its start and end positions are assigned to token index 0 in the course preprocessing function.

Training Procedure

The model was fine-tuned with the Hugging Face Trainer API.

Training Hyperparameters

Hyperparameter Value
Learning rate 2e-5
Train batch size 16
Evaluation batch size 16
Number of epochs 3
Weight decay 0.01
Seed 42
Evaluation strategy epoch
LR scheduler linear
Optimizer AdamW (ADAMW_TORCH_FUSED)

The training notebook does not compute SQuAD Exact Match or F1. Evaluation therefore reports loss only.

Training Results

Epoch Training Loss Validation Loss
1 — 2.5810
2 2.9258 1.8591
3 2.9258 1.7391

Final reported evaluation result:

  • Validation loss: 1.7391

Usage

import torch
from transformers import AutoTokenizer, AutoModelForQuestionAnswering

model_id = "Driw0x/my_awesome_qa_model"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForQuestionAnswering.from_pretrained(model_id)

question = "How many programming languages does BLOOM support?"
context = (
    "BLOOM has 176 billion parameters and can generate text in 46 natural "
    "languages and 13 programming languages."
)

inputs = tokenizer(question, context, return_tensors="pt")

with torch.no_grad():
    outputs = model(**inputs)

start = outputs.start_logits.argmax()
end = outputs.end_logits.argmax()

answer_tokens = inputs["input_ids"][0, start : end + 1]
answer = tokenizer.decode(answer_tokens, skip_special_tokens=True)
print(answer)

Intended Uses

This model is primarily intended for:

  • learning extractive question answering with Transformers;
  • experimenting with DistilBERT fine-tuning;
  • learning answer-span alignment with tokenizer offset mappings;
  • reproducing a Hugging Face question-answering workflow.

Limitations

Important limitations include:

  • training used only the first 5,000 examples of the SQuAD training split;
  • the official SQuAD validation split was not used;
  • no Exact Match or F1 metric was computed in the training notebook;
  • the model can only extract an answer span from the supplied context and cannot generate an answer that is absent from it;
  • contexts longer than the configured maximum length are truncated without a sliding-window evaluation strategy in this course workflow;
  • biases and errors inherited from DistilBERT and SQuAD may remain.

This model was created as a course exercise and has not been validated for production or high-stakes question answering.

Framework Versions

  • Transformers 5.17.0
  • PyTorch 2.11.0+cu130
  • Datasets 4.8.5
  • Tokenizers 0.23.2

Training Source

The complete training procedure is available in:

Driw0x/hf-ai-courses

Notebook:

llm-course/1-transformer-models/notebooks/question_answering.ipynb

Downloads last month
35
Safetensors
Model size
66.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Driw0x/my_awesome_qa_model

Finetuned
(12664)
this model

Dataset used to train Driw0x/my_awesome_qa_model