Flan-T5-Large + prefix-tuning on DialogSum

google/flan-t5-large adapted for dialogue summarization on DialogSum, trained as part of an IIT-D Gen-AI course project comparing four fine-tuning methods under identical conditions.

Method: Prefix-tuning (100 virtual tokens, no prefix projection)

Code, evaluation harness and the other three models: https://github.com/dipika-s/iitd-genai

Known issue

This adapter scores identically to the untuned base model on all 13 evaluation metrics, to four decimal places. Training loss did decrease (1.117 train / 0.948 eval), so the prefix learned something — but it had no effect on generated text, which points at the generation path failing to apply the prefix rather than at a property of prefix-tuning.

The adapter is published as-is for completeness and reproducibility. Do not cite the numbers below as evidence about prefix-tuning without first re-checking generation.

Usage

from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
from peft import PeftModel

base = AutoModelForSeq2SeqLM.from_pretrained("google/flan-t5-large")
model = PeftModel.from_pretrained(base, "daggar/flan-t5-dialogsum-prefix")
tokenizer = AutoTokenizer.from_pretrained("daggar/flan-t5-dialogsum-prefix")

dialogue = "#Person1#: Hi, how was your weekend?\n#Person2#: Great, I went hiking."
inputs = tokenizer("Summarize the following dialogue:\n" + dialogue,
                   return_tensors="pt", max_length=512, truncation=True)
print(tokenizer.decode(model.generate(**inputs, max_new_tokens=128)[0],
                       skip_special_tokens=True))

Inputs must use the training prompt — "Summarize the following dialogue:\n" followed by the dialogue — truncated to 512 tokens. Summaries were trained at up to 128 tokens.

Training

Base model google/flan-t5-large
Dataset knkarthick/dialogsum — 12,460 train / 500 validation / 1,500 test
Epochs 3
Optimizer steps 4,674
Trainable parameters 4,915,200 of 788,065,280 (0.6238%)
Peak GPU memory 8.85 GB
Training time 38.8 min on an AWS g5.2xlarge (A10G 24GB)
Final training loss 1.1167
Final eval loss 0.9478

Evaluation

Measured on the full 1,500-example DialogSum test split, against the untuned base model.

Metric Base This model
BLEU 5.7129 5.7129
ROUGE-1 37.3183 37.3183
ROUGE-2 15.1221 15.1221
ROUGE-L 30.8519 30.8519
METEOR 22.2824 22.2824
GLEU 9.7314 9.7314
BERTScore 89.3701 89.3701
CoSIM 34.366 34.366
Repetition rate 2.2415 2.2415
Flesch reading ease 75.0617 75.0617
Toxicity 0.7889 0.7889
Novelty 71.759 71.759
Diversity 22.809 22.809

All four methods

Method Trainable params Peak GPU ROUGE-L BERTScore
Full FT 783M (100%) 16.11 GB 37.42 91.64
LoRA 4.7M (0.60%) 4.51 GB 37.40 91.66
QLoRA 4.7M (0.60%) 2.52 GB 37.01 91.58
Prefix 4.9M (0.62%) 8.85 GB 30.85* 89.37*

* see the known issue on the prefix model card.

Limitations

Trained only on DialogSum, which is two-speaker English conversation transcripts using #Person1# / #Person2# speaker tags. Summaries of longer, multi-party, domain-specific or non-English dialogue will be unreliable. Dialogues over 512 tokens are truncated, so content late in a long conversation may be dropped. The model inherits any biases present in google/flan-t5-large and in DialogSum, and summaries can contain details not supported by the source dialogue.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for daggar/flan-t5-dialogsum-prefix

Adapter
(234)
this model

Dataset used to train daggar/flan-t5-dialogsum-prefix