Flan-T5-Large fully fine-tuned on DialogSum

google/flan-t5-large adapted for dialogue summarization on DialogSum, trained as part of an IIT-D Gen-AI course project comparing four fine-tuning methods under identical conditions.

Method: Full fine-tuning (all 783M parameters updated)

Code, evaluation harness and the other three models: https://github.com/dipika-s/iitd-genai

Usage

from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

model = AutoModelForSeq2SeqLM.from_pretrained("daggar/flan-t5-dialogsum-full-ft")
tokenizer = AutoTokenizer.from_pretrained("daggar/flan-t5-dialogsum-full-ft")

dialogue = "#Person1#: Hi, how was your weekend?\n#Person2#: Great, I went hiking."
inputs = tokenizer("Summarize the following dialogue:\n" + dialogue,
                   return_tensors="pt", max_length=512, truncation=True)
print(tokenizer.decode(model.generate(**inputs, max_new_tokens=128)[0],
                       skip_special_tokens=True))

Inputs must use the training prompt — "Summarize the following dialogue:\n" followed by the dialogue — truncated to 512 tokens. Summaries were trained at up to 128 tokens.

Training

Base model google/flan-t5-large
Dataset knkarthick/dialogsum — 12,460 train / 500 validation / 1,500 test
Epochs 1
Optimizer steps 1,558
Trainable parameters 783,150,080 of 783,150,080 (100.0%)
Peak GPU memory 16.11 GB
Training time 36.4 min on an AWS g5.2xlarge (A10G 24GB)
Final training loss 1.0257
Final eval loss 0.8793

Evaluation

Measured on the full 1,500-example DialogSum test split, against the untuned base model.

Metric Base This model
BLEU 5.7129 10.1951
ROUGE-1 37.3183 45.5045
ROUGE-2 15.1221 19.8113
ROUGE-L 30.8519 37.4236
METEOR 22.2824 34.2428
GLEU 9.7314 16.0062
BERTScore 89.3701 91.6383
CoSIM 34.366 41.4397
Repetition rate 2.2415 1.7513
Flesch reading ease 75.0617 66.986
Toxicity 0.7889 0.7124
Novelty 71.759 60.3228
Diversity 22.809 22.2354

All four methods

Method Trainable params Peak GPU ROUGE-L BERTScore
Full FT 783M (100%) 16.11 GB 37.42 91.64
LoRA 4.7M (0.60%) 4.51 GB 37.40 91.66
QLoRA 4.7M (0.60%) 2.52 GB 37.01 91.58
Prefix 4.9M (0.62%) 8.85 GB 30.85* 89.37*

* see the known issue on the prefix model card.

Limitations

Trained only on DialogSum, which is two-speaker English conversation transcripts using #Person1# / #Person2# speaker tags. Summaries of longer, multi-party, domain-specific or non-English dialogue will be unreliable. Dialogues over 512 tokens are truncated, so content late in a long conversation may be dropped. The model inherits any biases present in google/flan-t5-large and in DialogSum, and summaries can contain details not supported by the source dialogue.

Downloads last month
6
Safetensors
Model size
0.8B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for daggar/flan-t5-dialogsum-full-ft

Finetuned
(214)
this model

Dataset used to train daggar/flan-t5-dialogsum-full-ft