Flan-T5-Large + LoRA on DialogSum

google/flan-t5-large adapted for dialogue summarization on DialogSum, trained as part of an IIT-D Gen-AI course project comparing four fine-tuning methods under identical conditions.

Method: LoRA (r=16, alpha=32, dropout=0.05, target modules q & v)

Code, evaluation harness and the other three models: https://github.com/dipika-s/iitd-genai

Usage

from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
from peft import PeftModel

base = AutoModelForSeq2SeqLM.from_pretrained("google/flan-t5-large")
model = PeftModel.from_pretrained(base, "daggar/flan-t5-dialogsum-lora")
tokenizer = AutoTokenizer.from_pretrained("daggar/flan-t5-dialogsum-lora")

dialogue = "#Person1#: Hi, how was your weekend?\n#Person2#: Great, I went hiking."
inputs = tokenizer("Summarize the following dialogue:\n" + dialogue,
                   return_tensors="pt", max_length=512, truncation=True)
print(tokenizer.decode(model.generate(**inputs, max_new_tokens=128)[0],
                       skip_special_tokens=True))

Inputs must use the training prompt — "Summarize the following dialogue:\n" followed by the dialogue — truncated to 512 tokens. Summaries were trained at up to 128 tokens.

Training

Base model google/flan-t5-large
Dataset knkarthick/dialogsum — 12,460 train / 500 validation / 1,500 test
Epochs 3
Optimizer steps 4,674
Trainable parameters 4,718,592 of 787,868,672 (0.5989%)
Peak GPU memory 4.51 GB
Training time 94.7 min on an AWS g5.2xlarge (A10G 24GB)
Final training loss 1.0294
Final eval loss 0.8958

Evaluation

Measured on the full 1,500-example DialogSum test split, against the untuned base model.

Metric Base This model
BLEU 5.7129 9.8709
ROUGE-1 37.3183 45.5048
ROUGE-2 15.1221 19.8054
ROUGE-L 30.8519 37.4039
METEOR 22.2824 34.3109
GLEU 9.7314 15.7709
BERTScore 89.3701 91.6589
CoSIM 34.366 41.0212
Repetition rate 2.2415 1.6046
Flesch reading ease 75.0617 67.1588
Toxicity 0.7889 0.6659
Novelty 71.759 60.7914
Diversity 22.809 22.2496

All four methods

Method Trainable params Peak GPU ROUGE-L BERTScore
Full FT 783M (100%) 16.11 GB 37.42 91.64
LoRA 4.7M (0.60%) 4.51 GB 37.40 91.66
QLoRA 4.7M (0.60%) 2.52 GB 37.01 91.58
Prefix 4.9M (0.62%) 8.85 GB 30.85* 89.37*

* see the known issue on the prefix model card.

Limitations

Trained only on DialogSum, which is two-speaker English conversation transcripts using #Person1# / #Person2# speaker tags. Summaries of longer, multi-party, domain-specific or non-English dialogue will be unreliable. Dialogues over 512 tokens are truncated, so content late in a long conversation may be dropped. The model inherits any biases present in google/flan-t5-large and in DialogSum, and summaries can contain details not supported by the source dialogue.

Downloads last month
5
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for daggar/flan-t5-dialogsum-lora

Adapter
(234)
this model

Dataset used to train daggar/flan-t5-dialogsum-lora