contract-clause-phi3-lora

A QLoRA fine-tune of microsoft/Phi-3-mini-4k-instruct for legal contract clause classification, trained on the CUAD (Contract Understanding Atlas Dataset) dataset. Given a contract clause, it predicts which of 41 legal clause categories it belongs to (or "None").

Code, training pipeline, and full evaluation: github.com/D-L-C-S/contract-clause-qlora

Results

Evaluated on a held-out test set of 2,784 clauses, compared against the zero-shot base model:

Metric Zero-shot base model This adapter
Accuracy 21.6% 62.5%
Invalid (unparseable) output rate 58.5% 0.5%

Full per-category metrics, error analysis, and known limitations are in the eval notebook.

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel

base_model = AutoModelForCausalLM.from_pretrained(
    "microsoft/Phi-3-mini-4k-instruct",
    quantization_config=BitsAndBytesConfig(
        load_in_4bit=True,
        bnb_4bit_quant_type="nf4",
        bnb_4bit_compute_dtype=torch.float16,
    ),
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("microsoft/Phi-3-mini-4k-instruct")
model = PeftModel.from_pretrained(base_model, "DLCS/contract-clause-phi3-lora")
model.eval()

clause = "This Agreement shall be governed by and construed under the laws of the State of Delaware."
prompt_messages = [{
    "role": "user",
    "content": f'Classify the following contract clause into its category, or respond with "None" if it does not match any category. Respond with only the category name and nothing else.\n\nClause:\n{clause}',
}]
prompt = tokenizer.apply_chat_template(prompt_messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

with torch.inference_mode():
    output = model.generate(**inputs, max_new_tokens=18)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
# "Governing Law"

Training details

  • Base model: microsoft/Phi-3-mini-4k-instruct, loaded in 4-bit (NF4)
  • Method: LoRA (r=16, alpha=32, dropout=0.05) on qkv_proj, o_proj, gate_up_proj, down_proj — ~0.57% of total parameters trainable
  • Data: CUAD's annotated clause spans (positive examples) plus heuristically-constructed "None" examples from unannotated text gaps, split at the contract level to prevent leakage
  • Training run: 2 epochs, 376 steps, full fp32 (see the GitHub repo for why — a Turing-GPU bf16 limitation combined with a bitsandbytes/GradScaler bug forced this)

Limitations

Single-label classification only; several of the 41 categories have very few test examples, so per-category metrics for those are directional, not statistically reliable; the "None" class is heuristically constructed and carries some inherent label noise. See the eval notebook for a full, honest discussion.

Downloads last month
13
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DLCS/contract-clause-phi3-lora

Adapter
(865)
this model

Dataset used to train DLCS/contract-clause-phi3-lora