BSVGK's picture
Update README.md
ecc30d9 verified
|
Raw
History Blame Contribute Delete
2.02 kB
---
language:
- en
license: mit
base_model: microsoft/Phi-3.5-mini-instruct
tags:
- lora
- adapter
- text-generation
- knowledge-graph
- information-extraction
- rdf
- fine-tuned
- peft
- trustworthy-ai
- hallucination-detection
datasets:
- BSVGK/Text_to_KG_Construction_Dataset
pipeline_tag: text-generation
---
# Phi-3.5 Mini Instruct โ€” LoRA Adapter (Text-to-KG)
## Model Summary
This is the **LoRA adapter** for the Phi-3.5 Mini Instruct model fine-tuned to extract structured **RDF knowledge graph triples** from UK government procurement contract text.
> For the full merged model ready for inference, use:
> ๐Ÿ‘‰ [BSVGK/phi35-mini-lora-text2kg-merged](https://huggingface.co/BSVGK/phi35-mini-lora-text2kg-merged)
## Key Results
| Metric | Score |
|--------|-------|
| F1 Score | **0.9954** |
| BERTScore F1 | **0.9997** |
| Hallucination Rate | **0.00% (Zero)** |
| Test Contracts | 1,387 unseen contracts |
## Model Details
- **Base Model:** microsoft/Phi-3.5-mini-instruct
- **Adapter Type:** LoRA (Low-Rank Adaptation)
- **Task:** Text-to-KG โ€” RDF triple extraction from contract text
- **Domain:** UK Government Procurement Contracts
- **Training Dataset:** 9,244 verified UK contracts
- **Hardware:** NVIDIA A100
- **Framework:** PyTorch, Hugging Face PEFT, TRL, SFTTrainer
## Usage
```python
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
# Load base model
base_model = AutoModelForCausalLM.from_pretrained(
"microsoft/Phi-3.5-mini-instruct"
)
tokenizer = AutoTokenizer.from_pretrained(
"microsoft/Phi-3.5-mini-instruct"
)
# Load LoRA adapter
model = PeftModel.from_pretrained(
base_model,
"BSVGK/phi35-mini-lora-text2kg-adapter"
)
prompt = """Extract RDF triples from the following UK government contract:
Contract: [paste your contract text here]
RDF Triples:"""
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))