KORPatent-BGE / README.md
sttempler's picture
Update README.md
d760ffa verified
|
Raw
History Blame Contribute Delete
7.41 kB
---
license: mit
language:
- ko
base_model:
- BAAI/bge-m3
pipeline_tag: sentence-similarity
library_name: sentence-transformers
tags:
- patent
- feature-extraction
- sentence-similarity
---
# KORPatent-BGE
## Model Summary[EN]
This model is a KOREAN patent-domain embedding model fine-tuned from [BGE-M3](https://huggingface.co/BAAI/bge-m3) using curriculum triplet learning without relying on pseudo similarity scores (pseudo labels) generated by other models.
Instead of collecting dense, per-pair similarity scores for massive patent pairsโ€”which is costly and often infeasibleโ€”we define Anchor / Positive / Negative triplets using objective rules based on:
- Main IPC hierarchy
- Sub IPC hierarchy
- Keyword overlap constraints
Training progresses from easy discrimination to hard discrimination via 5-stage Curriculum Learning, encouraging the model to learn both coarse technical boundaries and fine-grained technical distinctions.
- Base model: [BGE-M3](https://huggingface.co/BAAI/bge-m3)
- Post-hoc weight interpolation: initial fine-tuned model + [KURE-v1](https://huggingface.co/nlpai-lab/KURE-v1)
## Model Summary[KO]
์ด ๋ชจ๋ธ์€ [BGE-M3](https://huggingface.co/BAAI/bge-m3)๋ฅผ ๊ธฐ๋ฐ˜์œผ๋กœ, ๋‹ค๋ฅธ ๋ชจ๋ธ์ด ์ƒ์„ฑํ•œ ์˜์‚ฌ ์œ ์‚ฌ๋„ ์ ์ˆ˜(์˜์‚ฌ ๋ ˆ์ด๋ธ”, pseudo label)์— ์˜์กดํ•˜์ง€ ์•Š๊ณ  ์ปค๋ฆฌํ˜๋Ÿผ ํŠธ๋ฆฌํ”Œ๋ › ํ•™์Šต(curriculum triplet learning)์œผ๋กœ ๋ฏธ์„ธ์กฐ์ •ํ•œ ํ•œ๊ตญ์–ด ํŠนํ—ˆ ๋„๋ฉ”์ธ ์ž„๋ฒ ๋”ฉ ๋ชจ๋ธ์ž…๋‹ˆ๋‹ค.
๋ฐฉ๋Œ€ํ•œ ํŠนํ—ˆ ์Œ์— ๋Œ€ํ•ด ์Œ๋งˆ๋‹ค ์กฐ๋ฐ€ํ•œ ์œ ์‚ฌ๋„ ์ ์ˆ˜๋ฅผ ์ˆ˜์ง‘ํ•˜๋Š” ๋ฐฉ์‹์€ ๋น„์šฉ์ด ํฌ๊ณ  ์ข…์ข… ์‹คํ˜„์ด ๋ถˆ๊ฐ€๋Šฅํ•ฉ๋‹ˆ๋‹ค. ์ด๋Ÿฐ ๋ฐฉ์‹ ๋Œ€์‹ , ๋‹ค์Œ์˜ ๊ฐ๊ด€์  ๊ทœ์น™์— ๊ธฐ๋ฐ˜ํ•ด ์•ต์ปค(Anchor) / ํฌ์ง€ํ‹ฐ๋ธŒ(Positive) / ๋„ค๊ฑฐํ‹ฐ๋ธŒ(Negative) ํŠธ๋ฆฌํ”Œ๋ ›์„ ์ •์˜ํ–ˆ์Šต๋‹ˆ๋‹ค.
- ์ฃผ(Main) IPC ๊ณ„์ธต
- ํ•˜์œ„(Sub) IPC ๊ณ„์ธต
- ํ‚ค์›Œ๋“œ ์ค‘๋ณต ์ œ์•ฝ
ํ•™์Šต์€ 5๋‹จ๊ณ„ ์ปค๋ฆฌํ˜๋Ÿผ ํ•™์Šต(Curriculum Learning)์„ ํ†ตํ•ด ์‰ฌ์šด ๊ตฌ๋ถ„์—์„œ ์–ด๋ ค์šด ๊ตฌ๋ถ„์œผ๋กœ ์ง„ํ–‰๋˜๋ฉฐ, ๋ชจ๋ธ์ด ๊ฑฐ์นœ ์ˆ˜์ค€์˜ ๊ธฐ์ˆ ์  ๊ฒฝ๊ณ„์™€ ๋ฏธ์„ธํ•œ ๊ธฐ์ˆ ์  ์ฐจ์ด๋ฅผ ๋ชจ๋‘ ํ•™์Šตํ•˜๋„๋ก ์œ ๋„ํ•ฉ๋‹ˆ๋‹ค.
- ๊ธฐ๋ฐ˜ ๋ชจ๋ธ: [BGE-M3](https://huggingface.co/BAAI/bge-m3)
- ์‚ฌํ›„ ๊ฐ€์ค‘์น˜ ๋ณด๊ฐ„(Post-hoc weight interpolation): ์ดˆ๊ธฐ ๋ฏธ์„ธ์กฐ์ • ๋ชจ๋ธ + [KURE-v1](https://huggingface.co/nlpai-lab/KURE-v1)
---
## Key Features[EN]
- **Patent-domain specialization**: robust semantic similarity in patent text.
- **No teacher-score dependency**: not used external model similarity scores as pseudo labels.
- **Curriculum triplet training**: with custom loss function.
- **Explicit hard negatives**: negatives are intentionally selected using IPC/keyword rules.
- **Post-hoc interpolation**: improves general-purpose behavior while retaining patent-domain gains.
## Key Features[KO]
- **ํŠนํ—ˆ ๋„๋ฉ”์ธ ํŠนํ™”**: ํŠนํ—ˆ ํ…์ŠคํŠธ์—์„œ ๊ฒฌ๊ณ ํ•œ ์˜๋ฏธ ์œ ์‚ฌ๋„๋ฅผ ์ œ๊ณตํ•ฉ๋‹ˆ๋‹ค.
- **๊ต์‚ฌ ์ ์ˆ˜ ๋น„์˜์กด**: ์™ธ๋ถ€ ๋ชจ๋ธ์˜ ์œ ์‚ฌ๋„ ์ ์ˆ˜๋ฅผ ์˜์‚ฌ ๋ ˆ์ด๋ธ”๋กœ ์‚ฌ์šฉํ•˜์ง€ ์•Š์•˜์Šต๋‹ˆ๋‹ค.
- **์ปค๋ฆฌํ˜๋Ÿผ ํŠธ๋ฆฌํ”Œ๋ › ํ•™์Šต**: ์ปค์Šคํ…€ ์†์‹ค ํ•จ์ˆ˜๋ฅผ ์‚ฌ์šฉํ•ฉ๋‹ˆ๋‹ค.
- **๋ช…์‹œ์  ํ•˜๋“œ ๋„ค๊ฑฐํ‹ฐ๋ธŒ**: ๋„ค๊ฑฐํ‹ฐ๋ธŒ๋ฅผ IPC/ํ‚ค์›Œ๋“œ ๊ทœ์น™์œผ๋กœ ์˜๋„์ ์œผ๋กœ ์„ ๋ณ„ํ•ฉ๋‹ˆ๋‹ค.
- **์‚ฌํ›„ ๋ณด๊ฐ„**: ํŠนํ—ˆ ๋„๋ฉ”์ธ ์„ฑ๋Šฅ์„ ์œ ์ง€ํ•˜๋ฉด์„œ ๋ฒ”์šฉ ์„ฑ๋Šฅ์„ ๊ฐœ์„ ํ•ฉ๋‹ˆ๋‹ค.
---
## Intended Use[EN]
**Recommended**
- Patent semantic search / prior art retrieval based on dense retrieval
- Patent clustering, de-duplication, and nearest-neighbor exploration
- Patent-to-patent similarity scoring (cosine similarity)
- For RAG
**Not recommended / out of scope**
- Legal advice, patent validity judgments, infringement or claim scope interpretation
- High-stakes decisions without expert review
## Intended Use[KO]
**๊ถŒ์žฅ**
- ๋ฐ€์ง‘ ๊ฒ€์ƒ‰(dense retrieval) ๊ธฐ๋ฐ˜ ํŠนํ—ˆ ์˜๋ฏธ ๊ฒ€์ƒ‰ / ์„ ํ–‰๊ธฐ์ˆ  ์กฐ์‚ฌ
- ํŠนํ—ˆ ๊ตฐ์ง‘ํ™”, ์ค‘๋ณต ์ œ๊ฑฐ, ์ตœ๊ทผ์ ‘ ์ด์›ƒ ํƒ์ƒ‰
- ํŠนํ—ˆ ๊ฐ„ ์œ ์‚ฌ๋„ ์‚ฐ์ •(์ฝ”์‚ฌ์ธ ์œ ์‚ฌ๋„)
- RAG ์šฉ๋„
**๋น„๊ถŒ์žฅ / ์ ์šฉ ๋ฒ”์œ„ ๋ฐ–**
- ๋ฒ•๋ฅ  ์ž๋ฌธ, ํŠนํ—ˆ ์œ ํšจ์„ฑ ํŒ๋‹จ, ์นจํ•ด ๋˜๋Š” ์ฒญ๊ตฌ๋ฒ”์œ„ ํ•ด์„
- ์ „๋ฌธ๊ฐ€ ๊ฒ€ํ†  ์—†๋Š” ์ค‘๋Œ€ํ•œ ์˜์‚ฌ๊ฒฐ์ •
---
## Training Data (Title + Abstract + Claims)
- [ํŠนํ—ˆ ๋ถ„์•ผ ์ž๋™๋ถ„๋ฅ˜ ๋ฐ์ดํ„ฐ](https://aihub.or.kr/aihubdata/data/view.do?dataSetSn=547)
- [๊ตญ๊ฐ€์ค‘์ ๊ธฐ์ˆ ๋Œ€์‘ ํŠนํ—ˆ ๋ฐ์ดํ„ฐ](https://aihub.or.kr/aihubdata/data/view.do?currMenu=115&dataSetSn=71739)
- [์‚ฐ์—…์ •๋ณด ์—ฐ๊ณ„ ์ฃผ์š”๊ตญ ํŠนํ—ˆ ์˜-ํ•œ ๋ฐ์ดํ„ฐ](https://aihub.or.kr/aihubdata/data/view.do?dataSetSn=563)
- [๊ณผํ•™๊ธฐ์ˆ ํ‘œ์ค€๋ถ„๋ฅ˜ ๋Œ€์‘ ํŠนํ—ˆ ๋ฐ์ดํ„ฐ](https://aihub.or.kr/aihubdata/data/view.do?dataSetSn=71531)
---
## UMAP Visualization by IPC Subclass
![image](https://cdn-uploads.huggingface.co/production/uploads/644b861641003d2ecf6e098b/p1vrxE46RKe3ZLa2Pl5G5.png)
Patent documents inherently share highly overlapping terminology, making it fundamentally difficult to separate them based on surface-level text alone.
As an illustrative example to demonstrate our model's capability to capture true semantic context, we selected five IPC subclasses for UMAP visualization: G01D (General Measurement), G01R (Measuring Electric Variables), H03K (Pulse Techniques), G05B (Control Systems), and H01L (Semiconductor Devices).
These subclasses represent a continuous technical workflow: Measurement โ†’ Signal Processing โ†’ Control โ†’ Hardware Manufacturing. While their textual vocabularies are heavily intertwined, they have distinctly different functional roles from a domain knowledge perspective.
ํŠนํ—ˆ ๋ฌธ์„œ๋Š” ๋ณธ์งˆ์ ์œผ๋กœ ๋งค์šฐ ์ค‘์ฒฉ๋œ ์šฉ์–ด๋ฅผ ๊ณต์œ ํ•˜๊ธฐ ๋•Œ๋ฌธ์—, ํ‘œ๋ฉด์ ์ธ ํ…์ŠคํŠธ๋งŒ์œผ๋กœ ๊ตฌ๋ถ„ํ•˜๊ธฐ๊ฐ€ ๊ทผ๋ณธ์ ์œผ๋กœ ์–ด๋ ต์Šต๋‹ˆ๋‹ค.
๋ชจ๋ธ์ด ์‹ค์ œ ์˜๋ฏธ์  ๋งฅ๋ฝ์„ ํฌ์ฐฉํ•˜๋Š” ๋Šฅ๋ ฅ์„ ๋ณด์—ฌ์ฃผ๊ธฐ ์œ„ํ•œ ์˜ˆ์‹œ๋กœ, UMAP ์‹œ๊ฐํ™”์— ๋‹ค์„ฏ ๊ฐœ์˜ IPC ์„œ๋ธŒํด๋ž˜์Šค๋ฅผ ์„ ํƒํ–ˆ์Šต๋‹ˆ๋‹ค: G01D(์ผ๋ฐ˜ ์ธก์ •), G01R(์ „๊ธฐ ๋ณ€๋Ÿ‰ ์ธก์ •), H03K(ํŽ„์Šค ๊ธฐ์ˆ ), G05B(์ œ์–ด ์‹œ์Šคํ…œ), H01L(๋ฐ˜๋„์ฒด ์†Œ์ž).
์ด ์„œ๋ธŒํด๋ž˜์Šค๋“ค์€ ์ธก์ • โ†’ ์‹ ํ˜ธ ์ฒ˜๋ฆฌ โ†’ ์ œ์–ด โ†’ ํ•˜๋“œ์›จ์–ด ์ œ์กฐ๋กœ ์ด์–ด์ง€๋Š” ์—ฐ์†์ ์ธ ๊ธฐ์ˆ  ํ๋ฆ„์„ ๋‚˜ํƒ€๋ƒ…๋‹ˆ๋‹ค. ํ…์ŠคํŠธ ์–ดํœ˜๋Š” ์„œ๋กœ ํฌ๊ฒŒ ์–ฝํ˜€ ์žˆ์ง€๋งŒ, ๋„๋ฉ”์ธ ์ง€์‹ ๊ด€์ ์—์„œ๋Š” ๊ธฐ๋Šฅ์  ์—ญํ• ์ด ๋šœ๋ ท์ด ๋‹ค๋ฆ…๋‹ˆ๋‹ค.
---
## Usage
### Sentence-Transformers
```python
from sentence_transformers import SentenceTransformer
import torch
model = SentenceTransformer("sttempler/KORPatent-BGE", device="cuda" if torch.cuda.is_available() else "cpu")
sentences = [
"์ธ๊ฐ„ ์„ ํ˜ธ ๋ฐ์ดํ„ฐ๋ฅผ ์ด์šฉํ•˜์—ฌ ์ƒ๋‹ด ์ฑ—๋ด‡ ์‘๋‹ต์„ ์ •๋ ฌํ•˜๋Š” RLHF ๊ธฐ๋ฐ˜ ํ•™์Šต ๋ฐฉ๋ฒ•.",
"์„ ํ˜ธ ํ”ผ๋“œ๋ฐฑ์œผ๋กœ ๋ณด์ƒ ์‹ ํ˜ธ๋ฅผ ๊ตฌ์„ฑํ•˜์—ฌ ์ƒ์„ฑ ๋ชจ๋ธ์„ ์›ํ•˜๋Š” ๋ฐฉํ–ฅ์œผ๋กœ ํ•™์Šต์‹œํ‚ค๋Š” ๋ฐฉ๋ฒ•."
]
emb = model.encode(sentences, normalize_embeddings=True)
score = float(emb[0] @ emb[1].T)
print(score)
# 0.7382504343986511
```
---
### Citation
If you use this model in your research, please cite the following doctoral dissertation:
```bibtex
@phdthesis{kim2026koreanpatentembedding,
title = {A Study on Domain-Specific Long-Context Embedding Models for Advancing Korean Patent Retrieval},
author = {Kim, Yongwoo},
school = {Hanyang University},
department = {Graduate School of Technology & Innovation Management},
year = {2026},
type = {Doctoral dissertation},
note = {Korean title: ํ•œ๊ตญ์–ด ํŠนํ—ˆ ๊ฒ€์ƒ‰์„ ์œ„ํ•œ ์žฅ๋ฌธ ์ž„๋ฒ ๋”ฉ ๋ชจ๋ธ}
}
```