KORPatent-BGE

Model Summary[EN]

This model is a KOREAN patent-domain embedding model fine-tuned from BGE-M3 using curriculum triplet learning without relying on pseudo similarity scores (pseudo labels) generated by other models.

Instead of collecting dense, per-pair similarity scores for massive patent pairsโ€”which is costly and often infeasibleโ€”we define Anchor / Positive / Negative triplets using objective rules based on:

  • Main IPC hierarchy
  • Sub IPC hierarchy
  • Keyword overlap constraints

Training progresses from easy discrimination to hard discrimination via 5-stage Curriculum Learning, encouraging the model to learn both coarse technical boundaries and fine-grained technical distinctions.

  • Base model: BGE-M3
  • Post-hoc weight interpolation: initial fine-tuned model + KURE-v1

Model Summary[KO]

์ด ๋ชจ๋ธ์€ BGE-M3๋ฅผ ๊ธฐ๋ฐ˜์œผ๋กœ, ๋‹ค๋ฅธ ๋ชจ๋ธ์ด ์ƒ์„ฑํ•œ ์˜์‚ฌ ์œ ์‚ฌ๋„ ์ ์ˆ˜(์˜์‚ฌ ๋ ˆ์ด๋ธ”, pseudo label)์— ์˜์กดํ•˜์ง€ ์•Š๊ณ  ์ปค๋ฆฌํ˜๋Ÿผ ํŠธ๋ฆฌํ”Œ๋ › ํ•™์Šต(curriculum triplet learning)์œผ๋กœ ๋ฏธ์„ธ์กฐ์ •ํ•œ ํ•œ๊ตญ์–ด ํŠนํ—ˆ ๋„๋ฉ”์ธ ์ž„๋ฒ ๋”ฉ ๋ชจ๋ธ์ž…๋‹ˆ๋‹ค.

๋ฐฉ๋Œ€ํ•œ ํŠนํ—ˆ ์Œ์— ๋Œ€ํ•ด ์Œ๋งˆ๋‹ค ์กฐ๋ฐ€ํ•œ ์œ ์‚ฌ๋„ ์ ์ˆ˜๋ฅผ ์ˆ˜์ง‘ํ•˜๋Š” ๋ฐฉ์‹์€ ๋น„์šฉ์ด ํฌ๊ณ  ์ข…์ข… ์‹คํ˜„์ด ๋ถˆ๊ฐ€๋Šฅํ•ฉ๋‹ˆ๋‹ค. ์ด๋Ÿฐ ๋ฐฉ์‹ ๋Œ€์‹ , ๋‹ค์Œ์˜ ๊ฐ๊ด€์  ๊ทœ์น™์— ๊ธฐ๋ฐ˜ํ•ด ์•ต์ปค(Anchor) / ํฌ์ง€ํ‹ฐ๋ธŒ(Positive) / ๋„ค๊ฑฐํ‹ฐ๋ธŒ(Negative) ํŠธ๋ฆฌํ”Œ๋ ›์„ ์ •์˜ํ–ˆ์Šต๋‹ˆ๋‹ค.

  • ์ฃผ(Main) IPC ๊ณ„์ธต
  • ํ•˜์œ„(Sub) IPC ๊ณ„์ธต
  • ํ‚ค์›Œ๋“œ ์ค‘๋ณต ์ œ์•ฝ ํ•™์Šต์€ 5๋‹จ๊ณ„ ์ปค๋ฆฌํ˜๋Ÿผ ํ•™์Šต(Curriculum Learning)์„ ํ†ตํ•ด ์‰ฌ์šด ๊ตฌ๋ถ„์—์„œ ์–ด๋ ค์šด ๊ตฌ๋ถ„์œผ๋กœ ์ง„ํ–‰๋˜๋ฉฐ, ๋ชจ๋ธ์ด ๊ฑฐ์นœ ์ˆ˜์ค€์˜ ๊ธฐ์ˆ ์  ๊ฒฝ๊ณ„์™€ ๋ฏธ์„ธํ•œ ๊ธฐ์ˆ ์  ์ฐจ์ด๋ฅผ ๋ชจ๋‘ ํ•™์Šตํ•˜๋„๋ก ์œ ๋„ํ•ฉ๋‹ˆ๋‹ค.
  • ๊ธฐ๋ฐ˜ ๋ชจ๋ธ: BGE-M3
  • ์‚ฌํ›„ ๊ฐ€์ค‘์น˜ ๋ณด๊ฐ„(Post-hoc weight interpolation): ์ดˆ๊ธฐ ๋ฏธ์„ธ์กฐ์ • ๋ชจ๋ธ + KURE-v1

Key Features[EN]

  • Patent-domain specialization: robust semantic similarity in patent text.
  • No teacher-score dependency: not used external model similarity scores as pseudo labels.
  • Curriculum triplet training: with custom loss function.
  • Explicit hard negatives: negatives are intentionally selected using IPC/keyword rules.
  • Post-hoc interpolation: improves general-purpose behavior while retaining patent-domain gains.

Key Features[KO]

  • ํŠนํ—ˆ ๋„๋ฉ”์ธ ํŠนํ™”: ํŠนํ—ˆ ํ…์ŠคํŠธ์—์„œ ๊ฒฌ๊ณ ํ•œ ์˜๋ฏธ ์œ ์‚ฌ๋„๋ฅผ ์ œ๊ณตํ•ฉ๋‹ˆ๋‹ค.
  • ๊ต์‚ฌ ์ ์ˆ˜ ๋น„์˜์กด: ์™ธ๋ถ€ ๋ชจ๋ธ์˜ ์œ ์‚ฌ๋„ ์ ์ˆ˜๋ฅผ ์˜์‚ฌ ๋ ˆ์ด๋ธ”๋กœ ์‚ฌ์šฉํ•˜์ง€ ์•Š์•˜์Šต๋‹ˆ๋‹ค.
  • ์ปค๋ฆฌํ˜๋Ÿผ ํŠธ๋ฆฌํ”Œ๋ › ํ•™์Šต: ์ปค์Šคํ…€ ์†์‹ค ํ•จ์ˆ˜๋ฅผ ์‚ฌ์šฉํ•ฉ๋‹ˆ๋‹ค.
  • ๋ช…์‹œ์  ํ•˜๋“œ ๋„ค๊ฑฐํ‹ฐ๋ธŒ: ๋„ค๊ฑฐํ‹ฐ๋ธŒ๋ฅผ IPC/ํ‚ค์›Œ๋“œ ๊ทœ์น™์œผ๋กœ ์˜๋„์ ์œผ๋กœ ์„ ๋ณ„ํ•ฉ๋‹ˆ๋‹ค.
  • ์‚ฌํ›„ ๋ณด๊ฐ„: ํŠนํ—ˆ ๋„๋ฉ”์ธ ์„ฑ๋Šฅ์„ ์œ ์ง€ํ•˜๋ฉด์„œ ๋ฒ”์šฉ ์„ฑ๋Šฅ์„ ๊ฐœ์„ ํ•ฉ๋‹ˆ๋‹ค.

Intended Use[EN]

Recommended

  • Patent semantic search / prior art retrieval based on dense retrieval
  • Patent clustering, de-duplication, and nearest-neighbor exploration
  • Patent-to-patent similarity scoring (cosine similarity)
  • For RAG

Not recommended / out of scope

  • Legal advice, patent validity judgments, infringement or claim scope interpretation
  • High-stakes decisions without expert review

Intended Use[KO]

๊ถŒ์žฅ

  • ๋ฐ€์ง‘ ๊ฒ€์ƒ‰(dense retrieval) ๊ธฐ๋ฐ˜ ํŠนํ—ˆ ์˜๋ฏธ ๊ฒ€์ƒ‰ / ์„ ํ–‰๊ธฐ์ˆ  ์กฐ์‚ฌ
  • ํŠนํ—ˆ ๊ตฐ์ง‘ํ™”, ์ค‘๋ณต ์ œ๊ฑฐ, ์ตœ๊ทผ์ ‘ ์ด์›ƒ ํƒ์ƒ‰
  • ํŠนํ—ˆ ๊ฐ„ ์œ ์‚ฌ๋„ ์‚ฐ์ •(์ฝ”์‚ฌ์ธ ์œ ์‚ฌ๋„)
  • RAG ์šฉ๋„

๋น„๊ถŒ์žฅ / ์ ์šฉ ๋ฒ”์œ„ ๋ฐ–

  • ๋ฒ•๋ฅ  ์ž๋ฌธ, ํŠนํ—ˆ ์œ ํšจ์„ฑ ํŒ๋‹จ, ์นจํ•ด ๋˜๋Š” ์ฒญ๊ตฌ๋ฒ”์œ„ ํ•ด์„
  • ์ „๋ฌธ๊ฐ€ ๊ฒ€ํ†  ์—†๋Š” ์ค‘๋Œ€ํ•œ ์˜์‚ฌ๊ฒฐ์ •

Training Data (Title + Abstract + Claims)


UMAP Visualization by IPC Subclass

image Patent documents inherently share highly overlapping terminology, making it fundamentally difficult to separate them based on surface-level text alone. As an illustrative example to demonstrate our model's capability to capture true semantic context, we selected five IPC subclasses for UMAP visualization: G01D (General Measurement), G01R (Measuring Electric Variables), H03K (Pulse Techniques), G05B (Control Systems), and H01L (Semiconductor Devices). These subclasses represent a continuous technical workflow: Measurement โ†’ Signal Processing โ†’ Control โ†’ Hardware Manufacturing. While their textual vocabularies are heavily intertwined, they have distinctly different functional roles from a domain knowledge perspective.

ํŠนํ—ˆ ๋ฌธ์„œ๋Š” ๋ณธ์งˆ์ ์œผ๋กœ ๋งค์šฐ ์ค‘์ฒฉ๋œ ์šฉ์–ด๋ฅผ ๊ณต์œ ํ•˜๊ธฐ ๋•Œ๋ฌธ์—, ํ‘œ๋ฉด์ ์ธ ํ…์ŠคํŠธ๋งŒ์œผ๋กœ ๊ตฌ๋ถ„ํ•˜๊ธฐ๊ฐ€ ๊ทผ๋ณธ์ ์œผ๋กœ ์–ด๋ ต์Šต๋‹ˆ๋‹ค. ๋ชจ๋ธ์ด ์‹ค์ œ ์˜๋ฏธ์  ๋งฅ๋ฝ์„ ํฌ์ฐฉํ•˜๋Š” ๋Šฅ๋ ฅ์„ ๋ณด์—ฌ์ฃผ๊ธฐ ์œ„ํ•œ ์˜ˆ์‹œ๋กœ, UMAP ์‹œ๊ฐํ™”์— ๋‹ค์„ฏ ๊ฐœ์˜ IPC ์„œ๋ธŒํด๋ž˜์Šค๋ฅผ ์„ ํƒํ–ˆ์Šต๋‹ˆ๋‹ค: G01D(์ผ๋ฐ˜ ์ธก์ •), G01R(์ „๊ธฐ ๋ณ€๋Ÿ‰ ์ธก์ •), H03K(ํŽ„์Šค ๊ธฐ์ˆ ), G05B(์ œ์–ด ์‹œ์Šคํ…œ), H01L(๋ฐ˜๋„์ฒด ์†Œ์ž). ์ด ์„œ๋ธŒํด๋ž˜์Šค๋“ค์€ ์ธก์ • โ†’ ์‹ ํ˜ธ ์ฒ˜๋ฆฌ โ†’ ์ œ์–ด โ†’ ํ•˜๋“œ์›จ์–ด ์ œ์กฐ๋กœ ์ด์–ด์ง€๋Š” ์—ฐ์†์ ์ธ ๊ธฐ์ˆ  ํ๋ฆ„์„ ๋‚˜ํƒ€๋ƒ…๋‹ˆ๋‹ค. ํ…์ŠคํŠธ ์–ดํœ˜๋Š” ์„œ๋กœ ํฌ๊ฒŒ ์–ฝํ˜€ ์žˆ์ง€๋งŒ, ๋„๋ฉ”์ธ ์ง€์‹ ๊ด€์ ์—์„œ๋Š” ๊ธฐ๋Šฅ์  ์—ญํ• ์ด ๋šœ๋ ท์ด ๋‹ค๋ฆ…๋‹ˆ๋‹ค.


Usage

Sentence-Transformers

from sentence_transformers import SentenceTransformer
import torch

model = SentenceTransformer("sttempler/KORPatent-BGE", device="cuda" if torch.cuda.is_available() else "cpu")

sentences = [
    "์ธ๊ฐ„ ์„ ํ˜ธ ๋ฐ์ดํ„ฐ๋ฅผ ์ด์šฉํ•˜์—ฌ ์ƒ๋‹ด ์ฑ—๋ด‡ ์‘๋‹ต์„ ์ •๋ ฌํ•˜๋Š” RLHF ๊ธฐ๋ฐ˜ ํ•™์Šต ๋ฐฉ๋ฒ•.",
    "์„ ํ˜ธ ํ”ผ๋“œ๋ฐฑ์œผ๋กœ ๋ณด์ƒ ์‹ ํ˜ธ๋ฅผ ๊ตฌ์„ฑํ•˜์—ฌ ์ƒ์„ฑ ๋ชจ๋ธ์„ ์›ํ•˜๋Š” ๋ฐฉํ–ฅ์œผ๋กœ ํ•™์Šต์‹œํ‚ค๋Š” ๋ฐฉ๋ฒ•."
]
emb = model.encode(sentences, normalize_embeddings=True)
score = float(emb[0] @ emb[1].T)
print(score)
# 0.7382504343986511

Citation

If you use this model in your research, please cite the following doctoral dissertation:

@phdthesis{kim2026koreanpatentembedding,
  title        = {A Study on Domain-Specific Long-Context Embedding Models for Advancing Korean Patent Retrieval},
  author       = {Kim, Yongwoo},
  school       = {Hanyang University},
  department   = {Graduate School of Technology & Innovation Management},
  year         = {2026},
  type         = {Doctoral dissertation},
  note         = {Korean title: ํ•œ๊ตญ์–ด ํŠนํ—ˆ ๊ฒ€์ƒ‰์„ ์œ„ํ•œ ์žฅ๋ฌธ ์ž„๋ฒ ๋”ฉ ๋ชจ๋ธ}
}
Downloads last month
143
Safetensors
Model size
0.6B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for sttempler/KORPatent-BGE

Base model

BAAI/bge-m3
Finetuned
(518)
this model