Sentence Similarity
PEFT
Safetensors
sentence-transformers
Korean
feature-extraction
korean
fiction
stylometry
authorship-analysis
lora
Instructions to use Baragi-AI/Munche-v2-768 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Baragi-AI/Munche-v2-768 with PEFT:
Task type is invalid.
- sentence-transformers
How to use Baragi-AI/Munche-v2-768 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("Baragi-AI/Munche-v2-768") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
| language: | |
| - ko | |
| license: gemma | |
| library_name: peft | |
| pipeline_tag: sentence-similarity | |
| base_model: google/embeddinggemma-300m | |
| tags: | |
| - sentence-transformers | |
| - feature-extraction | |
| - korean | |
| - fiction | |
| - stylometry | |
| - authorship-analysis | |
| - lora | |
| # Munche-v2-768 | |
| **Munche-v2-768**์ ํ๊ตญ์ด ์ฅ๋ฅด์์ค์ *๋ด์ฉ*๋ณด๋ค ๋ฌธ์ฅ ์ด์ฉ, ์์ ๋ฆฌ๋ฌ, ํํยท๊ธฐ๋ฅ์ด ์ฌ์ฉ๊ณผ ๊ฐ์ *๋ฌธ์ฒด*๋ฅผ ๋น๊ตํ๊ธฐ ์ํด ํ์ตํ 768์ฐจ์ ํ ์คํธ ์๋ฒ ๋ฉ ๋ชจ๋ธ์ ๋๋ค. [`google/embeddinggemma-300m`](https://huggingface.co/google/embeddinggemma-300m)์ ์๋ 768์ฐจ์ pooling/projection ๊ฒฝ๋ก๋ฅผ ์ ์งํ๋ฉด์, style LoRA๋ฅผ ํ์ตํ์ต๋๋ค. | |
| ์ด ๋ชจ๋ธ์ ์ผ๋ฐ ์๋ฏธ ๊ฒ์ ๋ชจ๋ธ์ ๋์ฒด์ฌ๊ฐ ์๋๋๋ค. ๋์ผยท์ ์ฌํ ๋ด์ฉ์ ์ฐพ๋ ๊ฒ๋ณด๋ค ์๋ก ๋ค๋ฅธ ์ํ์ ๋ฐ๋ณต๋๋ ์๊ฐ์ ๋ฌธ์ฒด๋ฅผ ๋น๊ตํ๋ ์ฉ๋๋ก ์ค๊ณํ์ต๋๋ค. | |
|  | |
| ## ์ฃผ์ ํน์ง | |
| - **์๋ณธ 768์ฐจ์ head ์ ์ง:** ์๋ก์ด projection head๋ฅผ ๋ง๋ถ์ด์ง ์๊ณ EmbeddingGemma์ mean pooling๊ณผ ๋ projection layer๋ฅผ ๊ทธ๋๋ก ์ฌ์ฉํฉ๋๋ค. | |
| - **Style LoRA:** ๋๊ฒฐ๋ backbone์ `q_proj`, `v_proj`, `o_proj`์ rank 16, alpha 32, dropout 0.05์ LoRA๋ฅผ ํ์ตํ์ต๋๋ค. ์๋ณธ pooling/projection layer๋ ๋๊ฒฐํ์ต๋๋ค. | |
| - **ํ์ค PEFT adapter:** LoRA๋ฅผ ๋ณํฉํ์ง ์๊ณ ํ์ต๋ adapter ๊ทธ๋๋ก ์ ๊ณตํฉ๋๋ค. | |
| - **ํ ๊ณต๊ฐ์์ ๊ณต๋ ํ์ต:** ์ํ, ์๊ฐ, ๋ค์ค prototype, content-hard, counterfactual ์ ํธ๊ฐ ๋ชจ๋ ์ต์ข 768์ฐจ์ cosine ๊ณต๊ฐ์ ์ง์ ์์ฉํฉ๋๋ค. | |
| - **๊ธด ํ ์คํธ:** ํ์ต ๊ตฌ๊ฐ์ 512/768/1024 token์ด๋ฉฐ, 1024 token์ ๋๋ ์ ๋ ฅ์ 512 stride sliding window์ overlap-corrected spherical pooling์ ๊ถ์ฅํฉ๋๋ค. | |
| - **๋ณด์กฐ ๊ณผ์ :** ์ฐ์ฌ ์๊ธฐ, Kiwi stylometry, Human/AI ๋ถ๋ฅ๋ ๋ณ๋ ๋ณด์กฐ head๋ก ํ์ตํ๋ encoder gradient๋ฅผ ์ ํํ๊ฑฐ๋ ํ๋ฐ์ ๊ฐ์ ์์ผฐ์ต๋๋ค. ๊ธฐ๋ณธ ์๋ฒ ๋ฉ API๋ ์ด ๋ณด์กฐ ์์ธก๊ฐ์ด ์๋๋ผ L2-normalized 768์ฐจ์ ๋ฒกํฐ๋ฅผ ๋ฐํํฉ๋๋ค. | |
| ## ์ฌ์ฉ๋ฒ | |
| EmbeddingGemma์ ๋ฌธ์ prompt๋ฅผ ํฌํจํด ์ ๋ ฅํ๋ ๊ฒ์ ๊ถ์ฅํฉ๋๋ค. | |
| ```python | |
| import torch | |
| from peft import PeftModel | |
| from sentence_transformers import SentenceTransformer | |
| model = SentenceTransformer("google/embeddinggemma-300m").to(torch.bfloat16) | |
| model[0].auto_model = PeftModel.from_pretrained( | |
| model[0].auto_model, | |
| "Baragi-AI/Munche-v2-768", | |
| ) | |
| model.max_seq_length = 1024 | |
| texts = [ | |
| "title: none | text: ๊ทธ๋ ๋๋ตํ์ง ์์๋ค. ์ฐฝ๋ฐ์ ๋น๊ฐ ์ค๋๋ ์ง๋ถ์ ๋๋๋ ธ๋ค.", | |
| "title: none | text: ๋๋ ๊ฒ์ ๋ด๋ ค๋์๋ค. ํด์ผ ํ ๋ง์ ์ด๋ฏธ ๋ชจ๋ ๋๋ ๋ค์๋ค.", | |
| ] | |
| embeddings = model.encode( | |
| texts, | |
| normalize_embeddings=True, | |
| convert_to_numpy=True, | |
| ) | |
| similarity = embeddings @ embeddings.T | |
| ``` | |
| ํ ์ํ ์ ์ฒด๋ฅผ ์๋ฒ ๋ฉํ ๋๋ ๋ค์ ์ ์ฐจ๋ฅผ ๊ถ์ฅํฉ๋๋ค. | |
| 1. ์ค์ tokenizer ๊ธฐ์ค 1024-token window์ 512-token stride๋ฅผ ์ฌ์ฉํฉ๋๋ค. | |
| 2. ๊ฐ window๋ฅผ ๊ฐ๋ณ์ ์ผ๋ก L2 normalizeํฉ๋๋ค. | |
| 3. ๊ฒน์น token์ด ์ฌ๋ฌ ๋ฒ ์ง๊ณ๋์ง ์๋๋ก window๋ณ token coverage ์ญ์๋ฅผ ๊ฐ์ค์น๋ก ์ฌ์ฉํฉ๋๋ค. | |
| 4. ๊ฐ์ค ํ๊ท ๊ฒฐ๊ณผ๋ฅผ ๋ค์ L2 normalizeํฉ๋๋ค. | |
| 5. ์ํ ๊ธธ์ด ํธํฅ์ ์ค์ด๋ ค๋ฉด ๋จผ์ ํ์ฐจ๋ณ๋ก poolingํ ๋ค ํ์ฐจ ๋ฒกํฐ๋ฅผ ๋์ผ ๊ฐ์ค ํ๊ท ํฉ๋๋ค. | |
| ์ด ๋ชจ๋ธ์ BF16์ผ๋ก ํ์ตยทํ๊ฐํ์ผ๋ฉฐ FP16 activation์ ์ง์ํ์ง ์์ต๋๋ค. | |
| ## ๋ชจ๋ธ ๊ตฌ์กฐ | |
| ```text | |
| text + document prompt | |
| โ frozen EmbeddingGemma 300M backbone | |
| + trainable Q/V/O LoRA | |
| โ frozen original mean pooling | |
| โ frozen original Dense โ Dense (768d) | |
| โ L2 normalization | |
| โ style embedding z โ R^768 | |
| โโ scalar ordinal publication head [training auxiliary] | |
| โโ Kiwi stylometry MLP [training auxiliary] | |
| โโ Human/AI binary head [training auxiliary] | |
| ``` | |
| ## ํ์ต ๋ฐฉ๋ฒ | |
| ### ๋ฐ์ดํฐ ๋ถํ ๊ณผ sampling | |
| - ์๊ฐ๊ฐ ํ์ธ๋ ๋ฐ์ดํฐ๋ ์๊ฐ ์ฐ๊ฒฐ์์ ๋จ์๋ก train/validation/test๋ฅผ ๋ถ๋ฆฌํ์ต๋๋ค. ๊ฐ์ ์๊ฐ์ ์ฌ๋ฌ ์ํ๊ณผ ๊ฐ์ ์ํ์ ๋ชจ๋ ํ์ window๋ ํ๋์ split์๋ง ์กด์ฌํฉ๋๋ค. | |
| - ์ ํ ์ค๋ณต๊ณผ near-duplicate ์ฐ๊ฒฐ์์๋ฅผ ๋จผ์ ์ฒ๋ฆฌํด `processed_data`์ ํ์ counterfactual ๋ฐ์ดํฐ์ ๋์๋ฅผ ์ค์์ต๋๋ค. | |
| - ๊ธด ์ํ์ด ํ์ต์ ๋ ์ ํ์ง ์๋๋ก ์ํ์ ๋จผ์ ๊ท ํ samplingํ๊ณ , ์ํ ์์์ ๋จ์ด์ง ์์น์ window๋ฅผ ์ ํํ์ต๋๋ค. | |
| - ์ผ๋ฐ metric batch๋ `8 authors ร 3 works ร 2 windows`์ ๋๋ค. ์ธ๊ฐ ์ํ metric loss๋ ๋งค ๋ ๋ฒ์งธ step์ ์ ์ฉํ์ต๋๋ค. | |
| - ์ต์ข ๋จ๊ณ์์๋ 10 step๋ง๋ค ํ ๋ฒ `4 authors ร 4 works ร 3 windows`์ prototype ์ ์ฉ batch๋ฅผ ์ฌ์ฉํ์ต๋๋ค. | |
| ### ์ต์ข embedding์ ์ง์ ์ ์ฉํ ๋ชฉ์ ํจ์ | |
| 1. **Work metric loss** โ ๊ฐ์ ์ํ์ ์๋ก ๋จ์ด์ง ๊ตฌ๊ฐ์ ๊ฐ๊น๊ฒ ํ์ตํฉ๋๋ค. ๊ฐ์ ์๊ฐ์ ๋ค๋ฅธ ์ํ์ ์ํ loss์ negative์์ ์ ์ธํฉ๋๋ค. | |
| 2. **Cross-work author loss** โ ๊ฐ์ ์๊ฐ์ ์๋ก ๋ค๋ฅธ ์ํ์ ๊ฐ๊น๊ฒ ํ๋, ํ๋ฐ์๋ ๋จ์ผ centroid ์๋ ฅ์ ๊ฐ์ ํฉ๋๋ค. | |
| 3. **Leave-one-work-out multi-prototype loss** โ ์๊ฐ๋น `N=3` prototype์ support ์ํ์ผ๋ก ๋ง๋ค๊ณ , ์ ์ธํ query ์ํ์ window๋ฅผ ๋ถ๋ฅํฉ๋๋ค. | |
| 4. **Work-balanced prototype construction** โ ์ํ๋ณ local assignment๋ฅผ ๋จผ์ ๊ณ์ฐํ๊ณ ์ํ๋ง๋ค ๊ฐ์ ๊ฐ์ค์น๋ฅผ ์ฃผ์ด, window๊ฐ ๋ง์ ์ํ์ด prototype์ ์ง๋ฐฐํ์ง ์๊ฒ ํฉ๋๋ค. | |
| 5. **Cross-work coverage + diversity** โ ๊ฐ prototype์ด ์ต์ ๋ ์ํ์์ ์ง์ง๋ฅผ ๋ฐ๋๋ก effective-work ๋ฐ second-work-mass hinge๋ฅผ ์ ์ฉํ๊ณ , ์ถฉ๋ถํ ์ง์ง๋๋ prototype๋ผ๋ฆฌ๋ง separation์ ์ ๋ํฉ๋๋ค. Prototype ์ ์ฉ batch์์๋ coverage ๊ธฐ์ฌ๋ฅผ 1.5๋ฐฐ๋ก ์ ์ฉํ์ต๋๋ค. | |
| 6. **Semantic hard negatives** โ ๋๊ฒฐ๋ ์๋ณธ EmbeddingGemma์์ ์๋ฏธ๊ฐ ๊ฐ๊น์ด ๋ค๋ฅธ ์๊ฐ์ ๊ตฌ๊ฐ 20๊ฐ๋ฅผ ์ฐพ์ style ๊ณต๊ฐ์์๋ ๋ฉ์ด์ง๊ฒ ํฉ๋๋ค. | |
| 7. **Conditional decorrelation** โ ๊ฐ์ ์๊ฐ ์์์ ๋ด์ฉ semantic embedding์ด ์์ง์ด๋ ๋ฐฉํฅ์ style embedding์ด ๊ทธ๋๋ก ๋ฐ๋ฅด์ง ์๋๋ก cross-covariance๋ฅผ ์ ํํฉ๋๋ค. ์ด๋ฐ์๋ ๋ฐฉํฅ ํ์ฑ์ ์ฌ์ฉํ๊ณ ํ๋ฐ์๋ guardrail๋ก ๋ฎ์ท์ต๋๋ค. | |
| 8. **Human/LLM counterfactual ranking** โ ์ธ๊ฐ ์๋ฌธ๊ณผ ๋ด์ฉ ๋ณด์กด LLM rewrite๋ฅผ ๊ตฌ๋ถํ๋๋ก, ์ธ๊ฐ ์๊ฐยท์ํ positive๊ฐ rewrite๋ณด๋ค ๊ฐ๊น๊ฒ ํ์ตํฉ๋๋ค. | |
| 9. **Synthetic hierarchy** โ ๋์ผ ๋ด์ฉ blueprint์์ `same recipe > same model/different prompt > different model/same prompt > different model/different prompt` ์์๋ฅผ ์ ๋ํ๊ณ ํ๋ฐ์๋ ๊ฐ์ ํฉ๋๋ค. | |
| ### ๋ณด์กฐ ๊ณผ์ ์ schedule | |
| - **Publication:** 5๊ฐ๋ก ๊ตฌ๋ถ๋ ์๊ธฐ๋ฅผ ๊ธฐ์ค์ผ๋ก ํ์ฌ ํ๋์ ์ฐ์ ์๊ธฐ scalar๋ฅผ ์์ธกํฉ๋๋ค. ํ์ต ๊ฐ๋ฅํ ordered cutpoint, interval-aware NLL/Huber, chronological ranking์ ํจ๊ป ์ฌ์ฉํ๋ฉฐ class-balanced ์ ์ฉ batch๋ฅผ 4 step๋ง๋ค ํ์ตํ์ต๋๋ค. | |
| - **Kiwi stylometry:** ์ธ๊ฐ train split์์ window ๋จ์ ์ ๋ขฐ๋๋ก 16โ24๊ฐ ํน์ง์ ์ ํํ๊ณ , hidden 256 MLP๋ก ์ธ๊ฐยทAI window์ ํ์คํ๋ ์งํ๋ฅผ ํ๊ทํ์ต๋๋ค. ๋ฌธ์ฒด ๋ฐฉํฅ์ ์ก๋ ์ด๊ธฐ ์ ํธ๋ก ์ฌ์ฉํ ๋ค ๊ฐ์ ํ์ต๋๋ค. | |
| - **Human/AI:** ์ธ๊ฐ ๋ณธ๋ฌธ, counterfactual rewrite, synthetic fiction์ ์ถ์ฒ๋ณ ๊ท ํ ๊ธฐ์ฌ๋ก ํ์ตํ์ต๋๋ค. ๋ณด์กฐ head์ encoder gradient๋ 0.3๋ฐฐ๋ก ์ ํํ์ต๋๋ค. | |
| - **Optimization:** BF16, AdamW, LoRA LR `2e-5`, auxiliary head LR `8e-5`/`2e-4`, weight decay `0.01`, max gradient norm `50`; gradient checkpointing์ ์ฌ์ฉํ์ง ์์์ต๋๋ค. | |
| - **Ramps/fades:** hard-negative, counterfactual, decorrelation, synthetic, Human/AI loss๋ฅผ ramp๋ก ๋์ ํ์ต๋๋ค. Stylometry์ synthetic์ ์ด๊ธฐ ์ ๋ ํ ๊ฐ์ ํ๊ณ , decorrelation์ ์คํ๋ฐ guardrail๋ก ์ ์งํ์ต๋๋ค. | |
| ## ์ธ๋ถ ํ๊ฐ | |
| ### ํ๋กํ ์ฝ | |
| - ํ๊ตญ์ด ์ฅ๋ฅด์์ค **11 authors / 80 works / 640 segments** | |
| - ์ํ๋ง๋ค ๋ฌด์์ ์์น์์ ๋์ผํ๊ฒ 8๊ฐ ๊ตฌ๊ฐ ์ถ์ถ | |
| - ์ ๋ ฅ ๊ธธ์ด 1024 tokens, ๋ชจ๋ ๋ชจ๋ธ์ ๋์ผํ query/gallery ์ฌ์ฉ | |
| - ๋น๊ต ๋ชจ๋ธ: ์ํ์ ๋ฌด์์ ๊ธฐ๋๊ฐ, ์๋ณธ EmbeddingGemma 300M, ์ ์ธ๋ [`Baragi-AI/Munche-768`](https://huggingface.co/Baragi-AI/Munche-768), Munche-v2-768 | |
| - ์ด์ ํ๊ฐ ๋ฐ์ดํฐ์ Munche-768์ ํ์ต ๋ ธ์ถ์ด ํ์ธ๋์ด ํด๋น ๊ฒฐ๊ณผ๋ ํ๊ธฐํ๊ณ , ๋ณ๋์ ์์ ์๊ฐ ๋ง๋ญ์น์์ ๋ค์ ํ๋ณธ์ ์ถ์ถํ์ต๋๋ค. | |
| - ๋ฌด์์ ๊ฒฐ๊ณผ๋ ๋์ ์๋ฎฌ๋ ์ด์ ์ด ์๋๋ผ ์ค์ candidate/positive ์์ ๋ฐ๋ฅธ closed-form expectation์ ๋๋ค. | |
| | Metric | Random | EmbeddingGemma 300M | Munche-768 | **Munche-v2-768** | | |
| |---|---:|---:|---:|---:| | |
| | Same-work mAP | 0.0221 | 0.5680 | 0.7979 | **0.8233** | | |
| | Same-work Recall@1 | 0.0120 | 0.8328 | **0.9484** | **0.9484** | | |
| | Cross-work author mAP | 0.0882 | 0.1973 | 0.2960 | **0.3433** | | |
| | Cross-work author Recall@1 | 0.0794 | 0.3726 | 0.5302 | **0.6395** | | |
| | Cross-work author MRR | 0.2161 | 0.5163 | 0.6368 | **0.7205** | | |
| | N=3 prototype, 2 support works, macro top1 | 0.0909 | 0.4599 | 0.5064 | **0.6116** | | |
| | N=3 prototype, 3 support works, macro top1 | 0.0909 | 0.4981 | 0.5482 | **0.6205** | | |
| | Content-hard pairwise accuracy | 0.5000 | 0.0888 | 0.5719 | **0.6213** | | |
| | Content-hard top1 | 0.6998 | 0.3726 | 0.7412 | **0.7981** | | |
| `Content-hard`์ negative๋ ์๋ณธ EmbeddingGemma semantic space์์ ๊ฐ์ฅ ๊ฐ๊น์ด ๋ค๋ฅธ ์๊ฐ ๊ตฌ๊ฐ์ ๋๋ค. ๋ฐ๋ผ์ EmbeddingGemma ์์ฒด์ ๋ฎ์ content-hard ์ ์๋ ์ผ๋ฐ ์๋ฏธ ๊ฒ์ ์ฑ๋ฅ ์ ํ๋ฅผ ๋ปํ์ง ์์ผ๋ฉฐ, ๊ฐ์ semantic space๋ก ๊ณ ๋ฅธ ์๋์ ์ธ adversarial baseline์ ๋๋ค. Content-hard top1์ ๋ฌด์์ ๊ธฐ๋๊ฐ์ด ๋์ ๊ฒ์ query๋น same-author positive๊ฐ ๋ค์์ธ ๋ฐ๋ฉด hard negative๋ฅผ 20๊ฐ๋ก ์ ํํ๊ธฐ ๋๋ฌธ์ ๋๋ค. | |
| ### ์๊ฐ ๋จ์ paired bootstrap | |
| Munche-768 ๋๋น Munche-v2-768์ cross-work ์ฐจ์ด๋ฅผ ์๊ฐ๋ฅผ ํ๋ณธ ๋จ์๋ก 20,000ํ ๋ณต์์ถ์ถํ์ต๋๋ค. | |
| | Metric | Paired difference | 95% bootstrap CI | Better authors | | |
| |---|---:|---:|---:| | |
| | mAP | **+0.0472** | `[+0.0107, +0.0850]` | 8 / 11 | | |
| | Recall@1 | **+0.1093** | `[+0.0339, +0.1795]` | 9 / 11 | | |
| | MRR | **+0.0836** | `[+0.0221, +0.1427]` | 9 / 11 | | |
| ## ํด์๊ณผ ์ ํ์ฌํญ | |
| - Same-work retrieval์ ์ธ๋ฌผยท์ธ๊ณ๊ดยท์ฌ๊ฑด ๋จ์๋ฅผ ์ฌ์ฉํ ์ ์์ผ๋ฏ๋ก ๋ฌธ์ฒด ๋ ๋ฆฝ์ฑ์ ๋จ๋ ์ผ๋ก ์ฆ๋ช ํ์ง ์์ต๋๋ค. **Cross-work**, **prototype**, **content-hard** ์งํ๋ฅผ ์ฐ์ ํด์ ๋ณด์ธ์. | |
| - ๋ชจ๋ธ์ ํ๊ตญ์ด ์ฅ๋ฅด์์ค์ ํนํ๋์ด ์์ต๋๋ค. ๋น๋ฌธํ, ๋ฒ์ญ๋ฌธ, ์งง์ ๋ฌธ์ฅ, ์, ์ฑํ , ์์ด ๋ฑ์์๋ ์ฑ๋ฅ์ ๋ณด์ฅํ์ง ์์ต๋๋ค. | |
| - ๋ฌธ์ฒด ์ ์ฌ๋๋ ์ ์ ์ ์์ ๋ฒ์ ยท์ฌ์ค์ ์ฆ๊ฑฐ๊ฐ ์๋๋๋ค. ๊ณต๋ ์งํ, ํธ์ง, ์ฅ๋ฅด ๊ด์ต, ์๋, ํ๋ซํผ ๊ท์น, ์๋์ ๋ชจ๋ฐฉ์ ์ํฅ์ ๋ฐ์ ์ ์์ต๋๋ค. | |
| - Human/AI ๋ณด์กฐ ํ์ต์ ํน์ ์์ฑ ๋ชจ๋ธ๊ณผ ๋ฐ์ดํฐ ๋ถํฌ์ ์์กดํฉ๋๋ค. ์ด ์๋ฒ ๋ฉ์ ๋จ๋ AI ํ์ง๊ธฐ๋ก ์ฌ์ฉํ์ง ๋ง์ธ์. | |
| - ์ ์ ์ถ์ , ์ต๋ช ์ฌ์ฉ์ ์๋ณ, ํ์ ๋จ์ ๋ฑ ๊ฐ์ธ์๊ฒ ๋ถ์ด์ต์ ์ค ์ ์๋ ์ฉ๋์๋ ์ธ๊ฐ ๊ฒํ ์ ๋ณ๋ ๊ฒ์ฆ์ด ํ์ํฉ๋๋ค. | |
| ## ๋ผ์ด์ ์ค | |
| ์ด ๋ชจ๋ธ์ EmbeddingGemma ํ์ ๋ชจ๋ธ์ด๋ฉฐ **Gemma Terms of Use**์ **Gemma Prohibited Use Policy**๋ฅผ ๋ฐ๋ฆ ๋๋ค. ๋ฒ ์ด์ค ๋ชจ๋ธ ํ์ผ์ ๋ฐ์ผ๋ ค๋ฉด Hugging Face์์ Google์ ์ฌ์ฉ ์กฐ๊ฑด์ ๋์ํด์ผ ํ ์ ์์ต๋๋ค. ์์ธํ ๋ด์ฉ์ [EmbeddingGemma ๋ชจ๋ธ ์นด๋](https://huggingface.co/google/embeddinggemma-300m)๋ฅผ ํ์ธํ์ธ์. | |
| ## Citation | |
| EmbeddingGemma๋ฅผ ์ฌ์ฉํ๋ ๊ฒฝ์ฐ ์ ๋ชจ๋ธ ๋ ผ๋ฌธ์ ์ธ์ฉํ์ธ์. | |
| ```bibtex | |
| @article{embedding_gemma_2025, | |
| title = {EmbeddingGemma: Powerful and Lightweight Text Representations}, | |
| author = {Schechter Vera, Henrique and others}, | |
| year = {2025}, | |
| url = {https://arxiv.org/abs/2509.20354} | |
| } | |
| ``` | |