initial release

Browse files

Files changed (7) hide show

README.md +45 -0
config.json +0 -0
pytorch_model.bin +3 -0
special_tokens_map.json +51 -0
supar.model +3 -0
tokenizer.json +0 -0
tokenizer_config.json +55 -0

README.md ADDED Viewed

	@@ -0,0 +1,45 @@

+---
+language:
+- "ko"
+tags:
+- "korean"
+- "token-classification"
+- "pos"
+- "dependency-parsing"
+datasets:
+- "universal_dependencies"
+license: "apache-2.0"
+pipeline_tag: "token-classification"
+widget:
+- text: "홍시 맛이 나서 홍시라 생각한다."
+- text: "紅柹 맛이 나서 紅柹라 生覺한다."
+---
+# deberta-base-korean-morph-upos
+## Model Description
+This is a RoBERTa model pre-trained on Korean texts for POS-tagging and dependency-parsing, derived from [deberta-v3-base-korean](https://huggingface.co/team-lucid/deberta-v3-base-korean) and [morphUD-korean](https://github.com/jungyeul/morphUD-korean). Every morpheme (형태소) is tagged by [UPOS](https://universaldependencies.org/u/pos/)(Universal Part-Of-Speech).
+## How to Use
+```py
+from transformers import AutoTokenizer,AutoModelForTokenClassification,TokenClassificationPipeline
+tokenizer=AutoTokenizer.from_pretrained("KoichiYasuoka/deberta-base-korean-morph-upos")
+model=AutoModelForTokenClassification.from_pretrained("KoichiYasuoka/deberta-base-korean-morph-upos")
+pipeline=TokenClassificationPipeline(tokenizer=tokenizer,model=model,aggregation_strategy="simple")
+nlp=lambda x:[(x[t["start"]:t["end"]],t["entity_group"]) for t in pipeline(x)]
+print(nlp("홍시 맛이 나서 홍시라 생각한다."))
+```
+or
+```py
+import esupar
+nlp=esupar.load("KoichiYasuoka/deberta-base-korean-morph-upos")
+print(nlp("홍시 맛이 나서 홍시라 생각한다."))
+```
+## See Also
+[esupar](https://github.com/KoichiYasuoka/esupar): Tokenizer POS-tagger and Dependency-parser with BERT/RoBERTa/DeBERTa models

config.json ADDED Viewed

The diff for this file is too large to render. See raw diff

pytorch_model.bin ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:0e3a83ab25f9059b283b593eb4ac9a0450b2bf11d949dd6741aca05a066b8f86
+size 542314863

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,51 @@

+{
+  "bos_token": {
+    "content": "[CLS]",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "cls_token": {
+    "content": "[CLS]",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "eos_token": {
+    "content": "[SEP]",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "mask_token": {
+    "content": "[MASK]",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "pad_token": {
+    "content": "[PAD]",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "sep_token": {
+    "content": "[SEP]",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "unk_token": {
+    "content": "[UNK]",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  }
+}

supar.model ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:23ba0acd11e8177739087bd90d1e18fc3cd32c4188c3206e8b5e1f92ab3a1a8d
+size 587550695

tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,55 @@

+{
+  "added_tokens_decoder": {
+    "0": {
+      "content": "[PAD]",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "1": {
+      "content": "[CLS]",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "2": {
+      "content": "[SEP]",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "3": {
+      "content": "[UNK]",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "64000": {
+      "content": "[MASK]",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "bos_token": "[CLS]",
+  "clean_up_tokenization_spaces": true,
+  "cls_token": "[CLS]",
+  "do_lower_case": false,
+  "eos_token": "[SEP]",
+  "mask_token": "[MASK]",
+  "pad_token": "[PAD]",
+  "sep_token": "[SEP]",
+  "split_by_punct": false,
+  "tokenizer_class": "DebertaV2TokenizerFast",
+  "unk_token": "[UNK]"
+}