Buckets:
| ## Token Classification | |
| Token classification is a task in which a label is assigned to some tokens in a text. Some popular token classification subtasks are Named Entity Recognition (NER) and Part-of-Speech (PoS) tagging. | |
| > [!TIP] | |
| > For more details about the `token-classification` task, check out its [dedicated page](https://huggingface.co/tasks/token-classification)! You will find examples and related materials. | |
| ### Recommended models | |
| - [dslim/bert-base-NER](https://huggingface.co/dslim/bert-base-NER): A robust performance model to identify people, locations, organizations and names of miscellaneous entities. | |
| - [FacebookAI/xlm-roberta-large-finetuned-conll03-english](https://huggingface.co/FacebookAI/xlm-roberta-large-finetuned-conll03-english): A strong model to identify people, locations, organizations and names in multiple languages. | |
| - [blaze999/Medical-NER](https://huggingface.co/blaze999/Medical-NER): A token classification model specialized on medical entity recognition. | |
| Explore all available models and find the one that suits you best [here](https://huggingface.co/models?inference=warm&pipeline_tag=token-classification&sort=trending), or from the terminal with the [`hf` CLI](https://huggingface.co/docs/huggingface_hub/package_reference/cli#hf-models-list): | |
| ```bash | |
| hf models ls --warm --pipeline-tag token-classification --sort trending_score | |
| ``` | |
| ### Using the API | |
| <InferenceSnippet | |
| pipeline=token-classification | |
| providersMapping={ {"hf-inference":{"modelId":"rizzoaiacademy/rizzo-pii-0.3B","providerModelId":"rizzoaiacademy/rizzo-pii-0.3B"}} } | |
| /> | |
| ### API specification | |
| #### Request | |
| | Headers | | | | |
| | :--- | :--- | :--- | | |
| | **authorization** | _string_ | Authentication header in the form `'Bearer: hf_****'` when `hf_****` is a personal user access token with "Inference Providers" permission. You can generate one from [your settings page](https://huggingface.co/settings/tokens/new?ownUserPermissions=inference.serverless.write&tokenType=fineGrained). | | |
| | Payload | | | | |
| | :--- | :--- | :--- | | |
| | **inputs*** | _string_ | The input text data | | |
| | **parameters** | _object_ | | | |
| | ** ignore_labels** | _string[]_ | A list of labels to ignore | | |
| | ** stride** | _integer_ | The number of overlapping tokens between chunks when splitting the input text. | | |
| | ** aggregation_strategy** | _string_ | One of the following: | | |
| | ** (#1)** | _'none'_ | Do not aggregate tokens | | |
| | ** (#2)** | _'simple'_ | Group consecutive tokens with the same label in a single entity. | | |
| | ** (#3)** | _'first'_ | Similar to "simple", also preserves word integrity (use the label predicted for the first token in a word). | | |
| | ** (#4)** | _'average'_ | Similar to "simple", also preserves word integrity (uses the label with the highest score, averaged across the word's tokens). | | |
| | ** (#5)** | _'max'_ | Similar to "simple", also preserves word integrity (uses the label with the highest score across the word's tokens). | | |
| #### Response | |
| | Body | | | |
| | :--- | :--- | :--- | | |
| | **(array)** | _object[]_ | Output is an array of objects. | | |
| | ** entity_group** | _string_ | The predicted label for a group of one or more tokens | | |
| | ** entity** | _string_ | The predicted label for a single token | | |
| | ** score** | _number_ | The associated score / probability | | |
| | ** word** | _string_ | The corresponding text | | |
| | ** start** | _integer_ | The character position in the input where this group begins. | | |
| | ** end** | _integer_ | The character position in the input where this group ends. | | |
Xet Storage Details
- Size:
- 4.49 kB
- Xet hash:
- 3c8be8a9ff8430cf7ec8193993ac812ac2d6636efcd40a7b43327b83362397d5
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.