Buckets:
Token Classification
Token classification is a task in which a label is assigned to some tokens in a text. Some popular token classification subtasks are Named Entity Recognition (NER) and Part-of-Speech (PoS) tagging.
For more details about the
token-classificationtask, check out its dedicated page! You will find examples and related materials.
Recommended models
- dslim/bert-base-NER: A robust performance model to identify people, locations, organizations and names of miscellaneous entities.
- FacebookAI/xlm-roberta-large-finetuned-conll03-english: A strong model to identify people, locations, organizations and names in multiple languages.
- blaze999/Medical-NER: A token classification model specialized on medical entity recognition.
Explore all available models and find the one that suits you best here, or from the terminal with the hf CLI:
hf models ls --warm --pipeline-tag token-classification --sort trending_score
Using the API
<InferenceSnippet pipeline=token-classification providersMapping={ {"hf-inference":{"modelId":"rizzoaiacademy/rizzo-pii-0.3B","providerModelId":"rizzoaiacademy/rizzo-pii-0.3B"}} } />
API specification
Request
| Headers | ||
|---|---|---|
| authorization | string | Authentication header in the form 'Bearer: hf_****' when hf_**** is a personal user access token with "Inference Providers" permission. You can generate one from your settings page. |
| Payload | ||
|---|---|---|
| inputs* | string | The input text data |
| parameters | object | |
| ignore_labels | string[] | A list of labels to ignore |
| stride | integer | The number of overlapping tokens between chunks when splitting the input text. |
| aggregation_strategy | string | One of the following: |
| (#1) | 'none' | Do not aggregate tokens |
| (#2) | 'simple' | Group consecutive tokens with the same label in a single entity. |
| (#3) | 'first' | Similar to "simple", also preserves word integrity (use the label predicted for the first token in a word). |
| (#4) | 'average' | Similar to "simple", also preserves word integrity (uses the label with the highest score, averaged across the word's tokens). |
| (#5) | 'max' | Similar to "simple", also preserves word integrity (uses the label with the highest score across the word's tokens). |
Response
| Body | | | :--- | :--- | :--- | | (array) | object[] | Output is an array of objects. | | entity_group | string | The predicted label for a group of one or more tokens | | entity | string | The predicted label for a single token | | score | number | The associated score / probability | | word | string | The corresponding text | | start | integer | The character position in the input where this group begins. | | end | integer | The character position in the input where this group ends. |
Xet Storage Details
- Size:
- 4.49 kB
- Xet hash:
- 3c8be8a9ff8430cf7ec8193993ac812ac2d6636efcd40a7b43327b83362397d5
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.