raydotac/OmniVoice-bucket / tokenizer_config.json
raydotac's picture
download
raw
533 Bytes
{
"add_prefix_space": false,
"backend": "tokenizers",
"bos_token": null,
"clean_up_tokenization_spaces": false,
"eos_token": "<|im_end|>",
"errors": "replace",
"extra_special_tokens": [
"<|denoise|>",
"<|lang_start|>",
"<|lang_end|>",
"<|instruct_start|>",
"<|instruct_end|>",
"<|text_start|>",
"<|text_end|>"
],
"is_local": true,
"model_max_length": 131072,
"pad_token": "<|endoftext|>",
"split_special_tokens": false,
"tokenizer_class": "Qwen2Tokenizer",
"unk_token": null
}

Xet Storage Details

Size:
533 Bytes
·
Xet hash:
5eb94f90374f1689413001a4c2487b4353a2ce0400a643c24046138df6014e4e

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.