3.78 MB

Ctrl+K

1 contributor

Add SPE tokenizer (1125 SMILES subword tokens) + <|start_of_smiles|>/<|end_of_smiles|> special tokens. Trained on 2M ZINC20 + 2M ChEMBL canonical SMILES. SPE vocab_size=1000, min_freq=4000.

f6eb1d8 verified 2 months ago

.gitattributes

1.52 kB
initial commit 2 months ago
README.md

5.17 kB
Add SPE tokenizer (1125 SMILES subword tokens) + <|start_of_smiles|>/<|end_of_smiles|> special tokens. Trained on 2M ZINC20 + 2M ChEMBL canonical SMILES. SPE vocab_size=1000, min_freq=4000. 2 months ago
tokenizer.json

3.77 MB
Add SPE tokenizer (1125 SMILES subword tokens) + <|start_of_smiles|>/<|end_of_smiles|> special tokens. Trained on 2M ZINC20 + 2M ChEMBL canonical SMILES. SPE vocab_size=1000, min_freq=4000. 2 months ago
tokenizer_config.json

480 Bytes
Add SPE tokenizer (1125 SMILES subword tokens) + <|start_of_smiles|>/<|end_of_smiles|> special tokens. Trained on 2M ZINC20 + 2M ChEMBL canonical SMILES. SPE vocab_size=1000, min_freq=4000. 2 months ago