We have switched from the Regex tokenizer to the SentencePiece tokenizer. (The updated code was written by OpenSoftware-World and corrected by ChatGPT.)
Training code for the SentencePiece tokenizer for the OpenSoftware-World-OSW1 AI model. (This code was written by ChatGPT and edited by OpenSoftware-World.)