tokenizer-lm Copyright 2026 Clara Meister This product includes software developed as part of the tokenizer-lm study, licensed under the Apache License, Version 2.0. See LICENSE. -------------------------------------------------------------------------------- Third-party components -------------------------------------------------------------------------------- nanochat https://github.com/karpathy/nanochat MIT License, Copyright (c) 2025 Andrej Karpathy The model architecture, optimizer and dataloader come from nanochat. This project uses a patched copy (https://github.com/cimeister/nanochat, branch tokenizer-lm-patches) that adds a configurable value-embedding dimension, a configurable logit soft cap, and a configurable learning-rate width exponent for muP. The released model class scripts/hf_release/nanochat_model.py reimplements that architecture for the HuggingFace transformers interface. The full MIT license text is at nanochat/LICENSE. -------------------------------------------------------------------------------- Training data -------------------------------------------------------------------------------- The released model weights were trained on a mixture of: FineWeb-Edu HuggingFaceFW/fineweb-edu ODC-By FineWeb-2 (filtered) HuggingFaceFW/fineweb-2 ODC-By FineMath 4plus HuggingFaceTB/finemath ODC-By StarCoderData bigcode/starcoderdata listed as "other" Licenses as stated on each dataset's Hugging Face card, read 2026-08-17. The multilingual and code portions were used through quality-filtered derivatives prepared inside the Swiss AI initiative; see EXPERIMENTAL_SETUP.md for the exact paths and filters. -------------------------------------------------------------------------------- Tokenizers not redistributed -------------------------------------------------------------------------------- Two off-the-shelf tokenizers were used as reference points and are cited rather than redistributed: the Apertus tokenizer (swiss-ai/Apertus-70B-2509, the same artifact as Mistral-Nemo's tokenizer) and the Llama-3 tokenizer (recorded at training time as NousResearch/Meta-Llama-3-8B). In the names of the tokenizers this project trained, "gpt4o", "claude" and "gpt2" refer to a borrowed pretokenization regular expression, not to those vendors' tokenizers.