| --- |
| title: Til-Qazyna |
| emoji: 📚 |
| colorFrom: indigo |
| colorTo: blue |
| sdk: static |
| pinned: false |
| --- |
| |
|  |
|
|
| > **Қазақша.** Ш. Шаяхметов атындағы «Тіл-Қазына» ұлттық ғылыми-практикалық орталығы қазақ тіліне арналған деректер мен модельдерді әзірлейді. Бұл парақта мәтін, сөйлеу, OCR, морфология, терминология және тілдік модель репозиторийлері жинақталған. |
|
|
| > **По-русски.** Национальный научно-практический центр «Тіл-Қазына» имени Ш. Шаяхметова выпускает данные и модели для казахского языка. Здесь собраны репозитории текстов, речи, OCR, морфологии, терминологии и языковых моделей. |
|
|
| The National Scientific and Practical Centre 'Til-Qazyna' named after Sh. Shayakhmetov (**Ш. Шаяхметов атындағы «Тіл-Қазына» ұлттық ғылыми-практикалық орталығы**) develops datasets and models for Kazakh. Its repositories link source corpora, curated datasets, tokenizers, base models and task-specific fine-tunes, allowing researchers to trace training data and developers to work with the released resources. |
|
|
| ## What's here |
|
|
| | Models | Datasets | Data on the Hub | |
| |---:|---:|---:| |
| | 42 | 57 | 2.2 TB | |
|
|
| The total includes all branches and revision history. Repository cards are available in Kazakh, Russian and English. |
|
|
| ## Flagship datasets |
|
|
| | Dataset | Records | Size | Contents | |
| |---|---:|---:|---| |
| | [Til-Corpus](https://huggingface.co/datasets/TilQazyna/Til-Corpus) | 58,850,639 | 204.20 GB | Tiered multilingual text for language-model pretraining. | |
| | [Til-Instruct](https://huggingface.co/datasets/TilQazyna/Til-Instruct) | 6,259,769 | 8.35 GB | Judged instruction records grouped into raw, clean and premium configs. | |
| | [Til-Parallel](https://huggingface.co/datasets/TilQazyna/Til-Parallel) | 23,294,865 | 72.16 GB | Tiered parallel text centered on Kazakh. | |
| | [Til-Books](https://huggingface.co/datasets/TilQazyna/Til-Books) | 16,847 books | 2.02 GB | Full-text books extracted from digital files or recovered with OCR. | |
| | [Til-Audio](https://huggingface.co/datasets/TilQazyna/Til-Audio) | 380,068 | 249.77 GB | Speech and transcripts in the downloadable config for ASR, TTS and audio tasks. | |
| | [Til-GEC](https://huggingface.co/datasets/TilQazyna/Til-GEC) | 4,624,548 | 1.23 GB | Error-to-correction pairs for Kazakh GEC training. | |
| | [Til-Morphology](https://huggingface.co/datasets/TilQazyna/Til-Morphology) | 3,767,518 | 207.8 MB | Kazakh words with morpheme segmentation and context. | |
| | [Til-Terminology](https://huggingface.co/datasets/TilQazyna/Til-Terminology) | 317,277 | 5.7 MB | Judged terminology records across subject areas. | |
| | [Til-Classification](https://huggingface.co/datasets/TilQazyna/Til-Classification) | 91,766 | 15.1 MB | Labeled text for classification tasks. | |
|
|
| ## Model families |
|
|
| ### Til Core |
|
|
| Kazakh-only models trained from scratch with a 256,000-token morpheme-aware BPE vocabulary. |
|
|
| | Model | Size | Role | |
| |---|---:|---| |
| | [Til-Core-0.5B](https://huggingface.co/TilQazyna/Til-Core-0.5B) | 2.02 GB | Base model; Qwen2, 18 layers and 32,768-token context. | |
| | [Til-Core-0.5B-Instruct](https://huggingface.co/TilQazyna/Til-Core-0.5B-Instruct) | 2.02 GB | Instruction fine-tune of the 0.5B base with the same architecture and context. | |
| | [Til-Core-0.5B-QazGramma-GEC](https://huggingface.co/TilQazyna/Til-Core-0.5B-QazGramma-GEC) | 1.02 GB | GEC fine-tune of the 0.5B base for grammar, spelling and punctuation correction. | |
| | [Til-Core-1B](https://huggingface.co/TilQazyna/Til-Core-1B) | 7.11 GB | Base model; Llama architecture, 16 layers and 2,048-token context. | |
| | [Til-Core-1B-Instruct](https://huggingface.co/TilQazyna/Til-Core-1B-Instruct) | 3.57 GB | Instruction version of the 1B base for Kazakh chat and tasks. | |
| | [Til-Core-1B-GEC](https://huggingface.co/TilQazyna/Til-Core-1B-GEC) | 5.01 GB | GEC version of the 1B base for minimal-edit correction. | |
|
|
| ### Til multilingual |
|
|
| Kazakh-first models covering Russian, English, code and mathematics with a 131,072-token vocabulary. |
|
|
| | Model | Size | Role | |
| |---|---:|---| |
| | [Til-mini-1B](https://huggingface.co/TilQazyna/Til-mini-1B) | 3.84 GB | Base model; DeepSeek-V3 architecture, 24 layers and 2,048-token context. | |
| | [Til-mini-1B-GEC](https://huggingface.co/TilQazyna/Til-mini-1B-GEC) | 1.93 GB | GEC fine-tune of Til-mini-1B. | |
| | [Til-2B](https://huggingface.co/TilQazyna/Til-2B) | 7.92 GB | Base model; DeepSeek-V3 architecture, 30 layers and 4,096-token context. | |
| | [Til-2B-instruct](https://huggingface.co/TilQazyna/Til-2B-instruct) | 3.97 GB | Instruction fine-tune of Til-2B for chat, QA, writing and rewriting. | |
| | [Til-2B-GEC](https://huggingface.co/TilQazyna/Til-2B-GEC) | 3.97 GB | GEC fine-tune of Til-2B for Kazakh correction. | |
| | [Til-0.5B-multilingual-base](https://huggingface.co/TilQazyna/Til-0.5B-multilingual-base) | 952.8 MB | Base model with 24 layers and 4,096-token context. | |
| | [Til-1B-multilingual-base](https://huggingface.co/TilQazyna/Til-1B-multilingual-base) | 1.89 GB | 1B base model with 24 layers and 4,096-token context. | |
| | [Til-1B-multilingual-base-GEC](https://huggingface.co/TilQazyna/Til-1B-multilingual-base-GEC) | 1.89 GB | GEC fine-tune of the 1B multilingual base. | |
|
|
| ### Til 256k research ladder |
|
|
| The exp080–exp085 sequence follows a Kazakh model through pretraining, instruction tuning and GEC. |
|
|
| | Repository | Size | Role | |
| |---|---:|---| |
| | [Til-kk-0.5B-256k-exp080](https://huggingface.co/TilQazyna/Til-kk-0.5B-256k-exp080) | 7.55 GB per checkpoint | Pretrained base; weights are in `epoch-1`, `epoch-2` and `epoch-3`, while `main` has no weights. | |
| | [Til-kk-0.5B-256k-sft-exp081](https://huggingface.co/TilQazyna/Til-kk-0.5B-256k-sft-exp081) | 4.75 GB per checkpoint | Instruction-tuned model; weights are in `epoch-1`, `epoch-2` and `epoch-3`, while `main` has no weights. | |
| | [Til-kk-0.5B-256k-gec-exp082](https://huggingface.co/TilQazyna/Til-kk-0.5B-256k-gec-exp082) | 4.75 GB per checkpoint | GEC model; weights are in `epoch-1`, `epoch-2` and `epoch-3`, while `main` has no weights. | |
| | [Til-kk-0.5B-256k-gec-rlaif-exp083](https://huggingface.co/TilQazyna/Til-kk-0.5B-256k-gec-rlaif-exp083) | 4.75 GB | RLAIF-trained GEC model; weights are in `epoch-1`, while `main` has no weights. | |
| | [Til-kk-0.5B-256k-gec-exp085](https://huggingface.co/TilQazyna/Til-kk-0.5B-256k-gec-exp085) | 1.72 GB | Published GEC model; DeepSeek-V3 architecture, 24 layers and 4,096-token context. | |
| | [Til-Corpus-exp078](https://huggingface.co/datasets/TilQazyna/Til-Corpus-exp078) | 139.07 GB | Corpus snapshot used alongside the experimental ladder. | |
|
|
| ## Collections |
|
|
| - [Til Flagship Datasets](https://huggingface.co/collections/TilQazyna/til-flagship-datasets-6a77766b2c61823dc5081b4e) — datasets for pretraining, instruction tuning, speech and language tasks; 9 repositories. |
| - [Til Core — Kazakh-only models](https://huggingface.co/collections/TilQazyna/til-core-kazakh-only-models-6a77767061d8d763891080c0) — base, instruct and GEC models for Kazakh; 6 repositories. |
| - [Til — multilingual models](https://huggingface.co/collections/TilQazyna/til-multilingual-models-6a7776bafb17f051a0041d0c) — Kazakh-first base models and task fine-tunes; 8 repositories. |
| - [Kazakh GEC — all models](https://huggingface.co/collections/TilQazyna/kazakh-gec-all-models-6a7776be500ada5b47c19c24) — correction models and their training dataset; 8 repositories. |
| - [Til 256k research ladder](https://huggingface.co/collections/TilQazyna/til-256k-research-ladder-6a7776c208f45d11a08da060) — checkpoints and data from the exp080–exp085 sequence; 6 repositories. |
| - [Tokenizers](https://huggingface.co/collections/TilQazyna/tokenizers-6a7776c5cc63232d1ef78d7d) — tokenizers and vocabulary data for model training; 2 repositories. |
| - [Til Instruct — task datasets](https://huggingface.co/collections/TilQazyna/til-instruct-task-datasets-6a7776c78558be3e6c2fe201) — instruction data for QA, summarization, classification and generation; 18 repositories. |
| - [Til Books — sources and scans](https://huggingface.co/collections/TilQazyna/til-books-sources-and-scans-6a7776d01da4dbd2b4777352) — extracted books, scans and source-specific subsets; 6 repositories. |
| - [Til Web, crawls and archive](https://huggingface.co/collections/TilQazyna/til-web-crawls-and-archive-6a7776d3936ff367bd5612ee) — web text, crawls and archived sources for pretraining corpora; 11 repositories. |
| - [Speech and OCR](https://huggingface.co/collections/TilQazyna/speech-and-ocr-6a7776d95c08c96b473af1d6) — audio data, ASR models and OCR tools; 6 repositories. |
| - [Kazakh morphology and POS](https://huggingface.co/collections/TilQazyna/kazakh-morphology-and-pos-6a7776dcdb40d73bb3ceae65) — models and annotated data for segmentation, morphology and POS tagging; 16 repositories. |
| - [Kazakh terminology](https://huggingface.co/collections/TilQazyna/kazakh-terminology-6a7776e4a1416f6f9c344a1e) — terminology extraction models, datasets and evaluations; 9 repositories. |
| - [First releases (2024)](https://huggingface.co/collections/TilQazyna/first-releases-2024-6a7776e99ac03965e0667217) — early language and speech models; 3 repositories. |
|
|
| ## Start here |
|
|
| **Run an instruction model** |
|
|
| ```python |
| from transformers import pipeline |
| generate = pipeline("text-generation", model="TilQazyna/Til-2B-instruct") |
| print(generate("Қазақ тілінде қысқаша жауап бер:", max_new_tokens=80)[0]["generated_text"]) |
| ``` |
|
|
| **Stream the pretraining corpus** |
|
|
| ```python |
| from datasets import load_dataset |
| ds = load_dataset("TilQazyna/Til-Corpus", "premium", split="train", streaming=True) |
| print(next(iter(ds))) |
| ``` |
|
|
| **Prepare corpus records with the released tokenizer** |
|
|
| ```python |
| from datasets import load_dataset |
| from transformers import AutoTokenizer |
| ds = load_dataset("TilQazyna/Til-Corpus", "premium", split="train", streaming=True) |
| tokenizer = AutoTokenizer.from_pretrained("TilQazyna/Til-Tokenizer-128k") |
| tokenized = ds.map(lambda row: tokenizer(row["text"], truncation=True)) |
| ``` |
|
|
| ## Access |
|
|
| All repositories require access approval. Cards and file lists are public; files can be downloaded after the TilQazyna team approves a request submitted through **Request access**. |
|
|
| ## Citation |
|
|
| ```bibtex |
| @misc{tilqazyna2026, |
| author = {{Ш. Шаяхметов атындағы «Тіл-Қазына» ұлттық ғылыми-практикалық орталығы}}, |
| title = {Til-Qazyna on Hugging Face}, |
| year = {2026}, |
| url = {https://huggingface.co/TilQazyna} |
| } |
| ``` |
|
|