Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
TilQazyna 's Collections
First Releases (2024)
Kazakh Terminology
Kazakh Morphology and POS
Speech and OCR
Til Web, Crawls and Archive
Til Instruct — Task Datasets
Tokenizers
Til 256k Research Ladder
Kazakh GEC — All Models
Til — Multilingual Models
Til Core — Kazakh-only Models
Til Flagship Datasets

Til Flagship Datasets

updated about 8 hours ago

Core datasets for pretraining, instruction tuning, speech and language tasks. Основные датасеты для предобучения, инструкций, речи и языковых задач.

Upvote
-

  • TilQazyna/Til-Corpus

    Updated about 4 hours ago • 31 • 1

    Note Start here for pretraining text.


  • TilQazyna/Til-Instruct

    Viewer • Updated 1 day ago • 6.26M • 6

  • TilQazyna/Til-Parallel

    Updated 1 day ago • 7

  • TilQazyna/Til-Audio

    Viewer • Updated about 20 hours ago • 380k • 14

  • TilQazyna/Til-GEC

    Viewer • Updated about 7 hours ago • 4.62M • 23

    Note Training data behind every GEC model in this collection.


  • TilQazyna/Til-Morphology

    Updated 1 day ago • 14

  • TilQazyna/Til-Terminology

    Viewer • Updated 1 day ago • 317k • 12

  • TilQazyna/Til-Classification

    Viewer • Updated 1 day ago • 91.8k • 7
Upvote
-
  • Collection guide
  • Browse collections
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs