pasketti-word 🍝

pasketti-word is an automatic speech recognition (ASR) model designed specifically for child speech. pasketti-word generates orthographic (per word) transcriptions from audio files.

Audio in, transcript out

🐣 Existing open-source ASR models trained on adult speech perform poorly on children because child speech is fundamentally different from adult speech. Children are still learning to produce speech sounds and developing fine motor skills. Common child speech errors include:

  • Metathesis: "elephant" → "ephelant"
  • Velar fronting: "cup" → "tup"
  • Syllable deletion: "banana" → "nana"
  • Combinations of errors: "spaghetti" → "pasketti" 🍝

🔉 pasketti-word is fine-tuned on a large corpus of child speech data including 372 hours of child speech spanning >3,600 different children. This is the largest dataset of its kind for use in developing and testing open ASR models for kids. The model is designed to advance early education assessments, teaching tools, and needs screening.

📈 pasketti-word achieves an overall word error rate (WER) of 17% on held-out evaluation data. When compared to leading open ASR models (Whisper and KidWhisper) on the same evaluation set, pasketti-word demonstrates improved performance across all demographic groups, cutting the error rate by more than half for children ages 3-7 and achieving close to adult-level performance for children ages 5 and above.

In this model card:

  1. Quickstart
  2. Model Details
  3. Uses
  4. Evaluation
    1. Performance by Subset
    2. Bias and Limitations
  5. Training Data

Quickstart

Prerequisites

  1. Hardware. Inference runs through vLLM. See the vLLM installation docs for supported hardware and operating systems. This model has been tested only on Linux with NVIDIA GPUs.

  2. Environment. Install qwen_asr with its vLLM extra:

    pip install "qwen_asr[vllm]==0.0.6"
    

    pip install transformers alone cannot load this model. The qwen3_asr architecture is not built into transformers; it is registered when qwen_asr is imported.

    This release was verified against qwen_asr==0.0.6, vllm==0.14.0, transformers==4.57.6 and torch==2.9.1, on Python 3.13. qwen_asr pins an exact transformers version, so if newer releases break the install, pin these versions to reconstruct a working environment.

  3. Data. Pass each clip to asr.transcribe(audio=...) in one of these forms:

    • a local file path, an https URL, or a base64-encoded audio string, which qwen_asr reads directly, or
    • a (waveform, sample_rate) tuple, where waveform is a 1-D (mono) or 2-D (multi-channel) NumPy array and sample_rate is its true sample rate in Hz. To match how the model was trained, load the audio with librosa, as in the example below.

qwen_asr converts every input to 16 kHz mono before transcription.

Inference

Pass the Hub repo id, drivendata/pasketti-word, to Qwen3ASRModel.LLM. The weights download automatically the first time the model loads.

Transcribe a single clip:

from qwen_asr import Qwen3ASRModel
import librosa

asr = Qwen3ASRModel.LLM(
    "drivendata/pasketti-word",
    gpu_memory_utilization=0.9,
    max_inference_batch_size=32,
    max_new_tokens=256,
)
wav, _ = librosa.load("clip.flac", sr=16000, mono=True)
result = asr.transcribe(audio=[(wav, 16000)], language="English")
print(result[0].text)

>> "the cat sat down"

For large datasets, load and transcribe the audio in chunks (for example, a few thousand clips at a time) rather than holding every clip in memory at once.

qwen_asr also offers a transformers backend through Qwen3ASRModel.from_pretrained(). That backend has not been evaluated with this model and is not supported.


Model Details

This model is a supervised fine-tune of Qwen3-ASR-1.7B. Qwen3-ASR pairs an audio encoder with a Qwen3 language model, which writes out the transcript one token at a time. All of the model's weights were updated during fine-tuning, using clips of up to 20 seconds of children's speech with orthographic (word-level) transcripts.

The model was trained for two epochs with a global batch size of 32 (4 per GPU across 8 A100s), using AdamW with a peak learning rate of 3e-5. The schedule used linear warmup over the first 2% of steps, followed by linear decay sized for five epochs.

The model design is based on the results of the On Top of Pasketti Challenge, in which participants competed to build child ASR models. All five prize-winning Word Track solutions fine-tuned Qwen3-ASR-1.7B, and pasketti-word adapts the 3rd-place solution, a full-parameter fine-tune.

Uses

This model is intended to be used for automatic speech recognition in service of improving educational and developmental outcomes for children.

Based on the model's performance, it is most suitable for downstream use cases that depend on child transcripts as an input. These downstream use cases often do not require exact transcripts to effectively inform decision-making.

For use cases that require exact transcription from child speech, model outputs should be human reviewed.

Downstream Use Cases

Example downstream use cases include:

  • Needs screening: Identifying children who may be struggling and better targeting additional support. For example, identifying children behind in literacy or who may have a speech pathology.
  • Formative assessment: Monitoring student learning based on naturally collected classroom speech data.
  • Educator feedback: Enabling better teacher or tutor support and growth by incorporating spoken interactions.

The model could be plugged into decision-making processes as is to generate transcripts for human review, or fine-tuned to predict in-scope outcomes directly. Intended users include education researchers, education technology developers, teachers, and school administrators.

Out-of-Scope Use

The model should not be used in ways that do not serve the interests of children whose speech it processes. Out-of-scope uses include:

  • Surveillance or profiling of children, such as building voice or behavioral profiles for advertising, identifying individual speakers, or monitoring children without informed consent from them and their guardians.
  • Engagement maximization, such as designing products that use speech signals to drive compulsive use or unhealthy dependence on technology.
  • Fully automated high-stakes decisions, such as diagnosis or educational placement, without review by a qualified educator or clinician.

Evaluation

pasketti-word performance: 17% word error rate (WER)

WER by model

The evaluation set includes 148 hours of child speech from >2,300 different children as young as 3. All models are scored on every utterance with Whisper's English normalizer.

The model is compared to existing state-of-the-art open ASR models openai/whisper-medium and aadel4/kid-whisper-medium-en-myst (Whisper fine-tuned on a small amount of child speech data).

All models were run with loop protection:

  • Whisper and KidWhisper are decoded greedily with the compression-ratio fallback from OpenAI's transcribe(). Retrying low-confidence predictions is not implemented.
  • pasketti-word implements loop protection by default. It uses qwen_asr's default decoding, which collapses repeats on every output, with a 256-token limit.

About Word Error Rate

Word Error Rate (WER) measures the minimum number of word-level substitutions (𝑆), deletions (𝐷), and insertions (𝐼) required to transform the predicted transcript into the reference transcript, divided by the total number of words in the reference (𝑁). WER indicates roughly what percent of words predicted are wrong.

WER=Substitutions+Deletions+InsertionsWords=S+D+IN \text{WER} = \frac{\text{Substitutions} + \text{Deletions} + \text{Insertions}}{\text{Words}} = \frac{S + D + I}{N}

A limitation of WER is that it only measures whether the predicted word is exactly correct — it does not reward similar predictions that are close to correct. As a result, WER often slightly underestimates how well a model can perform in a real-world setting.

Performance by Subset

The model was evaluated for bias in performance based on age, race, speech development pathology, sex, and source corpus. As expected, performance tends to improve for older children compared to younger children. However, performance gains are pronounced for younger children, with WER for children ages 3-4 improving by 30 percentage points compared to Whisper.

Compared to Whisper, pasketti-word decreases key disparities between white vs. non-white children as well as between children with speech pathologies vs. children without.

Performance by Age

WER by age

Whisper performance is likely poor on ages 12+ because those samples are derived from particularly challenging corpuses.

Representation in the evaluation set:

Age Utterances Hours Children
3-4 54,341 25.2 190
5-7 15,677 13 681
8-11 29,799 37.8 278
12+ 5,320 1.4 44
Unknown 78,735 70.5 1,148

Performance by Pathology

WER by speech pathology

Representation in the evaluation set:

Pathology Utterances Hours Children
Atypical development 38,005 14.5 175
Typical 41,891 23.8 154
Unknown 103,976 109.6 1,996

Performance by Race

Racial groups are compared within each individual corpus to remove any correlation between corpus difficulty and race that influences overall race-based performance. This is necessary because very few corpora have race populated.

WER by race and corpus

Representation in the evaluation set:

CorpusRaceUtterancesHoursChildren
SPROUTBlack19,3608.867
Hispanic1,6370.76
Multi-racial13,2855.347
White12,4575.745
ReadNetBlack3480.347
Multi-racial2340.230
White1,8891.8176

Performance by Setting

WER by setting

Representation in the evaluation set:

Setting Utterances Hours Children
In-classroom 12,261 6.3 NA
Not classroom 171,611 141.7 2,324

Performance by Sex

WER by sex

Representation in the evaluation set:

Sex Utterances Hours Children
Female 47,714 26 379
Male 43,731 22.3 336
Unknown 92,427 99.6 1,610

Performance by Duration

WER by utterance duration

Representation in the evaluation set:

Utterance duration (seconds) Utterances Hours Children
<1 34,831 7.3 553
[1-3) 88,682 44.7 2,052
[3-5) 36,962 40.2 1,991
[5-10) 18,942 34 1,941
[10-20) 3,330 12.6 278
[20-60) 1,125 9.1 193

Performance by Corpus

Performance is heavily influenced by corpus due to corpus-specific age ranges, data collection methods, and audio quality. Two corpora were put entirely in the test set: CSLU: Kids' Speech Version 1.1 and the CMU Kids Corpus.

WER by corpus

Representation in the evaluation set:

Corpus Utterances Hours Children
Arizona Child Acoustic Database Repository 497 0.8 5
CMU Kids Corpus† 2,477 3.7 74
CSLU: Kids' Speech Version 1.1† 64,430 62.8 1,117
Cameron 1,326 0.7 5
Edmonton Narrative Norms Instrument 2,415 2.6 34
Ellis Weismer Corpus 1,775 1.2 5
JIBO Kids 1,926 1.4 28
My Science Tutor 12,607 27.6 91
Ohio Child Speech Corpus 12,271 8.8 30
Other 12,261 6.3 NA
PERCEPT-GFTA 4,430 1 77
PERCEPT-R 11,836 3.5 29
ReadNet 5,329 5.1 661
Speech Production Repository for Optimizing Use of AI Technologies (SPROUT) 50,292 22.4 168

† Present only in the evaluation set and not in the training set.

Wherever "Other" appears as a corpus, it indicates an unnamed corpus of in-classroom data.

Bias and Limitations

The model performs best on audio data that is:

  • Already diarized and clipped to a single utterance. The model was trained on already-diarized audio clips of individual utterances with a single speaker. It was not trained to perform forced alignment or speaker identification.
  • Under 60 seconds in length. The model was trained on utterances that were 20 seconds in length or shorter, and evaluated on audio clips that were less than 60 seconds long. To perform inference on longer audio, we recommend either chunking the audio or increasing max_new_tokens.

The model performs better in some circumstances, and on some individuals, than others. As reflected in the evaluation graphs above, the model struggles most with:

  • Very young children (under 5). As expected, the model performs better on older children. By age 5, the model performs comparably to state-of-the-art performance on adult speech.
  • Noisy classroom settings. 43 hours of in-classroom audio are represented in the training data, compared to 330 hours of non-classroom audio. Error for noisy in-classroom audio is roughly three times that for other settings.

Performance could not be evaluated for all racial groups. For example, there was insufficient data for Asian children to determine model performance. Within the model's evaluation dataset, model performance is similar among children who identify as Caucasian, African-American, or Hispanic.


Training Data

The model was trained on 372 hours of read, prompted, and spontaneous child speech spanning >3,600 different children as young as 3.

Training data was intentionally gathered to target specific populations where performance of currently available models is poor:

  • Children with atypical speech development
  • Non-white children
  • Children from low-income backgrounds

Training Data Collection

Training data was compiled as part of the On Top of Pasketti Challenge. Training data comes from 12 different source corpora (an additional two corpora are used as evaluation data only).

Only a few corpora have pre-existing transcriptions. To supplement existing transcriptions, DrivenData managed a team of linguists to transcribe additional audio data. 10% of audio files were retranscribed for quality assurance.

Child speech data is sensitive and difficult to share. We are grateful for the work of the data providers who thoughtfully collected and provided the data used in this project under appropriate consents, and to the speakers represented whose voices have helped to advance this work.

Representation in the Training Data

Breakdown by age:

Age Utterances Hours Children
3-4 63,573 41.6 272
5-7 107,134 84.9 1,888
8-11 192,105 191.7 1,293
12+ 29,793 8 197
Unknown 84,669 46.3 82

Breakdown by pathology:

Pathology Utterances Hours Children
Atypical development 140,833 58.9 599
Typical 141,380 101.9 635
Unknown 195,061 211.6 2,386

Breakdown by race:

Race Utterances Hours Children
Asian 3,609 2.5 37
Black 12,668 7.5 192
Hispanic 3,261 2.2 9
Multi-racial 13,126 7.7 107
Native american 20 <0.1 3
Other 52 <0.1 9
Pacific islander 16 <0.1 1
Unknown 342,611 281.9 2,644
White 101,911 70.6 621

Breakdown by setting:

Setting Utterances Hours
In-classroom 78,979 42.8
Not classroom 398,295 329.6

Breakdown by corpus:

Corpus Utterances Hours Children
Arizona Child Acoustic Database Repository 8,122 13 46
Cameron 9,576 6.5 35
Edmonton Narrative Norms Instrument 21,907 25.5 311
Ellis Weismer Corpus 32,254 22 93
JIBO Kids 5,690 3.5 81
My Science Tutor 77,502 133 645
Ohio Child Speech Corpus 112,072 77.8 273
Other 78,979 42.8 N/A
PERCEPT-GFTA 15,098 3.2 265
PERCEPT-R 87,211 26.2 257
ReadNet 11,885 11.4 1,562
Speech Production Repository for Optimizing Use of AI Technologies (SPROUT) 16,978 7.6 50

Breakdown by utterance duration:

Utterance duration (seconds) Utterances Hours Children
<1 133,666 26.9 2,085
[1-3) 211,375 100.1 2,988
[3-5) 60,541 65 2,598
[5-10) 49,243 94.4 2,470
[10-20) 22,449 85.9 1,302

Data Processing

Data was prepared for retraining by:

  • Separating into utterances. Audio was segmented to create a single audio file per utterance clipped to the utterance boundaries. Each sample was a single utterance audio with a corresponding transcript.
  • Normalizing transcriptions. Transcripts were normalized to standardize spelling, casing, punctuation, and representations of things like numbers.
  • Resampling. Audio data was resampled to 16 kHz, 1-channel, 16-bit signed and converted to FLAC format.
  • Dropping long utterances. Utterances longer than 20 seconds were dropped from the training dataset.

Have questions or want to know more? Get in touch!


Citation

@misc{drivendata2026pasketti,
  author       = {{DrivenData}},
  title        = {{pasketti-word}: A speech recognition model for children},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {Machine learning model},
  url          = {https://huggingface.co/drivendata/pasketti-word}
}

We would like to thank Tang Yongqwei, the On Top of Pasketti Word Track 3rd-place winner whose solution provided the foundation for pasketti-word. We are also grateful to the many individuals, organizations, and teams who contributed expertise, data, and guidance throughout the development of this project. For full project acknowledgements, see the On Top of Pasketti Challenge page.

Downloads last month
54
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for drivendata/pasketti-word

Finetuned
(117)
this model

Collection including drivendata/pasketti-word