pasketti-word 🍝
pasketti-word is an automatic speech recognition (ASR) model designed specifically for child speech. pasketti-word generates orthographic (per word) transcriptions from audio files.
🐣 Existing open-source ASR models trained on adult speech perform poorly on children because child speech is fundamentally different from adult speech. Children are still learning to produce speech sounds and developing fine motor skills. Common child speech errors include:
- Metathesis: "elephant" → "ephelant"
- Velar fronting: "cup" → "tup"
- Syllable deletion: "banana" → "nana"
- Combinations of errors: "spaghetti" → "pasketti" 🍝
🔉 pasketti-word is fine-tuned on a large corpus of child speech data including 372 hours of child speech spanning >3,600 different children. This is the largest dataset of its kind for use in developing and testing open ASR models for kids. The model is designed to advance early education assessments, teaching tools, and needs screening.
📈 pasketti-word achieves an overall word error rate (WER) of 17% on held-out evaluation data. When compared to leading open ASR models (Whisper and KidWhisper) on the same evaluation set, pasketti-word demonstrates improved performance across all demographic groups, cutting the error rate by more than half for children ages 3-7 and achieving close to adult-level performance for children ages 5 and above.
In this model card:
Quickstart
Prerequisites
Hardware. Inference runs through vLLM. See the vLLM installation docs for supported hardware and operating systems. This model has been tested only on Linux with NVIDIA GPUs.
Environment. Install
qwen_asrwith its vLLM extra:pip install "qwen_asr[vllm]==0.0.6"pip install transformersalone cannot load this model. Theqwen3_asrarchitecture is not built intotransformers; it is registered whenqwen_asris imported.This release was verified against
qwen_asr==0.0.6,vllm==0.14.0,transformers==4.57.6andtorch==2.9.1, on Python 3.13.qwen_asrpins an exacttransformersversion, so if newer releases break the install, pin these versions to reconstruct a working environment.Data. Pass each clip to
asr.transcribe(audio=...)in one of these forms:- a local file path, an
httpsURL, or a base64-encoded audio string, whichqwen_asrreads directly, or - a
(waveform, sample_rate)tuple, wherewaveformis a 1-D (mono) or 2-D (multi-channel) NumPy array andsample_rateis its true sample rate in Hz. To match how the model was trained, load the audio withlibrosa, as in the example below.
- a local file path, an
qwen_asr converts every input to 16 kHz mono before transcription.
Inference
Pass the Hub repo id, drivendata/pasketti-word, to Qwen3ASRModel.LLM. The weights download automatically the first time the model loads.
Transcribe a single clip:
from qwen_asr import Qwen3ASRModel
import librosa
asr = Qwen3ASRModel.LLM(
"drivendata/pasketti-word",
gpu_memory_utilization=0.9,
max_inference_batch_size=32,
max_new_tokens=256,
)
wav, _ = librosa.load("clip.flac", sr=16000, mono=True)
result = asr.transcribe(audio=[(wav, 16000)], language="English")
print(result[0].text)
>> "the cat sat down"
For large datasets, load and transcribe the audio in chunks (for example, a few thousand clips at a time) rather than holding every clip in memory at once.
qwen_asr also offers a transformers backend through Qwen3ASRModel.from_pretrained(). That backend has not been evaluated with this model and is not supported.
Model Details
This model is a supervised fine-tune of Qwen3-ASR-1.7B. Qwen3-ASR pairs an audio encoder with a Qwen3 language model, which writes out the transcript one token at a time. All of the model's weights were updated during fine-tuning, using clips of up to 20 seconds of children's speech with orthographic (word-level) transcripts.
The model was trained for two epochs with a global batch size of 32 (4 per GPU across 8 A100s), using AdamW with a peak learning rate of 3e-5. The schedule used linear warmup over the first 2% of steps, followed by linear decay sized for five epochs.
The model design is based on the results of the On Top of Pasketti Challenge, in which participants competed to build child ASR models. All five prize-winning Word Track solutions fine-tuned Qwen3-ASR-1.7B, and pasketti-word adapts the 3rd-place solution, a full-parameter fine-tune.
Developed by: DrivenData, based on the 3rd-place Word Track solution by Tang Yongqwei (chuxiliyixiaosa)
References
Uses
This model is intended to be used for automatic speech recognition in service of improving educational and developmental outcomes for children.
Based on the model's performance, it is most suitable for downstream use cases that depend on child transcripts as an input. These downstream use cases often do not require exact transcripts to effectively inform decision-making.
For use cases that require exact transcription from child speech, model outputs should be human reviewed.
Downstream Use Cases
Example downstream use cases include:
- Needs screening: Identifying children who may be struggling and better targeting additional support. For example, identifying children behind in literacy or who may have a speech pathology.
- Formative assessment: Monitoring student learning based on naturally collected classroom speech data.
- Educator feedback: Enabling better teacher or tutor support and growth by incorporating spoken interactions.
The model could be plugged into decision-making processes as is to generate transcripts for human review, or fine-tuned to predict in-scope outcomes directly. Intended users include education researchers, education technology developers, teachers, and school administrators.
Out-of-Scope Use
The model should not be used in ways that do not serve the interests of children whose speech it processes. Out-of-scope uses include:
- Surveillance or profiling of children, such as building voice or behavioral profiles for advertising, identifying individual speakers, or monitoring children without informed consent from them and their guardians.
- Engagement maximization, such as designing products that use speech signals to drive compulsive use or unhealthy dependence on technology.
- Fully automated high-stakes decisions, such as diagnosis or educational placement, without review by a qualified educator or clinician.
Evaluation
pasketti-word performance: 17% word error rate (WER)
The evaluation set includes 148 hours of child speech from >2,300 different children as young as 3. All models are scored on every utterance with Whisper's English normalizer.
The model is compared to existing state-of-the-art open ASR models openai/whisper-medium and aadel4/kid-whisper-medium-en-myst (Whisper fine-tuned on a small amount of child speech data).
All models were run with loop protection:
- Whisper and KidWhisper are decoded greedily with the compression-ratio fallback from OpenAI's
transcribe(). Retrying low-confidence predictions is not implemented. pasketti-wordimplements loop protection by default. It usesqwen_asr's default decoding, which collapses repeats on every output, with a 256-token limit.
About Word Error Rate
Word Error Rate (WER) measures the minimum number of word-level substitutions (𝑆), deletions (𝐷), and insertions (𝐼) required to transform the predicted transcript into the reference transcript, divided by the total number of words in the reference (𝑁). WER indicates roughly what percent of words predicted are wrong.
A limitation of WER is that it only measures whether the predicted word is exactly correct — it does not reward similar predictions that are close to correct. As a result, WER often slightly underestimates how well a model can perform in a real-world setting.
Performance by Subset
The model was evaluated for bias in performance based on age, race, speech development pathology, sex, and source corpus. As expected, performance tends to improve for older children compared to younger children. However, performance gains are pronounced for younger children, with WER for children ages 3-4 improving by 30 percentage points compared to Whisper.
Compared to Whisper, pasketti-word decreases key disparities between white vs. non-white children as well as between children with speech pathologies vs. children without.
Performance by Age
Whisper performance is likely poor on ages 12+ because those samples are derived from particularly challenging corpuses.
Representation in the evaluation set:
| Age | Utterances | Hours | Children |
|---|---|---|---|
| 3-4 | 54,341 | 25.2 | 190 |
| 5-7 | 15,677 | 13 | 681 |
| 8-11 | 29,799 | 37.8 | 278 |
| 12+ | 5,320 | 1.4 | 44 |
| Unknown | 78,735 | 70.5 | 1,148 |
Performance by Pathology
Representation in the evaluation set:
| Pathology | Utterances | Hours | Children |
|---|---|---|---|
| Atypical development | 38,005 | 14.5 | 175 |
| Typical | 41,891 | 23.8 | 154 |
| Unknown | 103,976 | 109.6 | 1,996 |
Performance by Race
Racial groups are compared within each individual corpus to remove any correlation between corpus difficulty and race that influences overall race-based performance. This is necessary because very few corpora have race populated.
Representation in the evaluation set:
| Corpus | Race | Utterances | Hours | Children |
|---|---|---|---|---|
| SPROUT | Black | 19,360 | 8.8 | 67 |
| Hispanic | 1,637 | 0.7 | 6 | |
| Multi-racial | 13,285 | 5.3 | 47 | |
| White | 12,457 | 5.7 | 45 | |
| ReadNet | Black | 348 | 0.3 | 47 |
| Multi-racial | 234 | 0.2 | 30 | |
| White | 1,889 | 1.8 | 176 |
Performance by Setting
Representation in the evaluation set:
| Setting | Utterances | Hours | Children |
|---|---|---|---|
| In-classroom | 12,261 | 6.3 | NA |
| Not classroom | 171,611 | 141.7 | 2,324 |
Performance by Sex
Representation in the evaluation set:
| Sex | Utterances | Hours | Children |
|---|---|---|---|
| Female | 47,714 | 26 | 379 |
| Male | 43,731 | 22.3 | 336 |
| Unknown | 92,427 | 99.6 | 1,610 |
Performance by Duration
Representation in the evaluation set:
| Utterance duration (seconds) | Utterances | Hours | Children |
|---|---|---|---|
| <1 | 34,831 | 7.3 | 553 |
| [1-3) | 88,682 | 44.7 | 2,052 |
| [3-5) | 36,962 | 40.2 | 1,991 |
| [5-10) | 18,942 | 34 | 1,941 |
| [10-20) | 3,330 | 12.6 | 278 |
| [20-60) | 1,125 | 9.1 | 193 |
Performance by Corpus
Performance is heavily influenced by corpus due to corpus-specific age ranges, data collection methods, and audio quality. Two corpora were put entirely in the test set: CSLU: Kids' Speech Version 1.1 and the CMU Kids Corpus.
Representation in the evaluation set:
| Corpus | Utterances | Hours | Children |
|---|---|---|---|
| Arizona Child Acoustic Database Repository | 497 | 0.8 | 5 |
| CMU Kids Corpus† | 2,477 | 3.7 | 74 |
| CSLU: Kids' Speech Version 1.1† | 64,430 | 62.8 | 1,117 |
| Cameron | 1,326 | 0.7 | 5 |
| Edmonton Narrative Norms Instrument | 2,415 | 2.6 | 34 |
| Ellis Weismer Corpus | 1,775 | 1.2 | 5 |
| JIBO Kids | 1,926 | 1.4 | 28 |
| My Science Tutor | 12,607 | 27.6 | 91 |
| Ohio Child Speech Corpus | 12,271 | 8.8 | 30 |
| Other | 12,261 | 6.3 | NA |
| PERCEPT-GFTA | 4,430 | 1 | 77 |
| PERCEPT-R | 11,836 | 3.5 | 29 |
| ReadNet | 5,329 | 5.1 | 661 |
| Speech Production Repository for Optimizing Use of AI Technologies (SPROUT) | 50,292 | 22.4 | 168 |
† Present only in the evaluation set and not in the training set.
Wherever "Other" appears as a corpus, it indicates an unnamed corpus of in-classroom data.
Bias and Limitations
The model performs best on audio data that is:
- Already diarized and clipped to a single utterance. The model was trained on already-diarized audio clips of individual utterances with a single speaker. It was not trained to perform forced alignment or speaker identification.
- Under 60 seconds in length. The model was trained on utterances that were 20 seconds in length or shorter, and evaluated on audio clips that were less than 60 seconds long. To perform inference on longer audio, we recommend either chunking the audio or increasing
max_new_tokens.
The model performs better in some circumstances, and on some individuals, than others. As reflected in the evaluation graphs above, the model struggles most with:
- Very young children (under 5). As expected, the model performs better on older children. By age 5, the model performs comparably to state-of-the-art performance on adult speech.
- Noisy classroom settings. 43 hours of in-classroom audio are represented in the training data, compared to 330 hours of non-classroom audio. Error for noisy in-classroom audio is roughly three times that for other settings.
Performance could not be evaluated for all racial groups. For example, there was insufficient data for Asian children to determine model performance. Within the model's evaluation dataset, model performance is similar among children who identify as Caucasian, African-American, or Hispanic.
Training Data
The model was trained on 372 hours of read, prompted, and spontaneous child speech spanning >3,600 different children as young as 3.
Training data was intentionally gathered to target specific populations where performance of currently available models is poor:
- Children with atypical speech development
- Non-white children
- Children from low-income backgrounds
Training Data Collection
Training data was compiled as part of the On Top of Pasketti Challenge. Training data comes from 12 different source corpora (an additional two corpora are used as evaluation data only).
Only a few corpora have pre-existing transcriptions. To supplement existing transcriptions, DrivenData managed a team of linguists to transcribe additional audio data. 10% of audio files were retranscribed for quality assurance.
Child speech data is sensitive and difficult to share. We are grateful for the work of the data providers who thoughtfully collected and provided the data used in this project under appropriate consents, and to the speakers represented whose voices have helped to advance this work.
Representation in the Training Data
Breakdown by age:
| Age | Utterances | Hours | Children |
|---|---|---|---|
| 3-4 | 63,573 | 41.6 | 272 |
| 5-7 | 107,134 | 84.9 | 1,888 |
| 8-11 | 192,105 | 191.7 | 1,293 |
| 12+ | 29,793 | 8 | 197 |
| Unknown | 84,669 | 46.3 | 82 |
Breakdown by pathology:
| Pathology | Utterances | Hours | Children |
|---|---|---|---|
| Atypical development | 140,833 | 58.9 | 599 |
| Typical | 141,380 | 101.9 | 635 |
| Unknown | 195,061 | 211.6 | 2,386 |
Breakdown by race:
| Race | Utterances | Hours | Children |
|---|---|---|---|
| Asian | 3,609 | 2.5 | 37 |
| Black | 12,668 | 7.5 | 192 |
| Hispanic | 3,261 | 2.2 | 9 |
| Multi-racial | 13,126 | 7.7 | 107 |
| Native american | 20 | <0.1 | 3 |
| Other | 52 | <0.1 | 9 |
| Pacific islander | 16 | <0.1 | 1 |
| Unknown | 342,611 | 281.9 | 2,644 |
| White | 101,911 | 70.6 | 621 |
Breakdown by setting:
| Setting | Utterances | Hours |
|---|---|---|
| In-classroom | 78,979 | 42.8 |
| Not classroom | 398,295 | 329.6 |
Breakdown by corpus:
| Corpus | Utterances | Hours | Children |
|---|---|---|---|
| Arizona Child Acoustic Database Repository | 8,122 | 13 | 46 |
| Cameron | 9,576 | 6.5 | 35 |
| Edmonton Narrative Norms Instrument | 21,907 | 25.5 | 311 |
| Ellis Weismer Corpus | 32,254 | 22 | 93 |
| JIBO Kids | 5,690 | 3.5 | 81 |
| My Science Tutor | 77,502 | 133 | 645 |
| Ohio Child Speech Corpus | 112,072 | 77.8 | 273 |
| Other | 78,979 | 42.8 | N/A |
| PERCEPT-GFTA | 15,098 | 3.2 | 265 |
| PERCEPT-R | 87,211 | 26.2 | 257 |
| ReadNet | 11,885 | 11.4 | 1,562 |
| Speech Production Repository for Optimizing Use of AI Technologies (SPROUT) | 16,978 | 7.6 | 50 |
Breakdown by utterance duration:
| Utterance duration (seconds) | Utterances | Hours | Children |
|---|---|---|---|
| <1 | 133,666 | 26.9 | 2,085 |
| [1-3) | 211,375 | 100.1 | 2,988 |
| [3-5) | 60,541 | 65 | 2,598 |
| [5-10) | 49,243 | 94.4 | 2,470 |
| [10-20) | 22,449 | 85.9 | 1,302 |
Data Processing
Data was prepared for retraining by:
- Separating into utterances. Audio was segmented to create a single audio file per utterance clipped to the utterance boundaries. Each sample was a single utterance audio with a corresponding transcript.
- Normalizing transcriptions. Transcripts were normalized to standardize spelling, casing, punctuation, and representations of things like numbers.
- Resampling. Audio data was resampled to 16 kHz, 1-channel, 16-bit signed and converted to FLAC format.
- Dropping long utterances. Utterances longer than 20 seconds were dropped from the training dataset.
Have questions or want to know more? Get in touch!
Citation
@misc{drivendata2026pasketti,
author = {{DrivenData}},
title = {{pasketti-word}: A speech recognition model for children},
year = {2026},
publisher = {Hugging Face},
howpublished = {Machine learning model},
url = {https://huggingface.co/drivendata/pasketti-word}
}
We would like to thank Tang Yongqwei, the On Top of Pasketti Word Track 3rd-place winner whose solution provided the foundation for pasketti-word. We are also grateful to the many individuals, organizations, and teams who contributed expertise, data, and guidance throughout the development of this project. For full project acknowledgements, see the On Top of Pasketti Challenge page.
- Downloads last month
- 54
Model tree for drivendata/pasketti-word
Base model
Qwen/Qwen3-ASR-1.7B





