Distil-Whisper: Optimized for Qualcomm Devices
Distil-Whisper Small English is a distilled version of Whisper Small, optimized for fast and efficient automatic speech recognition.
This is based on the implementation of Distil-Whisper found here. This repository contains pre-exported model files optimized for Qualcomm® devices. You can use the Qualcomm® AI Hub Models library to export with custom configurations. More details on model performance across various devices, can be found here.
Qualcomm AI Hub Models uses Qualcomm AI Hub Workbench to compile, profile, and evaluate this model. Sign up to run these models on a hosted Qualcomm® device.
Getting Started
There are two ways to deploy this model on your device:
Option 1: Download Pre-Exported Models
Below are pre-exported model assets ready for deployment.
| Runtime | Precision | Chipset | SDK Versions | Download |
|---|---|---|---|---|
| ONNX | float | Universal | QAIRT 2.45, ONNX Runtime 1.27.1 | Download |
| QNN_DLC | float | Universal | QAIRT 2.45 | Download |
| TFLITE | float | Universal | Download |
For more device-specific assets and performance metrics, visit Distil-Whisper on Qualcomm® AI Hub.
Option 2: Export with Custom Configurations
Use the Qualcomm® AI Hub Models Python library to compile and export the model with your own:
- Custom weights (e.g., fine-tuned checkpoints)
- Custom input shapes
- Target device and runtime configurations
This option is ideal if you need to customize the model beyond the default configuration provided here.
See our repository for Distil-Whisper on GitHub for usage instructions.
Model Details
Model Type: Model_use_case.speech_recognition
Model Stats:
- Input resolution: 80x3000 (30 seconds audio)
- Max decoded sequence length: 200 tokens
- Model checkpoint: distil-whisper/distil-small.en
- Model size (decoder) (float): 450MB
- Model size (encoder) (float): 332 MB
- Number of parameters (decoder): 211M
- Number of parameters (encoder): 166M
Performance Summary
| Model | Runtime | Precision | Chipset | Inference Time (ms) | Peak Memory Range (MB) | Primary Compute Unit |
|---|---|---|---|---|---|---|
| decoder | ONNX | float | Snapdragon® X2 Elite | 5.245 ms | 40 - 40 MB | NPU |
| decoder | ONNX | float | Snapdragon® X Elite | 11.182 ms | 179 - 179 MB | NPU |
| decoder | ONNX | float | Snapdragon® 8 Gen 3 Mobile | 8.846 ms | 0 - 489 MB | NPU |
| decoder | ONNX | float | Snapdragon® 8 Gen 1 Mobile | 18.335 ms | 50 - 398 MB | NPU |
| decoder | ONNX | float | Qualcomm® Dragonwing™ IQ-8275 | 13.271 ms | 40 - 83 MB | NPU |
| decoder | ONNX | float | Qualcomm® Dragonwing™ QCS8550 (Proxy) | 11.87 ms | 40 - 43 MB | NPU |
| decoder | ONNX | float | Qualcomm® QCS8450 | 18.335 ms | 50 - 398 MB | NPU |
| decoder | ONNX | float | Qualcomm® Dragonwing™ IQ-9075 | 15.634 ms | 40 - 82 MB | NPU |
| decoder | ONNX | float | Qualcomm® Dragonwing™ IQ-X7181 | 11.182 ms | 179 - 179 MB | NPU |
| decoder | ONNX | float | Qualcomm® Dragonwing™ Q-8750 | 7.24 ms | 0 - 554 MB | NPU |
| decoder | ONNX | float | Snapdragon® 8 Elite Mobile | 7.24 ms | 0 - 554 MB | NPU |
| decoder | ONNX | float | Snapdragon® 8 Elite Gen 5 Mobile | 5.748 ms | 6 - 550 MB | NPU |
| decoder | QNN_DLC | float | Snapdragon® X2 Elite | 5.94 ms | 40 - 40 MB | NPU |
| decoder | QNN_DLC | float | Snapdragon® X Elite | 10.986 ms | 40 - 40 MB | NPU |
| decoder | QNN_DLC | float | Snapdragon® 8 Gen 3 Mobile | 8.61 ms | 0 - 602 MB | NPU |
| decoder | QNN_DLC | float | Snapdragon® 8 Gen 1 Mobile | 18.219 ms | 36 - 339 MB | NPU |
| decoder | QNN_DLC | float | Qualcomm® Dragonwing™ IQ-8275 | 12.82 ms | 40 - 87 MB | NPU |
| decoder | QNN_DLC | float | Qualcomm® Dragonwing™ IQ-8275 | 19.115 ms | 40 - 537 MB | NPU |
| decoder | QNN_DLC | float | Qualcomm® Dragonwing™ QCS8550 (Proxy) | 11.469 ms | 40 - 44 MB | NPU |
| decoder | QNN_DLC | float | Qualcomm® SA8775P | 12.96 ms | 30 - 522 MB | NPU |
| decoder | QNN_DLC | float | Qualcomm® SA8650P | 12.96 ms | 30 - 522 MB | NPU |
| decoder | QNN_DLC | float | Qualcomm® SA8255P | 12.96 ms | 30 - 522 MB | NPU |
| decoder | QNN_DLC | float | Qualcomm® QCS8450 | 18.219 ms | 36 - 339 MB | NPU |
| decoder | QNN_DLC | float | Qualcomm® Dragonwing™ IQ-9075 | 12.684 ms | 40 - 86 MB | NPU |
| decoder | QNN_DLC | float | Qualcomm® Dragonwing™ IQ-X7181 | 10.986 ms | 40 - 40 MB | NPU |
| decoder | QNN_DLC | float | Qualcomm® Dragonwing™ Q-8750 | 7.307 ms | 4 - 548 MB | NPU |
| decoder | QNN_DLC | float | Qualcomm® SA7255P | 19.115 ms | 40 - 537 MB | NPU |
| decoder | QNN_DLC | float | Qualcomm® SA8295P | 14.02 ms | 20 - 262 MB | NPU |
| decoder | QNN_DLC | float | Snapdragon® 8 Elite Mobile | 7.307 ms | 4 - 548 MB | NPU |
| decoder | QNN_DLC | float | Snapdragon® 8 Elite Gen 5 Mobile | 5.925 ms | 1 - 503 MB | NPU |
| decoder | TFLITE | float | Snapdragon® 8 Gen 3 Mobile | 8.591 ms | 3 - 746 MB | NPU |
| decoder | TFLITE | float | Snapdragon® 8 Gen 1 Mobile | 18.323 ms | 5 - 467 MB | NPU |
| decoder | TFLITE | float | Qualcomm® Dragonwing™ IQ-8275 | 12.899 ms | 0 - 266 MB | NPU |
| decoder | TFLITE | float | Qualcomm® Dragonwing™ IQ-8275 | 19.236 ms | 4 - 537 MB | NPU |
| decoder | TFLITE | float | Qualcomm® Dragonwing™ QCS8550 (Proxy) | 11.605 ms | 5 - 8 MB | NPU |
| decoder | TFLITE | float | Qualcomm® SA8775P | 12.992 ms | 5 - 538 MB | NPU |
| decoder | TFLITE | float | Qualcomm® SA8650P | 12.992 ms | 5 - 538 MB | NPU |
| decoder | TFLITE | float | Qualcomm® SA8255P | 12.992 ms | 5 - 538 MB | NPU |
| decoder | TFLITE | float | Qualcomm® QCS8450 | 18.323 ms | 5 - 467 MB | NPU |
| decoder | TFLITE | float | Qualcomm® Dragonwing™ IQ-9075 | 12.811 ms | 0 - 265 MB | NPU |
| decoder | TFLITE | float | Qualcomm® Dragonwing™ Q-8750 | 7.24 ms | 5 - 575 MB | NPU |
| decoder | TFLITE | float | Qualcomm® SA7255P | 19.236 ms | 4 - 537 MB | NPU |
| decoder | TFLITE | float | Qualcomm® SA8295P | 13.902 ms | 5 - 296 MB | NPU |
| decoder | TFLITE | float | Snapdragon® 8 Elite Mobile | 7.24 ms | 5 - 575 MB | NPU |
| decoder | TFLITE | float | Snapdragon® 8 Elite Gen 5 Mobile | 5.842 ms | 4 - 574 MB | NPU |
| encoder | ONNX | float | Snapdragon® X2 Elite | 60.4 ms | 133 - 133 MB | NPU |
| encoder | ONNX | float | Snapdragon® X Elite | 133.596 ms | 237 - 237 MB | NPU |
| encoder | ONNX | float | Snapdragon® 8 Gen 3 Mobile | 97.178 ms | 0 - 1079 MB | NPU |
| encoder | ONNX | float | Snapdragon® 8 Gen 1 Mobile | 270.258 ms | 76 - 1019 MB | NPU |
| encoder | ONNX | float | Qualcomm® Dragonwing™ IQ-8275 | 162.463 ms | 74 - 79 MB | NPU |
| encoder | ONNX | float | Qualcomm® Dragonwing™ QCS8550 (Proxy) | 129.246 ms | 0 - 258 MB | NPU |
| encoder | ONNX | float | Qualcomm® QCS8450 | 270.258 ms | 76 - 1019 MB | NPU |
| encoder | ONNX | float | Qualcomm® Dragonwing™ IQ-9075 | 152.735 ms | 76 - 79 MB | NPU |
| encoder | ONNX | float | Qualcomm® Dragonwing™ IQ-X7181 | 133.596 ms | 237 - 237 MB | NPU |
| encoder | ONNX | float | Qualcomm® Dragonwing™ Q-8750 | 71.573 ms | 79 - 789 MB | NPU |
| encoder | ONNX | float | Snapdragon® 8 Elite Mobile | 71.573 ms | 79 - 789 MB | NPU |
| encoder | ONNX | float | Snapdragon® 8 Elite Gen 5 Mobile | 58.128 ms | 68 - 803 MB | NPU |
| encoder | QNN_DLC | float | Snapdragon® X2 Elite | 60.245 ms | 1 - 1 MB | NPU |
| encoder | QNN_DLC | float | Snapdragon® X Elite | 139.945 ms | 1 - 1 MB | NPU |
| encoder | QNN_DLC | float | Snapdragon® 8 Gen 3 Mobile | 97.319 ms | 1 - 966 MB | NPU |
| encoder | QNN_DLC | float | Snapdragon® 8 Gen 1 Mobile | 270.339 ms | 0 - 822 MB | NPU |
| encoder | QNN_DLC | float | Qualcomm® Dragonwing™ IQ-8275 | 161.731 ms | 1 - 39 MB | NPU |
| encoder | QNN_DLC | float | Qualcomm® Dragonwing™ IQ-8275 | 438.051 ms | 1 - 697 MB | NPU |
| encoder | QNN_DLC | float | Qualcomm® Dragonwing™ QCS8550 (Proxy) | 135.911 ms | 0 - 6 MB | NPU |
| encoder | QNN_DLC | float | Qualcomm® SA8775P | 153.263 ms | 1 - 686 MB | NPU |
| encoder | QNN_DLC | float | Qualcomm® SA8650P | 153.263 ms | 1 - 686 MB | NPU |
| encoder | QNN_DLC | float | Qualcomm® SA8255P | 153.263 ms | 1 - 686 MB | NPU |
| encoder | QNN_DLC | float | Qualcomm® QCS8450 | 270.339 ms | 0 - 822 MB | NPU |
| encoder | QNN_DLC | float | Qualcomm® Dragonwing™ IQ-9075 | 153.492 ms | 3 - 41 MB | NPU |
| encoder | QNN_DLC | float | Qualcomm® Dragonwing™ IQ-X7181 | 139.945 ms | 1 - 1 MB | NPU |
| encoder | QNN_DLC | float | Qualcomm® Dragonwing™ Q-8750 | 71.457 ms | 1 - 691 MB | NPU |
| encoder | QNN_DLC | float | Qualcomm® SA7255P | 438.051 ms | 1 - 697 MB | NPU |
| encoder | QNN_DLC | float | Qualcomm® SA8295P | 192.968 ms | 1 - 610 MB | NPU |
| encoder | QNN_DLC | float | Snapdragon® 8 Elite Mobile | 71.457 ms | 1 - 691 MB | NPU |
| encoder | QNN_DLC | float | Snapdragon® 8 Elite Gen 5 Mobile | 57.819 ms | 0 - 718 MB | NPU |
| encoder | TFLITE | float | Snapdragon® 8 Gen 3 Mobile | 481.599 ms | 42 - 187 MB | GPU |
| encoder | TFLITE | float | Snapdragon® 8 Gen 1 Mobile | 839.03 ms | 41 - 194 MB | GPU |
| encoder | TFLITE | float | Qualcomm® Dragonwing™ IQ-8275 | 2998.255 ms | 0 - 39 MB | GPU |
| encoder | TFLITE | float | Qualcomm® Dragonwing™ IQ-8275 | 3123.092 ms | 18 - 62 MB | GPU |
| encoder | TFLITE | float | Qualcomm® Dragonwing™ QCS8550 (Proxy) | 649.127 ms | 0 - 311 MB | GPU |
| encoder | TFLITE | float | Qualcomm® SA8775P | 1324.913 ms | 17 - 61 MB | GPU |
| encoder | TFLITE | float | Qualcomm® SA8650P | 1324.913 ms | 17 - 61 MB | GPU |
| encoder | TFLITE | float | Qualcomm® SA8255P | 1324.913 ms | 17 - 61 MB | GPU |
| encoder | TFLITE | float | Qualcomm® QCS8450 | 839.03 ms | 41 - 194 MB | GPU |
| encoder | TFLITE | float | Qualcomm® Dragonwing™ IQ-9075 | 1281.735 ms | 0 - 39 MB | GPU |
| encoder | TFLITE | float | Qualcomm® Dragonwing™ Q-8750 | 405.666 ms | 42 - 81 MB | GPU |
| encoder | TFLITE | float | Qualcomm® SA7255P | 3123.092 ms | 18 - 62 MB | GPU |
| encoder | TFLITE | float | Qualcomm® SA8295P | 667.287 ms | 38 - 83 MB | GPU |
| encoder | TFLITE | float | Snapdragon® 8 Elite Mobile | 405.666 ms | 42 - 81 MB | GPU |
| encoder | TFLITE | float | Snapdragon® 8 Elite Gen 5 Mobile | 367.921 ms | 40 - 79 MB | GPU |
License
- The license for the original implementation of Distil-Whisper can be found here.
References
- Distil-Whisper - Robust Knowledge Distillation via Large-Scale Pseudo Labelling
- Source Model Implementation
Community
- Join our AI Hub Slack community to collaborate, post questions and learn more about on-device AI.
- For questions or feedback please reach out to us.
