Whisper-Small-Quantized: Optimized for Qualcomm Devices

We have applied w8a16 quantization to significantly enhance performance and efficiency. HuggingFace Whisper-Small ASR (Automatic Speech Recognition) model is a state-of-the-art system designed for transcribing spoken language into written text. This model is based on the transformer architecture and has been optimized for edge inference by replacing Multi-Head Attention (MHA) with Single-Head Attention (SHA) and linear layers with convolutional (conv) layers. It exhibits robust performance in realistic, noisy environments, making it highly reliable for real-world applications. Specifically, it excels in long-form transcription, capable of accurately transcribing audio clips up to 30 seconds long. Time to the first token is the encoder's latency, while time to each additional token is decoder's latency, where we assume a max decoded length specified below.

This is based on the implementation of Whisper-Small-Quantized found here. This repository contains pre-exported model files optimized for Qualcomm® devices. You can use the Qualcomm® AI Hub Models library to export with custom configurations. More details on model performance across various devices, can be found here.

Qualcomm AI Hub Models uses Qualcomm AI Hub Workbench to compile, profile, and evaluate this model. Sign up to run these models on a hosted Qualcomm® device.

Deploying Whisper-Small-Quantized on-device

This model is compatible with the Qualcomm Voice AI SDK. Download the SDK from the Qualcomm Package Manager to deploy this model on-device.

Getting Started

There are two ways to deploy this model on your device:

Option 1: Download Pre-Exported Models

Below are pre-exported model assets ready for deployment.

Runtime Precision Chipset SDK Versions Download
PRECOMPILED_QNN_ONNX w8a16 Snapdragon® 8 Elite Gen 5 For Galaxy Mobile QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX w8a16 Snapdragon® 8 Elite For Galaxy Mobile QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX w8a16 Snapdragon® X2 Elite QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX w8a16 Snapdragon® X Elite QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX w8a16 Snapdragon® 8 Gen 3 Mobile QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX w8a16 Snapdragon® 8 Gen 1 Mobile QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ QCS6490 QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ IQ-8275 QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ QCS8550 (Proxy) QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ Q-6690 QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ IQ-9075 QAIRT 2.50, ONNX Runtime 1.27.1 Download
PRECOMPILED_QNN_ONNX w8a16 Snapdragon® 7 Gen 4 Mobile QAIRT 2.50, ONNX Runtime 1.27.1 Download
QNN_CONTEXT_BINARY w8a16 Snapdragon® 8 Elite Gen 5 For Galaxy Mobile QAIRT 2.50 Download
QNN_CONTEXT_BINARY w8a16 Snapdragon® 8 Elite For Galaxy Mobile QAIRT 2.50 Download
QNN_CONTEXT_BINARY w8a16 Snapdragon® X2 Elite QAIRT 2.50 Download
QNN_CONTEXT_BINARY w8a16 Snapdragon® X Elite QAIRT 2.50 Download
QNN_CONTEXT_BINARY w8a16 Snapdragon® 8 Gen 3 Mobile QAIRT 2.50 Download
QNN_CONTEXT_BINARY w8a16 Snapdragon® 8 Gen 1 Mobile QAIRT 2.50 Download
QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ QCS6490 QAIRT 2.50 Download
QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ IQ-8275 QAIRT 2.50 Download
QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ QCS8550 (Proxy) QAIRT 2.50 Download
QNN_CONTEXT_BINARY w8a16 Qualcomm® SA8775P QAIRT 2.50 Download
QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ Q-6690 QAIRT 2.50 Download
QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ IQ-9075 QAIRT 2.50 Download
QNN_CONTEXT_BINARY w8a16 Qualcomm® SA7255P QAIRT 2.50 Download
QNN_CONTEXT_BINARY w8a16 Qualcomm® SA8295P QAIRT 2.50 Download
QNN_CONTEXT_BINARY w8a16 Snapdragon® 7 Gen 4 Mobile QAIRT 2.50 Download
VOICE_AI w8a16 Snapdragon® 8 Elite Gen 5 For Galaxy Mobile QAIRT 2.50 Download
VOICE_AI w8a16 Snapdragon® 8 Elite For Galaxy Mobile QAIRT 2.50 Download
VOICE_AI w8a16 Snapdragon® X2 Elite QAIRT 2.50 Download
VOICE_AI w8a16 Snapdragon® X Elite QAIRT 2.50 Download
VOICE_AI w8a16 Snapdragon® 8 Gen 3 Mobile QAIRT 2.50 Download
VOICE_AI w8a16 Snapdragon® 8 Gen 1 Mobile QAIRT 2.50 Download
VOICE_AI w8a16 Qualcomm® Dragonwing™ QCS6490 QAIRT 2.50 Download
VOICE_AI w8a16 Qualcomm® Dragonwing™ IQ-8275 QAIRT 2.50 Download
VOICE_AI w8a16 Qualcomm® Dragonwing™ QCS8550 (Proxy) QAIRT 2.50 Download
VOICE_AI w8a16 Qualcomm® SA8775P QAIRT 2.50 Download
VOICE_AI w8a16 Qualcomm® Dragonwing™ Q-6690 QAIRT 2.50 Download
VOICE_AI w8a16 Qualcomm® Dragonwing™ IQ-9075 QAIRT 2.50 Download
VOICE_AI w8a16 Qualcomm® SA7255P QAIRT 2.50 Download
VOICE_AI w8a16 Qualcomm® SA8295P QAIRT 2.50 Download
VOICE_AI w8a16 Snapdragon® 7 Gen 4 Mobile QAIRT 2.50 Download

For more device-specific assets and performance metrics, visit Whisper-Small-Quantized on Qualcomm® AI Hub.

Option 2: Export with Custom Configurations

Use the Qualcomm® AI Hub Models Python library to compile and export the model with your own:

  • Custom weights (e.g., fine-tuned checkpoints)
  • Custom input shapes
  • Target device and runtime configurations

This option is ideal if you need to customize the model beyond the default configuration provided here.

See our repository for Whisper-Small-Quantized on GitHub for usage instructions.

Model Details

Model Type: Model_use_case.speech_recognition

Model Stats:

  • Input resolution: 80x3000 (30 seconds audio)
  • Max decoded sequence length: 200 tokens
  • Model checkpoint: openai/whisper-small

Performance Summary

Model Runtime Precision Chipset Inference Time (ms) Peak Memory Range (MB) Primary Compute Unit
decoder PRECOMPILED_QNN_ONNX w8a16 Snapdragon® 8 Elite Gen 5 For Galaxy Mobile 3.954 ms 23 - 37 MB NPU
decoder PRECOMPILED_QNN_ONNX w8a16 Snapdragon® 8 Elite For Galaxy Mobile 4.71 ms 26 - 37 MB NPU
decoder PRECOMPILED_QNN_ONNX w8a16 Snapdragon® X2 Elite 3.716 ms 33 - 33 MB NPU
decoder PRECOMPILED_QNN_ONNX w8a16 Snapdragon® X Elite 7.17 ms 186 - 186 MB NPU
decoder PRECOMPILED_QNN_ONNX w8a16 Snapdragon® 8 Gen 3 Mobile 6.19 ms 37 - 50 MB NPU
decoder PRECOMPILED_QNN_ONNX w8a16 Snapdragon® 8 Gen 1 Mobile 9.516 ms 37 - 52 MB NPU
decoder PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ QCS6490 24.425 ms 30 - 63 MB NPU
decoder PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ IQ-8275 9.157 ms 25 - 59 MB NPU
decoder PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ QCS8550 (Proxy) 8.017 ms 0 - 193 MB NPU
decoder PRECOMPILED_QNN_ONNX w8a16 Qualcomm® QCS8450 9.516 ms 37 - 52 MB NPU
decoder PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ IQ-9075 8.804 ms 24 - 58 MB NPU
decoder PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ IQ-X7181 7.17 ms 186 - 186 MB NPU
decoder PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ Q-6690 29.332 ms 38 - 51 MB NPU
decoder PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ Q-7790 10.444 ms 30 - 38 MB NPU
decoder PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ Q-8750 4.71 ms 26 - 37 MB NPU
decoder PRECOMPILED_QNN_ONNX w8a16 Snapdragon® 7 Gen 4 Mobile 10.444 ms 30 - 38 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Snapdragon® 8 Elite Gen 5 For Galaxy Mobile 3.912 ms 24 - 33 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Snapdragon® 8 Elite For Galaxy Mobile 4.632 ms 22 - 30 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Snapdragon® X2 Elite 4.272 ms 30 - 30 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Snapdragon® X Elite 7.823 ms 30 - 30 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Snapdragon® 8 Gen 3 Mobile 6.179 ms 30 - 38 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Snapdragon® 8 Gen 1 Mobile 9.548 ms 28 - 38 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ QCS6490 23.463 ms 30 - 66 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ IQ-8275 8.945 ms 25 - 62 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ QCS8550 (Proxy) 7.934 ms 31 - 33 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Qualcomm® SA8650P 8.923 ms 22 - 32 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Qualcomm® SA8255P 8.923 ms 22 - 32 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Qualcomm® QCS8450 9.548 ms 28 - 38 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ IQ-9075 8.642 ms 27 - 63 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ IQ-X7181 7.823 ms 30 - 30 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ Q-6690 29.369 ms 31 - 38 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ Q-7790 10.29 ms 29 - 36 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ Q-8750 4.632 ms 22 - 30 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Qualcomm® SA8295P 9.527 ms 29 - 35 MB NPU
decoder QNN_CONTEXT_BINARY w8a16 Snapdragon® 7 Gen 4 Mobile 10.29 ms 29 - 36 MB NPU
decoder VOICE_AI w8a16 Snapdragon® 8 Elite Gen 5 For Galaxy Mobile 3.922 ms 24 - 32 MB NPU
decoder VOICE_AI w8a16 Snapdragon® 8 Elite For Galaxy Mobile 4.637 ms 18 - 26 MB NPU
decoder VOICE_AI w8a16 Snapdragon® X2 Elite 4.238 ms 30 - 30 MB NPU
decoder VOICE_AI w8a16 Snapdragon® X Elite 7.744 ms 30 - 30 MB NPU
decoder VOICE_AI w8a16 Snapdragon® 8 Gen 3 Mobile 6.135 ms 30 - 38 MB NPU
decoder VOICE_AI w8a16 Snapdragon® 8 Gen 1 Mobile 9.438 ms 27 - 40 MB NPU
decoder VOICE_AI w8a16 Qualcomm® Dragonwing™ QCS6490 23.521 ms 30 - 66 MB NPU
decoder VOICE_AI w8a16 Qualcomm® Dragonwing™ IQ-8275 8.929 ms 25 - 62 MB NPU
decoder VOICE_AI w8a16 Qualcomm® Dragonwing™ QCS8550 (Proxy) 7.872 ms 30 - 34 MB NPU
decoder VOICE_AI w8a16 Qualcomm® SA8650P 8.941 ms 29 - 38 MB NPU
decoder VOICE_AI w8a16 Qualcomm® SA8255P 8.941 ms 29 - 38 MB NPU
decoder VOICE_AI w8a16 Qualcomm® QCS8450 9.438 ms 27 - 40 MB NPU
decoder VOICE_AI w8a16 Qualcomm® Dragonwing™ IQ-9075 8.596 ms 25 - 61 MB NPU
decoder VOICE_AI w8a16 Qualcomm® Dragonwing™ IQ-X7181 7.744 ms 30 - 30 MB NPU
decoder VOICE_AI w8a16 Qualcomm® Dragonwing™ Q-6690 31.089 ms 30 - 37 MB NPU
decoder VOICE_AI w8a16 Qualcomm® Dragonwing™ Q-7790 10.237 ms 30 - 37 MB NPU
decoder VOICE_AI w8a16 Qualcomm® Dragonwing™ Q-8750 4.637 ms 18 - 26 MB NPU
decoder VOICE_AI w8a16 Qualcomm® SA8295P 9.554 ms 29 - 35 MB NPU
decoder VOICE_AI w8a16 Snapdragon® 7 Gen 4 Mobile 10.237 ms 30 - 37 MB NPU
encoder PRECOMPILED_QNN_ONNX w8a16 Snapdragon® 8 Elite Gen 5 For Galaxy Mobile 27.324 ms 64 - 72 MB NPU
encoder PRECOMPILED_QNN_ONNX w8a16 Snapdragon® 8 Elite For Galaxy Mobile 34.495 ms 63 - 75 MB NPU
encoder PRECOMPILED_QNN_ONNX w8a16 Snapdragon® X2 Elite 28.665 ms 67 - 67 MB NPU
encoder PRECOMPILED_QNN_ONNX w8a16 Snapdragon® X Elite 60.242 ms 108 - 108 MB NPU
encoder PRECOMPILED_QNN_ONNX w8a16 Snapdragon® 8 Gen 3 Mobile 44.446 ms 64 - 76 MB NPU
encoder PRECOMPILED_QNN_ONNX w8a16 Snapdragon® 8 Gen 1 Mobile 80.49 ms 63 - 79 MB NPU
encoder PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ QCS6490 235.433 ms 62 - 66 MB NPU
encoder PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ IQ-8275 56.374 ms 62 - 67 MB NPU
encoder PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ QCS8550 (Proxy) 59.93 ms 0 - 113 MB NPU
encoder PRECOMPILED_QNN_ONNX w8a16 Qualcomm® QCS8450 80.49 ms 63 - 79 MB NPU
encoder PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ IQ-9075 59.469 ms 63 - 67 MB NPU
encoder PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ IQ-X7181 60.242 ms 108 - 108 MB NPU
encoder PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ Q-6690 876.302 ms 30 - 42 MB NPU
encoder PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ Q-7790 94.392 ms 65 - 72 MB NPU
encoder PRECOMPILED_QNN_ONNX w8a16 Qualcomm® Dragonwing™ Q-8750 34.495 ms 63 - 75 MB NPU
encoder PRECOMPILED_QNN_ONNX w8a16 Snapdragon® 7 Gen 4 Mobile 94.392 ms 65 - 72 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Snapdragon® 8 Elite Gen 5 For Galaxy Mobile 27.207 ms 1 - 9 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Snapdragon® 8 Elite For Galaxy Mobile 34.453 ms 1 - 9 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Snapdragon® X2 Elite 35.197 ms 0 - 0 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Snapdragon® X Elite 59.766 ms 0 - 0 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Snapdragon® 8 Gen 3 Mobile 43.435 ms 1 - 8 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Snapdragon® 8 Gen 1 Mobile 80.125 ms 0 - 10 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ QCS6490 233.303 ms 2 - 32 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ IQ-8275 56.459 ms 0 - 30 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ QCS8550 (Proxy) 60.294 ms 1 - 34 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Qualcomm® SA8650P 315.467 ms 1 - 11 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Qualcomm® SA8255P 315.467 ms 1 - 11 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Qualcomm® QCS8450 80.125 ms 0 - 10 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ IQ-9075 58.878 ms 0 - 30 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ IQ-X7181 59.766 ms 0 - 0 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ Q-6690 736.074 ms 3 - 9 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ Q-7790 95.221 ms 1 - 7 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Qualcomm® Dragonwing™ Q-8750 34.453 ms 1 - 9 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Qualcomm® SA8295P 76.177 ms 0 - 5 MB NPU
encoder QNN_CONTEXT_BINARY w8a16 Snapdragon® 7 Gen 4 Mobile 95.221 ms 1 - 7 MB NPU
encoder VOICE_AI w8a16 Snapdragon® 8 Elite Gen 5 For Galaxy Mobile 27.197 ms 1 - 9 MB NPU
encoder VOICE_AI w8a16 Snapdragon® 8 Elite For Galaxy Mobile 34.502 ms 1 - 10 MB NPU
encoder VOICE_AI w8a16 Snapdragon® X2 Elite 28.396 ms 0 - 0 MB NPU
encoder VOICE_AI w8a16 Snapdragon® X Elite 59.757 ms 0 - 0 MB NPU
encoder VOICE_AI w8a16 Snapdragon® 8 Gen 3 Mobile 43.444 ms 3 - 11 MB NPU
encoder VOICE_AI w8a16 Snapdragon® 8 Gen 1 Mobile 78.738 ms 1 - 14 MB NPU
encoder VOICE_AI w8a16 Qualcomm® Dragonwing™ QCS6490 235.609 ms 0 - 29 MB NPU
encoder VOICE_AI w8a16 Qualcomm® Dragonwing™ IQ-8275 56.188 ms 0 - 30 MB NPU
encoder VOICE_AI w8a16 Qualcomm® Dragonwing™ QCS8550 (Proxy) 59.911 ms 0 - 3 MB NPU
encoder VOICE_AI w8a16 Qualcomm® SA8650P 315.503 ms 0 - 9 MB NPU
encoder VOICE_AI w8a16 Qualcomm® SA8255P 315.503 ms 0 - 9 MB NPU
encoder VOICE_AI w8a16 Qualcomm® QCS8450 78.738 ms 1 - 14 MB NPU
encoder VOICE_AI w8a16 Qualcomm® Dragonwing™ IQ-9075 59.305 ms 0 - 30 MB NPU
encoder VOICE_AI w8a16 Qualcomm® Dragonwing™ IQ-X7181 59.757 ms 0 - 0 MB NPU
encoder VOICE_AI w8a16 Qualcomm® Dragonwing™ Q-6690 772.348 ms 3 - 9 MB NPU
encoder VOICE_AI w8a16 Qualcomm® Dragonwing™ Q-7790 93.335 ms 1 - 7 MB NPU
encoder VOICE_AI w8a16 Qualcomm® Dragonwing™ Q-8750 34.502 ms 1 - 10 MB NPU
encoder VOICE_AI w8a16 Qualcomm® SA8295P 76.277 ms 0 - 5 MB NPU
encoder VOICE_AI w8a16 Snapdragon® 7 Gen 4 Mobile 93.335 ms 1 - 7 MB NPU

License

  • The license for the original implementation of Whisper-Small-Quantized can be found here.

References

Community

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support