Whisper Base for Allwinner A733 NPU (VIPLite) & CPU ONNX
This repository provides optimized models and pre-compiled binaries for running Whisper Base in a hybrid inference architecture on Single Board Computers powered by the Allwinner A733 SoC (e.g., Radxa Cubie A7A and compatible boards):
- Whisper Encoder: Compiled into a static NBG binary (
network_binary.nb) running directly on the Allwinner/VeriSilicon VIP9000 NPU via VIPLite runtime (ctypes, zero C-wrapper). - Whisper Decoder: Executed via
onnxruntimeon the CPU (base-decoder.onnx). - Wyoming Protocol: Ready for low-latency, private, on-device Speech-To-Text (ASR) integration with Home Assistant without PyTorch or OpenAI-Whisper dependencies.
π¦ Repository Files
| File | Description | Execution Target |
|---|---|---|
network_binary.nb |
15-second static Whisper encoder compiled for A733 NPU | NPU (VIPLite) |
base-decoder.onnx |
FP32 autoregressive Whisper decoder | CPU (ONNX Runtime) |
base-tokens.txt |
Vocabulary / Tokenizer token mapping | CPU / Python |
βοΈ Model Architecture & Specifications
To comply with the Allwinner A733 NPU engine requirements, the encoder has been exported with a fixed static tensor shape:
- Audio Duration: Fixed at 15 seconds (shorter audio must be zero-padded to 15s; longer audio must be truncated or split using VAD/external chunking).
- Audio Frontend: 16 kHz Mono $\to$ Log-Mel Spectrogram with 80 bins.
- Encoder Input Shape (NCHW):
[1, 80, 1500]FP16 (mapped to[1500, 80, 1]in VIPLite's[W,H,C,N]ordering). - Encoder Output Shape: 2 tensors (
cross_kandcross_v) with shape[6, 1, 750, 512]FP16 (mapped to[512, 750, 1, 6]in VIPLite), converted to FP32 for CPU decoding.
π οΈ System Prerequisites (Allwinner A733 SBC)
Ensure the board has the NPU kernel device and VIPLite dynamic libraries installed:
# Check device node
ls -l /dev/vipcore
# Verify VIPLite runtime shared libraries
/sbin/ldconfig -p | grep -E 'VIPhal|NBGlinker'
Required system paths:
/dev/vipcore
/lib/libVIPhal.so
/lib/libNBGlinker.so
If missing, obtain them from the official Radxa/Allwinner model zoo:
- Archive: allwinner-model-zoo.tar.gz
- Library path:
allwinner-model-zoo/common/npuruntime/lib_linux_aarch64/A733/
How network_binary.nb Was Built
The NBG model was generated from the k2-fsa/sherpa-onnx Whisper Base model using the Acuity/Pegasus toolchain inside a Docker/Podman container.
Step 1: Fix ONNX input shape to static (15s)
import onnx
model = onnx.load("base-encoder.onnx")
for inp in model.graph.input:
if inp.name == "mel":
dims = inp.type.tensor_type.shape.dim
dims[0].ClearField("dim_param")
dims[0].dim_value = 1 # Batch size 1
dims[2].ClearField("dim_param")
dims[2].dim_value = 1500 # 15 seconds (1500 frames)
onnx.checker.check_model(model)
onnx.save(model, "base-encoder-static-15s.onnx")
Step 2: Convert via Pegasus (Acuity Toolchain)
Inside container khalida5/ubuntu-npu:v2.0.10:
alias pegasus="python3 /usr/local/acuity_command_line_tools/pegasus.py"
# Import ONNX model
pegasus import onnx \
--model base-encoder-static-15s.onnx \
--output-data whisper_base_encoder.data \
--output-model whisper_base_encoder.json
# Generate metadata files
pegasus generate inputmeta \
--model whisper_base_encoder.json \
--input-meta-output model_inputmeta.yml
pegasus generate postprocess-file \
--model whisper_base_encoder.json \
--postprocess-file-output model_postprocess.yml
mkdir -p output_nbg
# Export compiled NBG for Allwinner A733 (VIP9000)
pegasus export ovxlib \
--model whisper_base_encoder.json \
--model-data whisper_base_encoder.data \
--dtype float \
--with-input-meta model_inputmeta.yml \
--postprocess-file model_postprocess.yml \
--target-ide-project linux64 \
--optimize VIP9000NANODI_PLUS_PID0X1000003B \
--viv-sdk "${VIV_SDK}" \
--pack-nbg-unify \
--output-path ./output_nbg/whisper_base_encoder
π·οΈ Credits
- Base Whisper architecture by OpenAI.
- ONNX export & vocab tokens by k2-fsa/sherpa-onnx.
- Allwinner A733 NPU SDK and toolchain references provided by Radxa.
Model tree for marcobarb94/a733-whisper-base
Base model
csukuangfj/sherpa-onnx-whisper-base