Performance Report
- Board: AX650N (
10.126.29.50) - Runtime: AX Engine 2.12.0s / pyaxengine
AxEngineExecutionProvider - Model: H-GTCRN neural core (
feat -> mask) - Sample duration for RTF: 10.0 s
- Input tensor:
feat[1, 6, 626, 257], FP32 - Output tensor:
mask[1, 2, 626, 257], FP32
| Runtime | First run | 20-run average | Core RTF |
|---|---|---|---|
| Python SDK | 31.831 ms | 21.142 ms | 0.002114 |
| C++ SDK | 22.023 ms | 20.123 ms | 0.002012 |
| Check | Value |
|---|---|
| Python/C++ output cosine | 1.0 |
| Python/C++ output MAE | 0.0 |
| Python/C++ max absolute difference | 0.0 |
| Board/reference mask cosine | 0.99876 |
| Board/original infer audio cosine | 0.999468371 |
| Board/original infer audio MAE | 0.000655536079 |
| ONNX/original infer audio cosine | 0.9999999998 |
| ONNX/original infer audio MAE | 0.0000000109 |
Audio sample:
- Input:
samples/Samples1_noisy.wav, 16 kHz, 10.0 s, 2 channels, RMS 0.09499 - Original
infer.pyoutput RMS: 0.02980 - Board output:
samples/Samples1_board_enhanced.wav, RMS 0.02932
The RTF number covers only the compiled NPU neural core. CPU-side STFT/WPE/IVA, feature construction, mask application, and ISTFT are not included.
The export was fixed on 2026-08-13: FoldedConvTranspose2d now preserves the
original activation (PReLU/Tanh) after folding BatchNorm into ConvTranspose.
Before this fix, ONNX and AXMODEL were mutually consistent but did not match the
original GTCRN_IVA.forward() audio.