File size: 1,488 Bytes
1ebfbbf
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
# Performance Report

- Board: AX650N (`10.126.29.50`)
- Runtime: AX Engine 2.12.0s / pyaxengine `AxEngineExecutionProvider`
- Model: H-GTCRN neural core (`feat -> mask`)
- Sample duration for RTF: 10.0 s
- Input tensor: `feat[1, 6, 626, 257]`, FP32
- Output tensor: `mask[1, 2, 626, 257]`, FP32

| Runtime | First run | 20-run average | Core RTF |
|---|---:|---:|---:|
| Python SDK | 31.831 ms | 21.142 ms | 0.002114 |
| C++ SDK | 22.023 ms | 20.123 ms | 0.002012 |

| Check | Value |
|---|---:|
| Python/C++ output cosine | 1.0 |
| Python/C++ output MAE | 0.0 |
| Python/C++ max absolute difference | 0.0 |
| Board/reference mask cosine | 0.99876 |
| Board/original infer audio cosine | 0.999468371 |
| Board/original infer audio MAE | 0.000655536079 |
| ONNX/original infer audio cosine | 0.9999999998 |
| ONNX/original infer audio MAE | 0.0000000109 |

Audio sample:

- Input: `samples/Samples1_noisy.wav`, 16 kHz, 10.0 s, 2 channels, RMS 0.09499
- Original `infer.py` output RMS: 0.02980
- Board output: `samples/Samples1_board_enhanced.wav`, RMS 0.02932

The RTF number covers only the compiled NPU neural core. CPU-side STFT/WPE/IVA,
feature construction, mask application, and ISTFT are not included.

The export was fixed on 2026-08-13: `FoldedConvTranspose2d` now preserves the
original activation (`PReLU`/`Tanh`) after folding BatchNorm into ConvTranspose.
Before this fix, ONNX and AXMODEL were mutually consistent but did not match the
original `GTCRN_IVA.forward()` audio.