Artem Plastinkin commited on
Commit ·
ec1e72e
1
Parent(s): 7a845cb
Add model
Browse files- README.md +6 -5
- compile_config/l2seg_csm_16mb.json +69 -0
- fp32/.metadata.yaml +5 -0
- fp32/centernet_r18_8xb16_crop512_140e_coco.onnx +3 -0
- int8/benchmarks/x5h_mwmx_npu_apm80_12core.yaml +48 -0
- int8/benchmarks/x5h_mwmx_npu_apm80_1core.yaml +22 -0
- int8/benchmarks/x5h_mwmx_npu_apm80_1core_2026-09-16.yaml +65 -0
README.md
CHANGED
|
@@ -15,8 +15,6 @@ tags:
|
|
| 15 |
|
| 16 |
# CenterNet-R18 (ONNX) – Renesas X5H
|
| 17 |
|
| 18 |
-
> ⏳ **Model file not yet uploaded.** Benchmark results on this page were published ahead of the model weights — see **Provided Artifacts** below. Download/deployment steps will not work until the file is added to this repository.
|
| 19 |
-
|
| 20 |
## Introduction
|
| 21 |
|
| 22 |
This repository hosts **CenterNet** with a **ResNet18** backbone, targeting the **Renesas
|
|
@@ -45,7 +43,7 @@ centernet_r18_..._optimized.onnx (FP32)
|
|
| 45 |
|
| 46 |
| Artifact | Status | Notes |
|
| 47 |
|----------|--------|-------|
|
| 48 |
-
| **FP32 (ONNX)** |
|
| 49 |
|
| 50 |
## Performance
|
| 51 |
|
|
@@ -56,6 +54,7 @@ Measured on **Renesas R-Car X5H** via the MWMX runtime (APM80 ship-performance C
|
|
| 56 |
| Runtime | Precision | Device | Latency (ms) | Type |
|
| 57 |
|---------|-----------|--------|---------------|------|
|
| 58 |
| MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | 11.978 | Measured |
|
|
|
|
| 59 |
| MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 12 Cores · 850 MHz | 3.467 | Measured |
|
| 60 |
|
| 61 |
### Accuracy
|
|
@@ -81,11 +80,13 @@ To run inference on Renesas R-Car X5H, you need:
|
|
| 81 |
|
| 82 |
1. **Renesas R-Car X5H board** with NPX6 NPU
|
| 83 |
2. **Renesas MWMX Runtime**
|
| 84 |
-
3. **Hugging Face CLI** to download the model
|
| 85 |
|
| 86 |
## Download
|
| 87 |
|
| 88 |
-
|
|
|
|
|
|
|
| 89 |
|
| 90 |
---
|
| 91 |
|
|
|
|
| 15 |
|
| 16 |
# CenterNet-R18 (ONNX) – Renesas X5H
|
| 17 |
|
|
|
|
|
|
|
| 18 |
## Introduction
|
| 19 |
|
| 20 |
This repository hosts **CenterNet** with a **ResNet18** backbone, targeting the **Renesas
|
|
|
|
| 43 |
|
| 44 |
| Artifact | Status | Notes |
|
| 45 |
|----------|--------|-------|
|
| 46 |
+
| **FP32 (ONNX)** | ✅ Published | `fp32/centernet_r18_8xb16_crop512_140e_coco.onnx` — auto-cast to INT8 by the MWMX toolchain at compile time (see Deployment Flow above); no separate INT8 file is shipped |
|
| 47 |
|
| 48 |
## Performance
|
| 49 |
|
|
|
|
| 54 |
| Runtime | Precision | Device | Latency (ms) | Type |
|
| 55 |
|---------|-----------|--------|---------------|------|
|
| 56 |
| MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | 11.978 | Measured |
|
| 57 |
+
| MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | 11.989 | Measured (2026-09-16) |
|
| 58 |
| MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 12 Cores · 850 MHz | 3.467 | Measured |
|
| 59 |
|
| 60 |
### Accuracy
|
|
|
|
| 80 |
|
| 81 |
1. **Renesas R-Car X5H board** with NPX6 NPU
|
| 82 |
2. **Renesas MWMX Runtime**
|
| 83 |
+
3. **Hugging Face CLI** to download the model
|
| 84 |
|
| 85 |
## Download
|
| 86 |
|
| 87 |
+
```bash
|
| 88 |
+
hf download Renesas/CenterNet-R18-ONNX --repo-type=model --include "fp32/*"
|
| 89 |
+
```
|
| 90 |
|
| 91 |
---
|
| 92 |
|
compile_config/l2seg_csm_16mb.json
ADDED
|
@@ -0,0 +1,69 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"segments": [
|
| 3 |
+
{
|
| 4 |
+
"name": "Segment_0",
|
| 5 |
+
"inputs": [
|
| 6 |
+
"input"
|
| 7 |
+
],
|
| 8 |
+
"outputs": [
|
| 9 |
+
"/neck/deconv_layers/deconv_layers.5/activate/Relu_output_0"
|
| 10 |
+
]
|
| 11 |
+
},
|
| 12 |
+
{
|
| 13 |
+
"name": "Segment_4",
|
| 14 |
+
"inputs": [
|
| 15 |
+
"/neck/deconv_layers/deconv_layers.5/activate/Relu_output_0"
|
| 16 |
+
],
|
| 17 |
+
"outputs": [
|
| 18 |
+
"/bbox_head/Sigmoid_output_0"
|
| 19 |
+
]
|
| 20 |
+
},
|
| 21 |
+
{
|
| 22 |
+
"name": "Segment_5",
|
| 23 |
+
"inputs": [
|
| 24 |
+
"/neck/deconv_layers/deconv_layers.5/activate/Relu_output_0"
|
| 25 |
+
],
|
| 26 |
+
"outputs": [
|
| 27 |
+
"/bbox_head/wh_head/wh_head.2/Conv_output_0"
|
| 28 |
+
]
|
| 29 |
+
},
|
| 30 |
+
{
|
| 31 |
+
"name": "Segment_6",
|
| 32 |
+
"inputs": [
|
| 33 |
+
"/neck/deconv_layers/deconv_layers.5/activate/Relu_output_0"
|
| 34 |
+
],
|
| 35 |
+
"outputs": [
|
| 36 |
+
"/bbox_head/offset_head/offset_head.2/Conv_output_0"
|
| 37 |
+
]
|
| 38 |
+
}
|
| 39 |
+
],
|
| 40 |
+
"layer_groups": [
|
| 41 |
+
{
|
| 42 |
+
"name": "LG0",
|
| 43 |
+
"inputs": [
|
| 44 |
+
"/backbone/maxpool/MaxPool_output_0"
|
| 45 |
+
],
|
| 46 |
+
"outputs": [
|
| 47 |
+
"/backbone/layer1/layer1.0/relu_1/Relu_output_0"
|
| 48 |
+
]
|
| 49 |
+
},
|
| 50 |
+
{
|
| 51 |
+
"name": "LG1",
|
| 52 |
+
"inputs": [
|
| 53 |
+
"/backbone/layer1/layer1.0/relu_1/Relu_output_0"
|
| 54 |
+
],
|
| 55 |
+
"outputs": [
|
| 56 |
+
"/backbone/layer1/layer1.1/relu_1/Relu_output_0"
|
| 57 |
+
]
|
| 58 |
+
},
|
| 59 |
+
{
|
| 60 |
+
"name": "LG2",
|
| 61 |
+
"inputs": [
|
| 62 |
+
"/backbone/layer1/layer1.1/relu_1/Relu_output_0"
|
| 63 |
+
],
|
| 64 |
+
"outputs": [
|
| 65 |
+
"/backbone/layer2/layer2.0/conv2/Conv_output_0"
|
| 66 |
+
]
|
| 67 |
+
}
|
| 68 |
+
]
|
| 69 |
+
}
|
fp32/.metadata.yaml
ADDED
|
@@ -0,0 +1,5 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
variant:
|
| 2 |
+
id: onnx_fp32
|
| 3 |
+
format: onnx
|
| 4 |
+
precision: fp32
|
| 5 |
+
method: onnx_export # exported straight from the OpenMMLab checkpoint, no quantization
|
fp32/centernet_r18_8xb16_crop512_140e_coco.onnx
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:279c1b2d971c9d2488b137a6115805115c4c875f505b160d8ff968031152f888
|
| 3 |
+
size 56884157
|
int8/benchmarks/x5h_mwmx_npu_apm80_12core.yaml
CHANGED
|
@@ -40,3 +40,51 @@ memory:
|
|
| 40 |
|
| 41 |
power:
|
| 42 |
avg_w: null
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 40 |
|
| 41 |
power:
|
| 42 |
avg_w: null
|
| 43 |
+
|
| 44 |
+
# Exact commands verified against the NNAC "Getting Started" chapter for this
|
| 45 |
+
# model (it is that chapter's own worked example), plus the network_configuration
|
| 46 |
+
# block verified against the internal model_zoo compile config for this model's
|
| 47 |
+
# 12-core slice (see compile_config/l2seg_csm_16mb.json). Rendered by the
|
| 48 |
+
# AI-Dashboard in place of the generic placeholder flow — see
|
| 49 |
+
# downloadRunHTML() / parse_reproduce(). NOTE: the 1-core vs. 12-core split
|
| 50 |
+
# (see `configuration` above) is set at compile time via
|
| 51 |
+
# ${WORKDIR}/nnac_config/config.yaml, not a host_app flag — the on-board run
|
| 52 |
+
# command below is identical to the 1-core run.
|
| 53 |
+
reproduce:
|
| 54 |
+
steps:
|
| 55 |
+
- title: Download the ONNX model and compile config
|
| 56 |
+
command: hf download Renesas/CenterNet-R18-ONNX --repo-type=model --include "fp32/*" --include "compile_config/*"
|
| 57 |
+
- title: Configure the network for the 12-core build
|
| 58 |
+
command: |
|
| 59 |
+
cat >> ${WORKDIR}/nnac_config/config.yaml << 'EOF'
|
| 60 |
+
network_configuration:
|
| 61 |
+
centernet_r18_8xb16_crop512_140e_coco:
|
| 62 |
+
last_nodes:
|
| 63 |
+
- /Reshape_output_0
|
| 64 |
+
- /Flatten_output_0
|
| 65 |
+
- /Reshape_1_output_0
|
| 66 |
+
global_config:
|
| 67 |
+
weight_symmetric: true
|
| 68 |
+
optimization_options:
|
| 69 |
+
- -segmentation_json ./compile_config/l2seg_csm_16mb.json
|
| 70 |
+
- -enable_shallow_conv_opt
|
| 71 |
+
- -enable_dma_fusion
|
| 72 |
+
- -stu_weight
|
| 73 |
+
- -distributed_coef_stu
|
| 74 |
+
EOF
|
| 75 |
+
# If config.yaml already has a network_configuration: key (e.g. from a
|
| 76 |
+
# previous model), merge this block into it instead of appending a duplicate key.
|
| 77 |
+
- title: Compile with the NNAC toolchain (INT8 auto-cast from the FP32 graph)
|
| 78 |
+
command: |
|
| 79 |
+
python3 nnac_frontend/legalize.py -d binary/nnx ./fp32/centernet_r18_8xb16_crop512_140e_coco.onnx
|
| 80 |
+
- title: Set up the R-Car X5H board
|
| 81 |
+
command: Configure the board per the AI Compiler (NNAC) "Getting Started" guide, section 3.4 (host TFTP/NFS setup, bootloader flashing, U-Boot, Linux boot, login) -- exact steps depend on your board/network setup.
|
| 82 |
+
- title: Copy the compiled artifact to the board
|
| 83 |
+
command: Copy ${WORKDIR}/binary/nnx to the board -- method may vary (NFS mount, scp, USB, etc.).
|
| 84 |
+
- title: Run on R-Car X5H (single NPU cluster, 12 AI cores)
|
| 85 |
+
command: |
|
| 86 |
+
cd binary
|
| 87 |
+
./host_app ./arc_prog_npus ./nnx/centernet_r18_8xb16_crop512_140e_coco
|
| 88 |
+
notes: >-
|
| 89 |
+
hash[n] = 0x...(OK) means the run's output matches the reference hash in
|
| 90 |
+
hash.txt; latency is the NPX execution time reported in cycles and ms.
|
int8/benchmarks/x5h_mwmx_npu_apm80_1core.yaml
CHANGED
|
@@ -40,3 +40,25 @@ memory:
|
|
| 40 |
|
| 41 |
power:
|
| 42 |
avg_w: null
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 40 |
|
| 41 |
power:
|
| 42 |
avg_w: null
|
| 43 |
+
|
| 44 |
+
# Exact commands verified against the NNAC "Getting Started" chapter for this
|
| 45 |
+
# model (it is that chapter's own worked example). Rendered by the AI-Dashboard
|
| 46 |
+
# in place of the generic placeholder flow — see downloadRunHTML() / parse_reproduce().
|
| 47 |
+
reproduce:
|
| 48 |
+
steps:
|
| 49 |
+
- title: Download the ONNX model
|
| 50 |
+
command: hf download Renesas/CenterNet-R18-ONNX --repo-type=model --include "fp32/*"
|
| 51 |
+
- title: Compile with the NNAC toolchain (INT8 auto-cast from the FP32 graph)
|
| 52 |
+
command: |
|
| 53 |
+
python3 nnac_frontend/legalize.py -d binary/nnx ./fp32/centernet_r18_8xb16_crop512_140e_coco.onnx
|
| 54 |
+
- title: Set up the R-Car X5H board
|
| 55 |
+
command: Configure the board per the AI Compiler (NNAC) "Getting Started" guide, section 3.4 (host TFTP/NFS setup, bootloader flashing, U-Boot, Linux boot, login) -- exact steps depend on your board/network setup.
|
| 56 |
+
- title: Copy the compiled artifact to the board
|
| 57 |
+
command: Copy ${WORKDIR}/binary/nnx to the board -- method may vary (NFS mount, scp, USB, etc.).
|
| 58 |
+
- title: Run on R-Car X5H (single NPU cluster)
|
| 59 |
+
command: |
|
| 60 |
+
cd binary
|
| 61 |
+
./host_app ./arc_prog_npus ./nnx/centernet_r18_8xb16_crop512_140e_coco
|
| 62 |
+
notes: >-
|
| 63 |
+
hash[n] = 0x...(OK) means the run's output matches the reference hash in
|
| 64 |
+
hash.txt; latency is the NPX execution time reported in cycles and ms.
|
int8/benchmarks/x5h_mwmx_npu_apm80_1core_2026-09-16.yaml
ADDED
|
@@ -0,0 +1,65 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Re-run captured 2026-09-16 (gf_model_result.csv export) — see also the original file in this directory
|
| 2 |
+
hardware:
|
| 3 |
+
vendor: renesas
|
| 4 |
+
chip: rcar-x5h
|
| 5 |
+
cpu: arm-cortex-a720
|
| 6 |
+
npu: npx6-48k
|
| 7 |
+
npu_count: 2
|
| 8 |
+
npu_cores: 12
|
| 9 |
+
npu_default_freq_mhz: 1066
|
| 10 |
+
accelerator:
|
| 11 |
+
- npu
|
| 12 |
+
|
| 13 |
+
runtime:
|
| 14 |
+
engine: mwmx
|
| 15 |
+
toolchain_version: "MWMX SDK v4.35.0"
|
| 16 |
+
format: onnx
|
| 17 |
+
execution_provider: npu
|
| 18 |
+
execution_precision: int8
|
| 19 |
+
|
| 20 |
+
configuration:
|
| 21 |
+
npu_instances: 1
|
| 22 |
+
npu_cores_per_instance: 1
|
| 23 |
+
npu_freq_mhz: 850
|
| 24 |
+
|
| 25 |
+
benchmark:
|
| 26 |
+
type: hil
|
| 27 |
+
parameters:
|
| 28 |
+
batch_size: 1
|
| 29 |
+
input_resolution: [1, 3, 512, 512] # inferred from "crop512" in the source checkpoint name
|
| 30 |
+
|
| 31 |
+
performance:
|
| 32 |
+
fps: null
|
| 33 |
+
latency: 11.989098 # APM80 pipeline run (metawaremx_runtime CI)
|
| 34 |
+
|
| 35 |
+
metrics:
|
| 36 |
+
accuracy: null
|
| 37 |
+
top5_accuracy: null
|
| 38 |
+
|
| 39 |
+
memory:
|
| 40 |
+
peak_mb: null
|
| 41 |
+
|
| 42 |
+
power:
|
| 43 |
+
avg_w: null
|
| 44 |
+
|
| 45 |
+
# Exact commands verified against the NNAC "Getting Started" chapter for this
|
| 46 |
+
# model (it is that chapter's own worked example). Rendered by the AI-Dashboard
|
| 47 |
+
# in place of the generic placeholder flow — see downloadRunHTML() / parse_reproduce().
|
| 48 |
+
reproduce:
|
| 49 |
+
steps:
|
| 50 |
+
- title: Download the ONNX model
|
| 51 |
+
command: hf download Renesas/CenterNet-R18-ONNX --repo-type=model --include "fp32/*"
|
| 52 |
+
- title: Compile with the NNAC toolchain (INT8 auto-cast from the FP32 graph)
|
| 53 |
+
command: |
|
| 54 |
+
python3 nnac_frontend/legalize.py -d binary/nnx ./fp32/centernet_r18_8xb16_crop512_140e_coco.onnx
|
| 55 |
+
- title: Set up the R-Car X5H board
|
| 56 |
+
command: Configure the board per the AI Compiler (NNAC) "Getting Started" guide, section 3.4 (host TFTP/NFS setup, bootloader flashing, U-Boot, Linux boot, login) -- exact steps depend on your board/network setup.
|
| 57 |
+
- title: Copy the compiled artifact to the board
|
| 58 |
+
command: Copy ${WORKDIR}/binary/nnx to the board -- method may vary (NFS mount, scp, USB, etc.).
|
| 59 |
+
- title: Run on R-Car X5H (single NPU cluster)
|
| 60 |
+
command: |
|
| 61 |
+
cd binary
|
| 62 |
+
./host_app ./arc_prog_npus ./nnx/centernet_r18_8xb16_crop512_140e_coco
|
| 63 |
+
notes: >-
|
| 64 |
+
hash[n] = 0x...(OK) means the run's output matches the reference hash in
|
| 65 |
+
hash.txt; latency is the NPX execution time reported in cycles and ms.
|