Artem Plastinkin commited on
Commit
ec1e72e
·
1 Parent(s): 7a845cb

Add model

Browse files
README.md CHANGED
@@ -15,8 +15,6 @@ tags:
15
 
16
  # CenterNet-R18 (ONNX) – Renesas X5H
17
 
18
- > ⏳ **Model file not yet uploaded.** Benchmark results on this page were published ahead of the model weights — see **Provided Artifacts** below. Download/deployment steps will not work until the file is added to this repository.
19
-
20
  ## Introduction
21
 
22
  This repository hosts **CenterNet** with a **ResNet18** backbone, targeting the **Renesas
@@ -45,7 +43,7 @@ centernet_r18_..._optimized.onnx (FP32)
45
 
46
  | Artifact | Status | Notes |
47
  |----------|--------|-------|
48
- | **FP32 (ONNX)** | ⏳ Not yet uploaded | Benchmark numbers below exist; the model file has not been published to this repo yet |
49
 
50
  ## Performance
51
 
@@ -56,6 +54,7 @@ Measured on **Renesas R-Car X5H** via the MWMX runtime (APM80 ship-performance C
56
  | Runtime | Precision | Device | Latency (ms) | Type |
57
  |---------|-----------|--------|---------------|------|
58
  | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | 11.978 | Measured |
 
59
  | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 12 Cores · 850 MHz | 3.467 | Measured |
60
 
61
  ### Accuracy
@@ -81,11 +80,13 @@ To run inference on Renesas R-Car X5H, you need:
81
 
82
  1. **Renesas R-Car X5H board** with NPX6 NPU
83
  2. **Renesas MWMX Runtime**
84
- 3. **Hugging Face CLI** to download the model (once the model file is published)
85
 
86
  ## Download
87
 
88
- TBD — model file not yet published to this repository.
 
 
89
 
90
  ---
91
 
 
15
 
16
  # CenterNet-R18 (ONNX) – Renesas X5H
17
 
 
 
18
  ## Introduction
19
 
20
  This repository hosts **CenterNet** with a **ResNet18** backbone, targeting the **Renesas
 
43
 
44
  | Artifact | Status | Notes |
45
  |----------|--------|-------|
46
+ | **FP32 (ONNX)** | ✅ Published | `fp32/centernet_r18_8xb16_crop512_140e_coco.onnx` — auto-cast to INT8 by the MWMX toolchain at compile time (see Deployment Flow above); no separate INT8 file is shipped |
47
 
48
  ## Performance
49
 
 
54
  | Runtime | Precision | Device | Latency (ms) | Type |
55
  |---------|-----------|--------|---------------|------|
56
  | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | 11.978 | Measured |
57
+ | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | 11.989 | Measured (2026-09-16) |
58
  | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 12 Cores · 850 MHz | 3.467 | Measured |
59
 
60
  ### Accuracy
 
80
 
81
  1. **Renesas R-Car X5H board** with NPX6 NPU
82
  2. **Renesas MWMX Runtime**
83
+ 3. **Hugging Face CLI** to download the model
84
 
85
  ## Download
86
 
87
+ ```bash
88
+ hf download Renesas/CenterNet-R18-ONNX --repo-type=model --include "fp32/*"
89
+ ```
90
 
91
  ---
92
 
compile_config/l2seg_csm_16mb.json ADDED
@@ -0,0 +1,69 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "segments": [
3
+ {
4
+ "name": "Segment_0",
5
+ "inputs": [
6
+ "input"
7
+ ],
8
+ "outputs": [
9
+ "/neck/deconv_layers/deconv_layers.5/activate/Relu_output_0"
10
+ ]
11
+ },
12
+ {
13
+ "name": "Segment_4",
14
+ "inputs": [
15
+ "/neck/deconv_layers/deconv_layers.5/activate/Relu_output_0"
16
+ ],
17
+ "outputs": [
18
+ "/bbox_head/Sigmoid_output_0"
19
+ ]
20
+ },
21
+ {
22
+ "name": "Segment_5",
23
+ "inputs": [
24
+ "/neck/deconv_layers/deconv_layers.5/activate/Relu_output_0"
25
+ ],
26
+ "outputs": [
27
+ "/bbox_head/wh_head/wh_head.2/Conv_output_0"
28
+ ]
29
+ },
30
+ {
31
+ "name": "Segment_6",
32
+ "inputs": [
33
+ "/neck/deconv_layers/deconv_layers.5/activate/Relu_output_0"
34
+ ],
35
+ "outputs": [
36
+ "/bbox_head/offset_head/offset_head.2/Conv_output_0"
37
+ ]
38
+ }
39
+ ],
40
+ "layer_groups": [
41
+ {
42
+ "name": "LG0",
43
+ "inputs": [
44
+ "/backbone/maxpool/MaxPool_output_0"
45
+ ],
46
+ "outputs": [
47
+ "/backbone/layer1/layer1.0/relu_1/Relu_output_0"
48
+ ]
49
+ },
50
+ {
51
+ "name": "LG1",
52
+ "inputs": [
53
+ "/backbone/layer1/layer1.0/relu_1/Relu_output_0"
54
+ ],
55
+ "outputs": [
56
+ "/backbone/layer1/layer1.1/relu_1/Relu_output_0"
57
+ ]
58
+ },
59
+ {
60
+ "name": "LG2",
61
+ "inputs": [
62
+ "/backbone/layer1/layer1.1/relu_1/Relu_output_0"
63
+ ],
64
+ "outputs": [
65
+ "/backbone/layer2/layer2.0/conv2/Conv_output_0"
66
+ ]
67
+ }
68
+ ]
69
+ }
fp32/.metadata.yaml ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ variant:
2
+ id: onnx_fp32
3
+ format: onnx
4
+ precision: fp32
5
+ method: onnx_export # exported straight from the OpenMMLab checkpoint, no quantization
fp32/centernet_r18_8xb16_crop512_140e_coco.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:279c1b2d971c9d2488b137a6115805115c4c875f505b160d8ff968031152f888
3
+ size 56884157
int8/benchmarks/x5h_mwmx_npu_apm80_12core.yaml CHANGED
@@ -40,3 +40,51 @@ memory:
40
 
41
  power:
42
  avg_w: null
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
40
 
41
  power:
42
  avg_w: null
43
+
44
+ # Exact commands verified against the NNAC "Getting Started" chapter for this
45
+ # model (it is that chapter's own worked example), plus the network_configuration
46
+ # block verified against the internal model_zoo compile config for this model's
47
+ # 12-core slice (see compile_config/l2seg_csm_16mb.json). Rendered by the
48
+ # AI-Dashboard in place of the generic placeholder flow — see
49
+ # downloadRunHTML() / parse_reproduce(). NOTE: the 1-core vs. 12-core split
50
+ # (see `configuration` above) is set at compile time via
51
+ # ${WORKDIR}/nnac_config/config.yaml, not a host_app flag — the on-board run
52
+ # command below is identical to the 1-core run.
53
+ reproduce:
54
+ steps:
55
+ - title: Download the ONNX model and compile config
56
+ command: hf download Renesas/CenterNet-R18-ONNX --repo-type=model --include "fp32/*" --include "compile_config/*"
57
+ - title: Configure the network for the 12-core build
58
+ command: |
59
+ cat >> ${WORKDIR}/nnac_config/config.yaml << 'EOF'
60
+ network_configuration:
61
+ centernet_r18_8xb16_crop512_140e_coco:
62
+ last_nodes:
63
+ - /Reshape_output_0
64
+ - /Flatten_output_0
65
+ - /Reshape_1_output_0
66
+ global_config:
67
+ weight_symmetric: true
68
+ optimization_options:
69
+ - -segmentation_json ./compile_config/l2seg_csm_16mb.json
70
+ - -enable_shallow_conv_opt
71
+ - -enable_dma_fusion
72
+ - -stu_weight
73
+ - -distributed_coef_stu
74
+ EOF
75
+ # If config.yaml already has a network_configuration: key (e.g. from a
76
+ # previous model), merge this block into it instead of appending a duplicate key.
77
+ - title: Compile with the NNAC toolchain (INT8 auto-cast from the FP32 graph)
78
+ command: |
79
+ python3 nnac_frontend/legalize.py -d binary/nnx ./fp32/centernet_r18_8xb16_crop512_140e_coco.onnx
80
+ - title: Set up the R-Car X5H board
81
+ command: Configure the board per the AI Compiler (NNAC) "Getting Started" guide, section 3.4 (host TFTP/NFS setup, bootloader flashing, U-Boot, Linux boot, login) -- exact steps depend on your board/network setup.
82
+ - title: Copy the compiled artifact to the board
83
+ command: Copy ${WORKDIR}/binary/nnx to the board -- method may vary (NFS mount, scp, USB, etc.).
84
+ - title: Run on R-Car X5H (single NPU cluster, 12 AI cores)
85
+ command: |
86
+ cd binary
87
+ ./host_app ./arc_prog_npus ./nnx/centernet_r18_8xb16_crop512_140e_coco
88
+ notes: >-
89
+ hash[n] = 0x...(OK) means the run's output matches the reference hash in
90
+ hash.txt; latency is the NPX execution time reported in cycles and ms.
int8/benchmarks/x5h_mwmx_npu_apm80_1core.yaml CHANGED
@@ -40,3 +40,25 @@ memory:
40
 
41
  power:
42
  avg_w: null
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
40
 
41
  power:
42
  avg_w: null
43
+
44
+ # Exact commands verified against the NNAC "Getting Started" chapter for this
45
+ # model (it is that chapter's own worked example). Rendered by the AI-Dashboard
46
+ # in place of the generic placeholder flow — see downloadRunHTML() / parse_reproduce().
47
+ reproduce:
48
+ steps:
49
+ - title: Download the ONNX model
50
+ command: hf download Renesas/CenterNet-R18-ONNX --repo-type=model --include "fp32/*"
51
+ - title: Compile with the NNAC toolchain (INT8 auto-cast from the FP32 graph)
52
+ command: |
53
+ python3 nnac_frontend/legalize.py -d binary/nnx ./fp32/centernet_r18_8xb16_crop512_140e_coco.onnx
54
+ - title: Set up the R-Car X5H board
55
+ command: Configure the board per the AI Compiler (NNAC) "Getting Started" guide, section 3.4 (host TFTP/NFS setup, bootloader flashing, U-Boot, Linux boot, login) -- exact steps depend on your board/network setup.
56
+ - title: Copy the compiled artifact to the board
57
+ command: Copy ${WORKDIR}/binary/nnx to the board -- method may vary (NFS mount, scp, USB, etc.).
58
+ - title: Run on R-Car X5H (single NPU cluster)
59
+ command: |
60
+ cd binary
61
+ ./host_app ./arc_prog_npus ./nnx/centernet_r18_8xb16_crop512_140e_coco
62
+ notes: >-
63
+ hash[n] = 0x...(OK) means the run's output matches the reference hash in
64
+ hash.txt; latency is the NPX execution time reported in cycles and ms.
int8/benchmarks/x5h_mwmx_npu_apm80_1core_2026-09-16.yaml ADDED
@@ -0,0 +1,65 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Re-run captured 2026-09-16 (gf_model_result.csv export) — see also the original file in this directory
2
+ hardware:
3
+ vendor: renesas
4
+ chip: rcar-x5h
5
+ cpu: arm-cortex-a720
6
+ npu: npx6-48k
7
+ npu_count: 2
8
+ npu_cores: 12
9
+ npu_default_freq_mhz: 1066
10
+ accelerator:
11
+ - npu
12
+
13
+ runtime:
14
+ engine: mwmx
15
+ toolchain_version: "MWMX SDK v4.35.0"
16
+ format: onnx
17
+ execution_provider: npu
18
+ execution_precision: int8
19
+
20
+ configuration:
21
+ npu_instances: 1
22
+ npu_cores_per_instance: 1
23
+ npu_freq_mhz: 850
24
+
25
+ benchmark:
26
+ type: hil
27
+ parameters:
28
+ batch_size: 1
29
+ input_resolution: [1, 3, 512, 512] # inferred from "crop512" in the source checkpoint name
30
+
31
+ performance:
32
+ fps: null
33
+ latency: 11.989098 # APM80 pipeline run (metawaremx_runtime CI)
34
+
35
+ metrics:
36
+ accuracy: null
37
+ top5_accuracy: null
38
+
39
+ memory:
40
+ peak_mb: null
41
+
42
+ power:
43
+ avg_w: null
44
+
45
+ # Exact commands verified against the NNAC "Getting Started" chapter for this
46
+ # model (it is that chapter's own worked example). Rendered by the AI-Dashboard
47
+ # in place of the generic placeholder flow — see downloadRunHTML() / parse_reproduce().
48
+ reproduce:
49
+ steps:
50
+ - title: Download the ONNX model
51
+ command: hf download Renesas/CenterNet-R18-ONNX --repo-type=model --include "fp32/*"
52
+ - title: Compile with the NNAC toolchain (INT8 auto-cast from the FP32 graph)
53
+ command: |
54
+ python3 nnac_frontend/legalize.py -d binary/nnx ./fp32/centernet_r18_8xb16_crop512_140e_coco.onnx
55
+ - title: Set up the R-Car X5H board
56
+ command: Configure the board per the AI Compiler (NNAC) "Getting Started" guide, section 3.4 (host TFTP/NFS setup, bootloader flashing, U-Boot, Linux boot, login) -- exact steps depend on your board/network setup.
57
+ - title: Copy the compiled artifact to the board
58
+ command: Copy ${WORKDIR}/binary/nnx to the board -- method may vary (NFS mount, scp, USB, etc.).
59
+ - title: Run on R-Car X5H (single NPU cluster)
60
+ command: |
61
+ cd binary
62
+ ./host_app ./arc_prog_npus ./nnx/centernet_r18_8xb16_crop512_140e_coco
63
+ notes: >-
64
+ hash[n] = 0x...(OK) means the run's output matches the reference hash in
65
+ hash.txt; latency is the NPX execution time reported in cycles and ms.