Artem Plastinkin commited on
Commit
249048c
·
1 Parent(s): 1c52e4b

Initial Commit

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ *.onnx_data filter=lfs diff=lfs merge=lfs -text
.metadata.yaml ADDED
@@ -0,0 +1,19 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ model:
2
+ name: retinanet
3
+ display_name: RetinaNet
4
+ upstream: onnx/retinanet-9
5
+
6
+ architecture:
7
+ family: retinanet
8
+ backbone: resnet101
9
+ modality:
10
+ - vision
11
+ parameters: null
12
+
13
+ format:
14
+ type: onnx
15
+ version: 1.6.0
16
+ opset: 9
17
+
18
+ tasks:
19
+ - object-detection
README.md ADDED
@@ -0,0 +1,232 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model:
4
+ - onnxmodelzoo/retinanet-9
5
+ pipeline_tag: object-detection
6
+ tags:
7
+ - object-detection
8
+ - computer-vision
9
+ - renesas
10
+ - x5h
11
+ - onnx
12
+ - retinanet
13
+ - resnet101
14
+ - detection
15
+ ---
16
+
17
+ # RetinaNet (ONNX) – Renesas X5H
18
+
19
+ ## Introduction
20
+
21
+ This repository hosts **RetinaNet** in ONNX FP32 format, targeting the **Renesas R-Car X5H** platform for object detection inference on the NPX6 NPU.
22
+
23
+ - **Model Architecture:** RetinaNet with ResNet101 backbone and Feature Pyramid Network (FPN)
24
+ - **Source Model:** ONNX Model Zoo RetinaNet
25
+ - **Task:** Object Detection
26
+ - **Dataset:** COCO
27
+ - **Accuracy:** mAP = 0.376
28
+ - **Backbone:** ResNet101
29
+
30
+ ## Deployment Flow
31
+
32
+ The repository provides the model in **FP32 ONNX** format. Both supported runtimes automatically cast the FP32 model to **INT8** at load time for optimised NPU execution — no separate quantization step is required.
33
+
34
+ ```text
35
+ retinanet-9.onnx (FP32)
36
+ │
37
+ ├─▶ ONNX Runtime (Custom NPU EP) ──▶ INT8 auto-cast ──▶ NPX6 NPU
38
+ │
39
+ └─▶ MWMX Runtime ──▶ INT8 auto-cast ──▶ NPX6 NPU
40
+ ```
41
+
42
+ ## Provided Artifacts
43
+
44
+ | Artifact | Status | Notes |
45
+ |----------|---------|---------|
46
+ | **FP32 (ONNX)** | ✅ Provided | Reference model from ONNX Model Zoo |
47
+
48
+ > INT8 execution is handled automatically by the NPU runtime — no additional quantized model file is needed.
49
+
50
+ ## Performance
51
+
52
+ All HIL results were measured on **Renesas R-Car X5H** physical hardware.
53
+ The FP32 ONNX model is auto-cast to INT8 by the runtime before NPU execution.
54
+ PPA Estimator results are software estimates based on model characteristics and hardware configuration.
55
+
56
+ > **Benchmark configuration:** Single NPU · Single AI Core · Input: 3 × 480 × 640 · Batch size: 1
57
+
58
+ ### Inference Latency & Throughput
59
+
60
+ | Runtime | Precision | Device | Latency (ms) | Throughput (fps) | Type |
61
+ |----------|----------|----------|----------|----------|----------|
62
+ | ORT Custom NPU EP | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | TBD | TBD | Measured |
63
+ | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | TBD | TBD | Measured |
64
+ | PPA Estimator | INT8 | X5H · 1× NPU · 1 Core · 1066 MHz | TBD | — | Estimated |
65
+
66
+ ### Accuracy (COCO Validation Set)
67
+
68
+ | Runtime / Precision | mAP (IoU=0.50:0.95) | Notes |
69
+ |----------|----------|----------|
70
+ | FP32 Reference | 0.376 | ONNX Model Zoo reference |
71
+ | ORT Custom NPU EP (INT8) | TBD | NPU execution |
72
+ | MWMX Runtime (INT8) | TBD | NPU execution |
73
+
74
+ ---
75
+
76
+ ## Runtime Details
77
+
78
+ ### ONNX Runtime – Custom NPU Execution Provider
79
+
80
+ - **Engine:** ONNX Runtime with Renesas Custom NPU Execution Provider
81
+ - **Input format:** FP32 ONNX (`.onnx`)
82
+ - **NPU execution precision:** INT8 (auto-cast at load time)
83
+ - **Execution target:** NPX6-48K NPU on R-Car X5H
84
+
85
+ ### MWMX Runtime
86
+
87
+ - **Engine:** Renesas MWMX (Middleware MX) native inference runtime
88
+ - **Input format:** FP32 ONNX (ingested and compiled by the MWMX toolchain)
89
+ - **NPU execution precision:** INT8 (auto-cast by MWMX toolchain)
90
+ - **Execution target:** NPX6-48K NPU on R-Car X5H
91
+
92
+ ### PPA Estimator
93
+
94
+ - **Engine:** Renesas PPA Estimator
95
+ - **Input format:** FP32 ONNX
96
+ - **NPU execution precision:** INT8
97
+ - **Type:** Software performance estimate — not measured on physical silicon
98
+
99
+ ---
100
+
101
+ ## Model Input
102
+
103
+ ### Input Tensor
104
+
105
+ - Shape: `(N, 3, H, W)`
106
+ - Format: RGB
107
+ - Data Type: FP32
108
+ - Pixel Range: `[0, 1]`
109
+
110
+ ### Preprocessing
111
+
112
+ ```python
113
+ from torchvision import transforms
114
+
115
+ preprocess = transforms.Compose([
116
+ transforms.ToTensor(),
117
+ transforms.Normalize(
118
+ mean=[0.485, 0.456, 0.406],
119
+ std=[0.229, 0.224, 0.225]
120
+ ),
121
+ ])
122
+ ```
123
+
124
+ ---
125
+
126
+ ## Model Outputs
127
+
128
+ The model produces **10 output tensors** corresponding to RetinaNet's multi-scale detection heads.
129
+
130
+ ### Classification Heads
131
+
132
+ Five tensors corresponding to object classification on feature pyramid levels P3–P7.
133
+
134
+ Example shapes for an input image of size `1 × 3 × 480 × 640`:
135
+
136
+ ```text
137
+ [1, 720, 60, 80]
138
+ [1, 720, 30, 40]
139
+ [1, 720, 15, 20]
140
+ [1, 720, 8, 10]
141
+ [1, 720, 4, 5]
142
+ ```
143
+
144
+ ### Bounding Box Regression Heads
145
+
146
+ Five tensors corresponding to anchor-box regression outputs.
147
+
148
+ ```text
149
+ [1, 36, 60, 80]
150
+ [1, 36, 30, 40]
151
+ [1, 36, 15, 20]
152
+ [1, 36, 8, 10]
153
+ [1, 36, 4, 5]
154
+ ```
155
+
156
+ ### Postprocessing
157
+
158
+ RetinaNet requires the following postprocessing steps:
159
+
160
+ 1. Anchor generation
161
+ 2. Bounding box decoding
162
+ 3. Confidence threshold filtering
163
+ 4. Non-Maximum Suppression (NMS)
164
+
165
+ These steps produce the final object detections:
166
+
167
+ - Bounding boxes
168
+ - Confidence scores
169
+ - Class labels
170
+
171
+ ---
172
+
173
+ ## Prerequisites
174
+
175
+ To run inference on Renesas R-Car X5H, you need:
176
+
177
+ 1. **Renesas R-Car X5H board** with NPX6 NPU
178
+ 2. **ONNX Runtime** with Renesas NPU Custom Execution Provider, or the **Renesas MWMX Runtime**
179
+ 3. **Hugging Face CLI** to download the model
180
+
181
+ ## Download
182
+
183
+ ```bash
184
+ huggingface-cli download Renesas/RetinaNet-ONNX fp32/retinanet-9.onnx
185
+ ```
186
+
187
+ ## Inference
188
+
189
+ ### ONNX Runtime (Custom NPU Execution Provider)
190
+
191
+ ```python
192
+ import onnxruntime as ort
193
+ import numpy as np
194
+
195
+ providers = [
196
+ ("RenesasNPUExecutionProvider", {}),
197
+ "CPUExecutionProvider"
198
+ ]
199
+
200
+ sess = ort.InferenceSession(
201
+ "fp32/retinanet-9.onnx",
202
+ providers=providers
203
+ )
204
+
205
+ input_data = np.random.rand(
206
+ 1, 3, 480, 640
207
+ ).astype(np.float32)
208
+
209
+ outputs = sess.run(
210
+ None,
211
+ {"images": input_data}
212
+ )
213
+
214
+ # outputs[0:5] -> classification heads
215
+ # outputs[5:10] -> box regression heads
216
+ ```
217
+
218
+ ### MWMX Runtime
219
+
220
+ Refer to the Renesas MWMX Runtime documentation for compilation and inference scripts targeting the NPX6 NPU on R-Car X5H. The MWMX toolchain ingests the FP32 ONNX model and automatically compiles it for INT8 NPU execution.
221
+
222
+ ---
223
+
224
+ ## Benchmark Methodology
225
+
226
+ - **HIL runs:** Hardware-in-the-loop — measured on physical R-Car X5H silicon; single NPU, single AI core, 850 MHz NPU clock
227
+ - **Estimation:** PPA Estimator software estimate; single NPU, single AI core, 1066 MHz NPU clock
228
+ - **Precision:** FP32 ONNX input; INT8 execution (auto-cast by runtime)
229
+ - **Latency:** Median over 1000 consecutive inference runs with warm cache
230
+ - **Throughput:** Computed as `1000 / latency_ms`
231
+ - **Accuracy:** Evaluated using the COCO validation dataset
232
+ - **Postprocessing:** Includes anchor generation, bounding-box decoding, confidence filtering, and NMS
fp32/.metadata.yaml ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ variant:
2
+ id: onnx_fp32
3
+ format: onnx
4
+ precision: fp32 # full-precision FP32 reference model
5
+ method: onnx_export # exported from ONNX Model Zoo v1.12
6
+
7
+ quantization:
8
+ datatype: fp32
9
+ scope:
10
+ - weights
11
+ granularity: none # not applicable for FP32
12
+ calibration: none # no calibration for FP32 reference
13
+ toolchain: onnx
fp32/benchmarks/x5h_ort_npu.yaml ADDED
@@ -0,0 +1,40 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ hardware:
2
+ vendor: renesas
3
+ chip: rcar-x5h
4
+ cpu: arm-cortex-a720
5
+ npu: npx6-48k
6
+ npu_count: 2
7
+ npu_cores: 12
8
+ npu_default_freq_mhz: 1066
9
+ accelerator:
10
+ - npu
11
+
12
+ runtime:
13
+ engine: onnxruntime
14
+ format: onnx
15
+ execution_provider: npu
16
+ execution_precision: fp32
17
+
18
+ configuration:
19
+ npu_instances: 1
20
+ npu_cores_per_instance: 1
21
+ npu_freq_mhz: 850
22
+
23
+ benchmark:
24
+ type: hil
25
+ parameters:
26
+ batch_size: 1
27
+ input_resolution: [1, 3, 480, 640]
28
+
29
+ performance:
30
+ fps: null
31
+ latency: null
32
+
33
+ metrics:
34
+ map: null # COCO validation set mAP (IoU=0.50:0.95) FP32 reference baseline
35
+
36
+ memory:
37
+ peak_mb: null
38
+
39
+ power:
40
+ avg_w: null
fp32/retinanet-9.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:06742923960ec4d9899e6fe407d4d2df013fe6962504f099463ca1b8cba45e44
3
+ size 228369343
int8/.metadata.yaml ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ variant:
2
+ id: onnx_int8
3
+ format: onnx
4
+ precision: int8 # INT8 quantized — auto-cast by NPU execution provider at runtime
5
+ method: onnx_export # base model exported from ONNX Model Zoo v1.12; quantization applied by runtime
6
+
7
+ quantization:
8
+ datatype: int8
9
+ scope:
10
+ - weights
11
+ - activations
12
+ granularity: per-tensor # typical for runtime-cast INT8
13
+ calibration: ptq # post-training quantization applied by the NPU EP / MWMX toolchain
14
+ toolchain: onnx
int8/benchmarks/x5h_mwmx_npu.yaml ADDED
@@ -0,0 +1,41 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ hardware:
2
+ vendor: renesas
3
+ chip: rcar-x5h
4
+ cpu: arm-cortex-a720
5
+ npu: npx6-48k
6
+ npu_count: 2
7
+ npu_cores: 12
8
+ npu_default_freq_mhz: 1066
9
+ accelerator:
10
+ - npu
11
+
12
+ runtime:
13
+ engine: mwmx # Renesas MWMX (Middleware MX) native inference runtime
14
+ format: onnx # input model format (FP32 ONNX)
15
+ execution_provider: npu
16
+ execution_precision: int8 # FP32 model is auto-cast to INT8 by the MWMX toolchain
17
+
18
+ configuration:
19
+ npu_instances: 1 # single NPU
20
+ npu_cores_per_instance: 12 # single AI core
21
+ npu_freq_mhz: 850 # 850 MHz NPU clock frequency
22
+
23
+ benchmark:
24
+ type: hil # Hardware-in-the-loop — measured on physical X5H silicon
25
+ parameters:
26
+ batch_size: 1
27
+ input_resolution: [1, 3, 480, 640]
28
+
29
+ performance:
30
+ fps: null # throughput: 1000 / latency
31
+ latency: 7.30 # median inference latency (ms) over 1000 warm runs
32
+
33
+ metrics:
34
+ accuracy: null
35
+ top5_accuracy: null
36
+
37
+ memory:
38
+ peak_mb: null # peak NPU memory during inference
39
+
40
+ power:
41
+ avg_w: null # not yet characterized
int8/benchmarks/x5h_ppa_npu.yaml ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ hardware:
2
+ vendor: renesas
3
+ chip: rcar-x5h
4
+ cpu: arm-cortex-a720
5
+ npu: npx6-48k
6
+ npu_count: 2
7
+ npu_cores: 12
8
+ npu_default_freq_mhz: 1066
9
+ accelerator:
10
+ - npu
11
+
12
+ runtime:
13
+ engine: ppa-estimator
14
+ format: onnx
15
+ execution_provider: npu
16
+ execution_precision: int8
17
+
18
+ configuration:
19
+ npu_instances: 12
20
+ npu_cores_per_instance: 1
21
+ npu_freq_mhz: 850
22
+
23
+ benchmark:
24
+ type: estimation
25
+ source: ppa-estimator
26
+ parameters:
27
+ batch_size: 1
28
+ input_resolution: [1, 3, 480, 640]
29
+
30
+ performance:
31
+ fps: null
32
+ latency: 3.67
33
+
34
+ metrics:
35
+ accuracy: null
36
+ top5_accuracy: null
37
+
38
+ memory:
39
+ peak_mb: null
40
+
41
+ power:
42
+ avg_w: null