YOLO-Master-EsMoE
This repository provides the Axera NPU deployment of YOLO-Master-EsMoE, a COCO object detection model built on a customized Ultralytics pipeline. The model replaces part of the YOLO backbone/head with mixture-of-experts (MoE) modules (ES_MOE). During ONNX export the MoE routing is converted to a dense fallback (all experts participate), producing a static computation graph that can be compiled by Pulsar2 into axmodel for Axera’s NPU-based AX650 Series. Three model sizes are provided: YOLO-Master-EsMoE-S, YOLO-Master-EsMoE-M and YOLO-Master-EsMoE-N, with an input size of 640×640 and 80 COCO classes.
References links:
For those who are interested in model conversion, you can try to export axmodel through
Support Platform
Performance
| Model | Input Shape | Latency (ms) | CMM Usage (MB) |
|---|---|---|---|
| YOLO-Master-EsMoE-M.axmodel | 1 x 640 x 640 x 3 | 23.446 | 70.40 |
| YOLO-Master-EsMoE-S.axmodel | 1 x 640 x 640 x 3 | 9.453 | 54.81 |
| YOLO-Master-EsMoE-N.axmodel | 1 x 640 x 640 x 3 | 4.406 | 46.68 |
Models
Download all files from this repository to the device
root@ax650 ~/root/yolo-master-esmoe # tree -L 3
.
|-- README.md
|-- requirement.txt
|-- models
| `-- ax650
| |-- config.json
| |-- YOLO-Master-EsMoE-M
| | |-- 1_host_post_YOLO-Master-EsMoE-M.onnx
| | `-- YOLO-Master-EsMoE-M.axmodel
| |-- YOLO-Master-EsMoE-N
| | |-- 1_host_post_YOLO-Master-EsMoE-N.onnx
| | `-- YOLO-Master-EsMoE-N.axmodel
| `-- YOLO-Master-EsMoE-S
| |-- 1_host_post_YOLO-Master-EsMoE-S.onnx
| `-- YOLO-Master-EsMoE-S.axmodel
|-- web
| |-- app.py
| |-- detector.py
| `-- README.md
|-- infer
| |-- infer_axmodel.py
| `-- infer_onnx.py
`-- tools
|-- convert_opset.py
|-- eval_map.py
|-- export_onnx.py
`-- split_onnx.py
python env requirement
pip install -r requirement.txt
Inference with AX650 Host
root@ax650 ~/root/yolo-master-esmoe # python web/app.py
* Running on local URL: http://0.0.0.0:7860
* To create a public link, set `share=True` in `launch()`.
Use the device IP address and port 7860 to access the WebApp, for example http://192.168.1.100:7860. In the page, select the device ax650 and the model YOLO-Master-EsMoE-M, upload an image and click Detect.
Input image
Result:
Inference with YOLO-Master-EsMoE-S and YOLO-Master-EsMoE-N
The demo for YOLO-Master-EsMoE-S and YOLO-Master-EsMoE-N is the same as YOLO-Master-EsMoE-M, just select the corresponding model in the WebApp.
Inference scripts
Besides the Gradio WebApp, two inference scripts are provided.
Run the FP32 ONNX baseline on the development machine:
python infer/infer_onnx.py --model <model.onnx> --image <image.jpg>
Run the axmodel inference on the target board:
python infer/infer_axmodel.py --axmodel models/ax650/YOLO-Master-EsMoE-M/YOLO-Master-EsMoE-M.axmodel --post-onnx models/ax650/YOLO-Master-EsMoE-M/1_host_post_YOLO-Master-EsMoE-M.onnx --image football.jpg
