library_name: onnx
license: apache-2.0
tags:
- foundation
- amd
- rocm
- object-detection
pipeline_tag: object-detection
PointPillars: Optimized for AMD ROCm
PointPillars is a lightweight 3D object detector that converts LIDAR point clouds into a 2D pseudo-image via pillar-based feature encoding, then runs a standard 2D detection backbone. This repository packages inference for 3D object detection using ONNX Runtime, exported and validated for AMD ROCm so it runs efficiently on AMD CPUs and NPUs.
This repository contains configurations and scripts optimized for AMD® ROCm™ platforms. You can use the pointpillars AMD scripts to reproduce results or export with custom configurations. More details on model performance can be found here.
Task Overview
Task: 3D object detection (LIDAR point clouds)
Dataset: KITTI (3 classes: Car, Pedestrian, Cyclist)
Output metrics: 3D AP, BEV AP, mAP, per-class AP (Car/Pedestrian/Cyclist), per-difficulty AP (Easy/Moderate/Hard)
ONNX Runtime note: GPU is not supported due to memory constraints and MIGraphX compatibility issues. NPU targets exist but are under beta — the VitisAI EP compiles the model, however the ScatterND/Where operations cause hardware execution timeouts at runtime.
AMD ROCm Optimization
This model export has been adapted and validated for AMD CPUs, with a beta AMD Ryzen AI NPU path. Key points:
- Validated backend: ONNX Runtime on CPU (FP32); NPU (VitisAI execution provider) compiles but may hit runtime hardware timeouts.
- No code changes required versus the upstream PointPillars implementation — only environment/runtime configuration differs.
- GPU is not supported due to memory constraints and MIGraphX compatibility issues.
| Runtime | Precision | Backend | Hardware | Notes |
|---|---|---|---|---|
| ONNX Runtime | FP32 | CPU Execution Provider | AMD CPU | — |
| ONNX Runtime | Auto | VitisAI Execution Provider | AMD Ryzen AI NPU | Beta — compiles but may hit runtime timeouts (ScatterND/Where ops) |
Getting Started
For setup instructions, evaluation scripts, and custom configuration options, see the pointpillars on GitHub.
Model Details
Model Type: 3D object detection (pillar-based point cloud encoder + 2D detection backbone)
Model Stats:
- Input (voxel_feats):
(12000, 64, 10)float32; (voxel_coords):(12000, 3)int64; (num_points):(12000,)int64 - Output (cls_preds):
(1, 53568, 18)float32; (box_preds):(1, 53568, 42)float32 - Precision tested: FP32 (CPU); auto-quantized (NPU, beta)
Accuracy Pipeline
Higher AP means the model's predicted 3D bounding boxes and classes agree more closely with ground truth — 100% would be perfect detection, 0% means no correct detections. BEV AP evaluates localization in the bird's-eye view plane only; 3D AP additionally requires correct height estimation. In practice, the published PointPillars paper reports ~77% BEV AP (Car, Moderate) on the full KITTI val split.
Metrics Explained
| Metric | Description |
|---|---|
| 3D AP | Average Precision for 3D bounding box detection at the class-specific IoU threshold. The strictest metric — predicted boxes must overlap ground truth in all three dimensions (x, y, z, plus size and orientation). Higher means better 3D localization. |
| BEV AP | Average Precision in bird's-eye view at the class-specific IoU threshold. Only evaluates the XY plane overlap, ignoring height — typically higher than 3D AP and reflects "did it find and laterally localize the object" rather than full 3D accuracy. |
| mAP | Mean AP averaged across Car, Pedestrian, and Cyclist. A single number summarizing overall detection quality; higher is better. |
| Per-class AP (Car/Pedestrian/Cyclist) | AP broken down per class — exposes class-specific weaknesses. Cars are the easiest (large, frequent); Cyclists are the hardest (rare, small). |
| Difficulty levels (Easy/Moderate/Hard) | AP at three KITTI difficulty levels, defined by 2D bounding box height, occlusion, and truncation. Easy includes only fully visible, large objects; Hard includes heavily occluded and small objects. |
Accuracy Results
KITTI Accuracy Results (CPU, 10 samples) — official AP using the R40 (40-point interpolation) protocol at standard IoU thresholds (0.7 for Car, 0.5 for Pedestrian/Cyclist):
| Metric | Difficulty | Car (IoU=0.7) | Pedestrian (IoU=0.5) | Cyclist (IoU=0.5) | mAP |
|---|---|---|---|---|---|
| 3D AP | Easy | 52.83 | 25.61 | 0.00 | 26.15 |
| 3D AP | Moderate | 57.35 | 25.61 | 100.00 | 60.99 |
| 3D AP | Hard | 57.35 | 25.61 | 100.00 | 60.99 |
| BEV AP | Easy | 61.87 | 25.61 | 0.00 | 29.16 |
| BEV AP | Moderate | 73.52 | 25.61 | 100.00 | 66.38 |
| BEV AP | Hard | 73.52 | 25.61 | 100.00 | 66.38 |
Note: These results are from a 10-sample KITTI subset. Full KITTI val (3,769 samples) results would be more representative but require downloading the complete dataset. Cyclist AP is unstable across repeated runs (observed 0%/50%/100% on identical CPU re-runs) because this subset contains only 1 ground-truth cyclist — a single detection's score threshold swings AP by its full range. Car and Pedestrian AP are stable. This is a sample-size artifact, not a determinism bug; it resolves on the full 3,769-sample val split.
Dig Deeper
Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?
📂 View the full project on GitHub
The GitHub repository includes:
- KITTI dataset staging and BEV visualization pipeline
- Benchmark and profile scripts for CPU and NPU
- Full official KITTI AP evaluation protocol (R40, per-class, per-difficulty)
- Benchmarking and reproduction instructions
