PointPillars: Optimized for AMD ROCm

PointPillars is a lightweight 3D object detector that converts LIDAR point clouds into a 2D pseudo-image via pillar-based feature encoding, then runs a standard 2D detection backbone. This repository packages inference for 3D object detection using ONNX Runtime, exported and validated for AMD ROCm so it runs efficiently on AMD CPUs and NPUs.

This repository contains configurations and scripts optimized for AMDยฎ ROCmโ„ข platforms. You can use the pointpillars AMD scripts to reproduce results or export with custom configurations. More details on model performance can be found here.


Task Overview

Task: 3D object detection (LIDAR point clouds)

Dataset: KITTI (3 classes: Car, Pedestrian, Cyclist)

Output metrics: 3D AP, BEV AP, mAP, per-class AP (Car/Pedestrian/Cyclist), per-difficulty AP (Easy/Moderate/Hard)

ONNX Runtime note: GPU is not supported due to memory constraints and MIGraphX compatibility issues. NPU targets exist but are under beta โ€” the VitisAI EP compiles the model, however the ScatterND/Where operations cause hardware execution timeouts at runtime.


AMD ROCm Optimization

This model export has been adapted and validated for AMD CPUs, with a beta AMD Ryzen AI NPU path. Key points:

  • Validated backend: ONNX Runtime on CPU (FP32); NPU (VitisAI execution provider) compiles but may hit runtime hardware timeouts.
  • No code changes required versus the upstream PointPillars implementation โ€” only environment/runtime configuration differs.
  • GPU is not supported due to memory constraints and MIGraphX compatibility issues.
Runtime Precision Backend Hardware Notes
ONNX Runtime FP32 CPU Execution Provider AMD CPU โ€”
ONNX Runtime Auto VitisAI Execution Provider AMD Ryzen AI NPU Beta โ€” compiles but may hit runtime timeouts (ScatterND/Where ops)

Getting Started

For setup instructions, evaluation scripts, and custom configuration options, see the pointpillars on GitHub.


Model Details

Model Type: 3D object detection (pillar-based point cloud encoder + 2D detection backbone)

Model Stats:

  • Input (voxel_feats): (12000, 64, 10) float32; (voxel_coords): (12000, 3) int64; (num_points): (12000,) int64
  • Output (cls_preds): (1, 53568, 18) float32; (box_preds): (1, 53568, 42) float32
  • Precision tested: FP32 (CPU); auto-quantized (NPU, beta)

Accuracy Pipeline

Higher AP means the model's predicted 3D bounding boxes and classes agree more closely with ground truth โ€” 100% would be perfect detection, 0% means no correct detections. BEV AP evaluates localization in the bird's-eye view plane only; 3D AP additionally requires correct height estimation. In practice, the published PointPillars paper reports ~77% BEV AP (Car, Moderate) on the full KITTI val split.

Metrics Explained

Metric Description
3D AP Average Precision for 3D bounding box detection at the class-specific IoU threshold. The strictest metric โ€” predicted boxes must overlap ground truth in all three dimensions (x, y, z, plus size and orientation). Higher means better 3D localization.
BEV AP Average Precision in bird's-eye view at the class-specific IoU threshold. Only evaluates the XY plane overlap, ignoring height โ€” typically higher than 3D AP and reflects "did it find and laterally localize the object" rather than full 3D accuracy.
mAP Mean AP averaged across Car, Pedestrian, and Cyclist. A single number summarizing overall detection quality; higher is better.
Per-class AP (Car/Pedestrian/Cyclist) AP broken down per class โ€” exposes class-specific weaknesses. Cars are the easiest (large, frequent); Cyclists are the hardest (rare, small).
Difficulty levels (Easy/Moderate/Hard) AP at three KITTI difficulty levels, defined by 2D bounding box height, occlusion, and truncation. Easy includes only fully visible, large objects; Hard includes heavily occluded and small objects.

Accuracy Results

KITTI Accuracy Results (CPU, 10 samples) โ€” official AP using the R40 (40-point interpolation) protocol at standard IoU thresholds (0.7 for Car, 0.5 for Pedestrian/Cyclist):

Metric Difficulty Car (IoU=0.7) Pedestrian (IoU=0.5) Cyclist (IoU=0.5) mAP
3D AP Easy 52.83 25.61 0.00 26.15
3D AP Moderate 57.35 25.61 100.00 60.99
3D AP Hard 57.35 25.61 100.00 60.99
BEV AP Easy 61.87 25.61 0.00 29.16
BEV AP Moderate 73.52 25.61 100.00 66.38
BEV AP Hard 73.52 25.61 100.00 66.38

Note: These results are from a 10-sample KITTI subset. Full KITTI val (3,769 samples) results would be more representative but require downloading the complete dataset. Cyclist AP is unstable across repeated runs (observed 0%/50%/100% on identical CPU re-runs) because this subset contains only 1 ground-truth cyclist โ€” a single detection's score threshold swings AP by its full range. Car and Pedestrian AP are stable. This is a sample-size artifact, not a determinism bug; it resolves on the full 3,769-sample val split.


Dig Deeper

Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?

๐Ÿ“‚ View the full project on GitHub

The GitHub repository includes:

  • KITTI dataset staging and BEV visualization pipeline
  • Benchmark and profile scripts for CPU and NPU
  • Full official KITTI AP evaluation protocol (R40, per-class, per-difficulty)
  • Benchmarking and reproduction instructions
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support