Object Detection
PyTorch
ONNX
TensorRT
3d-object-detection
autonomous-driving
streampetr
carla
nuscenes
bev
Instructions to use shin0412/StreamPETR_nuCarla with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- TensorRT
How to use shin0412/StreamPETR_nuCarla with TensorRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
| license: apache-2.0 | |
| tags: | |
| - object-detection | |
| - 3d-object-detection | |
| - autonomous-driving | |
| - streampetr | |
| - carla | |
| - nuscenes | |
| - bev | |
| - tensorrt | |
| - onnx | |
| datasets: | |
| - zhijieq/nuCarla | |
| library_name: pytorch | |
| pipeline_tag: object-detection | |
| # StreamPETR — nuCarla (Town04) | |
| Camera-only 3D object detection for autonomous driving. Six surround-view | |
| cameras in, 3D bounding boxes out, with a temporal memory carried across | |
| frames. Trained on the [nuCarla](https://huggingface.co/datasets/zhijieq/nuCarla) | |
| dataset, Town04 subset. | |
| Ships as a PyTorch checkpoint plus ready-to-run ONNX graphs and TensorRT FP16 | |
| engines. | |
| | | | | |
| |---|---| | |
| | Architecture | [StreamPETR](https://github.com/exiawsh/StreamPETR), ResNet-50 + CPFPN | | |
| | Classes | `car, truck, bus, motorcycle, bicycle, pedestrian` | | |
| | Input | 6 cameras, 1600×900 → 704×256 | | |
| | mAP @0.25 3D IoU | **0.5120** (nuCarla Town04 val) | | |
| | Latency | 4.5 ms/frame, RTX 5070 Ti FP16 (~220 fps) | | |
| | Init | nuScenes-pretrained, class head sliced to 6 classes | | |
| ## Demo | |
| `demo/carla_live_detection.mp4` — the model run over a clip recorded live in | |
| CARLA 0.9.16: an ego vehicle driving on autopilot through 150 spawned vehicles | |
| and 16 pedestrians in Town04. Six camera views with projected 3D boxes plus a | |
| bird's-eye panel. Averages 6.7 detections per frame. | |
| The clip is unlabelled, so it is a qualitative check, not a measured score. | |
| ## Checkpoints | |
| | File | Size | mAP @0.25 | What it is | | |
| |---|---|---|---| | |
| | `checkpoints/best_mAP_0.25_iter_8512.pth` | 432M | **0.5120** | Best epoch (28). Everything below — ONNX, engines, the demo video, all reported numbers — comes from this one. | | |
| | `checkpoints/final_iter_13376_epoch44.pth` | 432M | 0.4887 | Last checkpoint of the run. mAP had plateaued; kept for continuing training. | | |
| | `checkpoints/init_nuscenes_6class_sliced.pth` | 148M | — | Starting point: the nuScenes StreamPETR weights with the class head sliced from 10 to 6. Use this to reproduce training. | | |
| Intermediate epoch checkpoints are not published — they are rotation artifacts | |
| from `max_keep_ckpts`, not separately evaluated models. | |
| ## Files | |
| ``` | |
| checkpoints/ the three above | |
| stream_petr_r50_nucarla_town04.py training config | |
| DEPLOYMENT.md full inference guide — read this | |
| demo/carla_live_detection.mp4 12M detection demo | |
| deployment/ | |
| stream_petr_6cam_encoder.onnx 95M | |
| stream_petr_6cam_temporal_head.onnx 47M | |
| *_fp16.engine TensorRT 10.9, RTX 5070 Ti | |
| streampetr_trt_runner.py runtime, handles temporal memory | |
| carla_stream_petr_postprocess.py decoder | |
| build_streampetr_engines.py rebuild engines for another GPU | |
| test_inputs/, reference_outputs/ one frame, to verify your setup | |
| ``` | |
| ## Quick start (TensorRT) | |
| ```bash | |
| pip install tensorrt==10.9.0.34 numpy torch | |
| ``` | |
| ```python | |
| import numpy as np | |
| from streampetr_trt_runner import StreamPETRRunner | |
| from carla_stream_petr_postprocess import decode_stream_petr | |
| runner = StreamPETRRunner( | |
| "deployment/stream_petr_6cam_encoder_fp16.engine", | |
| "deployment/stream_petr_6cam_temporal_head_fp16.engine", | |
| ) | |
| runner.reset() # call at the start of every sequence | |
| for frame in sequence: | |
| out = runner( | |
| images=frame.images, # [1, 6, 3, 256, 704] float32, normalized | |
| timestamp=frame.seconds, | |
| ego_pose=frame.lidar2global, # [1, 4, 4] | |
| ego_pose_inv=frame.global2lidar, # [1, 4, 4] | |
| ) | |
| det = decode_stream_petr( | |
| out["class_logits"], out["bbox_predictions"], score_threshold=0.35 | |
| ) | |
| ``` | |
| The engines are tied to the GPU and TensorRT version they were built on. | |
| Rebuild from the ONNX for anything else — it takes about 15 seconds: | |
| ```bash | |
| python deployment/build_streampetr_engines.py --deployment-dir deployment --cameras 6 | |
| ``` | |
| ## Camera rig — must match | |
| The 3D position embedding is **baked into the exported graph**. Images from a | |
| different camera layout produce meaningless geometry with no error raised. | |
| Offsets are ego-frame metres (x forward, y left, z up), yaw counter-clockwise: | |
| | Camera | x | y | z | yaw | | |
| |---|---|---|---|---| | |
| | `CAM_FRONT` | 1.901 | 0.016 | 1.511 | 0.3° | | |
| | `CAM_FRONT_LEFT` | 1.524 | 0.495 | 1.509 | 55.2° | | |
| | `CAM_FRONT_RIGHT` | 1.551 | −0.493 | 1.496 | −56.4° | | |
| | `CAM_BACK` | 0.028 | 0.003 | 1.579 | 179.9° | | |
| | `CAM_BACK_LEFT` | 1.036 | 0.485 | 1.591 | 108.6° | | |
| | `CAM_BACK_RIGHT` | 1.015 | −0.481 | 1.562 | −110.8° | | |
| All six share 1600×900 at 65° horizontal FOV (`fx = fy = 1255.7484`, | |
| `cx = 800`, `cy = 450`) — what CARLA emits for `image_size_x=1600, | |
| image_size_y=900, fov=65`. | |
| Preprocessing, tensor shapes and the temporal memory protocol are all in | |
| `DEPLOYMENT.md`. | |
| ## Results | |
| nuCarla Town04 validation, 29 scenes / 1160 frames: | |
| | Class | AP @0.25 IoU | | |
| |---|---| | |
| | car | 0.630 | | |
| | bus | 0.623 | | |
| | truck | 0.612 | | |
| | bicycle | 0.358 | | |
| | motorcycle | 0.317 | | |
| | pedestrian | 0.018 | | |
| | **mAP** | **0.5120** | | |
| Pedestrian scores near zero **under 3D IoU** because the metric is brutal on | |
| small boxes: a 0.42 m wide pedestrian tolerates ~0.25 m of lateral error to | |
| reach IoU 0.25, against ~1.2 m for a car. Under the nuScenes center-distance | |
| metric the paper actually reports, pedestrian AP is **0.309** and overall mAP | |
| **0.4638** — the model does detect pedestrians. | |
| ### Not comparable to the nuCarla paper's 0.745 | |
| The paper reports mAP 0.745 for PETR, but under different conditions: | |
| | | Paper (PETR) | This model | | |
| |---|---|---| | |
| | Training data | 700 scenes / 7 maps | 91 scenes / **1 map** | | |
| | Backbone | VoVNet | ResNet-50 | | |
| | Resolution | 1600×640 | 704×256 (**5.7× fewer pixels**) | | |
| | GPU-hours | 150 | ~5 | | |
| | Metric | center-distance | 3D IoU | | |
| Town04 is also CARLA's highway map, where objects sit far from the ego (cars | |
| average 50 m away), making it harder than the urban maps in the full set. | |
| ## Training | |
| - Initialized from the official nuScenes StreamPETR checkpoint. All six target | |
| classes exist in nuScenes' ten, so the classification rows were **sliced | |
| across** rather than randomly re-initialized — nothing in the network starts | |
| from scratch. This lifted first-epoch mAP from 0.0497 to 0.2701. | |
| - 50 epochs planned, batch 12, AdamW at 1.2e-4 with cosine annealing, FP16. | |
| The run reached epoch 45; mAP had plateaued since epoch 23 and the LR was | |
| annealed to 3.6e-6. | |
| - Best checkpoint at iteration 8512 (epoch 28). | |
| ## Verification | |
| The ONNX wrapper reproduces the native PyTorch head exactly (`max |Δ| = 0.0` | |
| on both outputs). TensorRT FP16 against PyTorch FP32 on the bundled frame: | |
| | tensor | correlation | rel. error | | |
| |---|---|---| | |
| | `image_features` | 0.999996 | 2.9e-3 | | |
| | `bbox_predictions` | 0.999992 | 2.8e-2 | | |
| | `class_logits` | 0.999546 | 7.6e-2 | | |
| Decoded detections agree: **17 vs 17** above score 0.2, every one placed within | |
| 0.02–0.10 m. Two pairs swap rank order, both between detections whose scores | |
| tie to within 0.003. | |
| **Known FP16 effect.** 768 of the 1024 memory slots are bit-exact frame to | |
| frame; within the 256 slots selected fresh each frame, FP16 scoring picks a | |
| slightly different subset near the top-256 cut-off. Its accumulation over long | |
| sequences has not been measured — build the FP32 engine if you need to rule it | |
| out. | |
| ## Limitations | |
| - Trained on **one map** (Town04). Expect degradation on urban maps or dense | |
| intersections. | |
| - Pedestrian localization is weak in absolute terms; suitable for perception | |
| research, not safety-critical use. | |
| - Simulation only — no real-world images were used, and no sim-to-real transfer | |
| has been evaluated. | |
| - The rig is fixed. Any change to camera placement, count or intrinsics | |
| requires re-exporting from the checkpoint. | |
| ## Citation | |
| ```bibtex | |
| @article{nucarla, | |
| title={nuCarla: A nuScenes-Style Bird's-Eye View Perception Dataset for CARLA Simulation}, | |
| author={Qiao, Zhijie and Cao, Zhong and Liu, Henry X.}, | |
| year={2025}, | |
| url={https://arxiv.org/abs/2511.13744} | |
| } | |
| @article{streampetr, | |
| title={Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object Detection}, | |
| author={Wang, Shihao and Liu, Yingfei and Wang, Tiancai and Li, Ying and Zhang, Xiangyu}, | |
| journal={arXiv preprint arXiv:2303.11926}, | |
| year={2023} | |
| } | |
| ``` | |
| ## License | |
| Apache 2.0, following [StreamPETR](https://github.com/exiawsh/StreamPETR) and | |
| the nuCarla dataset. | |