bevfusion-p150
BEVFusion-L (Autoware autoware_bevfusion, lidar-only): the LiDAR-only BEVFusion network that Autoware's autoware_bevfusion node runs by default (bevfusion_lidar.onnx of release v2.0), ported to one Tenstorrent Blackhole p150 with tt-nn. One LiDAR point cloud in base_link in (x, y, z; one sweep, or Autoware's 2-sweep densified cloud with a time_lag column; intensity is not used); Autoware DetectedObjects-equivalent 3D boxes of CAR, TRUCK, BUS, TRAILER, BICYCLE and PEDESTRIAN out, with the node's pre- and post-processing. In Autoware it is an opt-in alternative: CenterPoint is the default LiDAR detector. The package's camera-lidar profile is not in this release yet.
Weights: AutowareFoundation/bevfusion v2.0 · Paper: arXiv:2205.13542 · Autoware package: autoware_bevfusion · Training code: tier4/AWML (projects/BEVFusion; the lidar-only release behind v2.0 is not published) · Port: code/
Runs on p150 (mesh P150). Configuration: dispatch on the ETH cores, 1 command queue, 12×10 compute grid. Precision: fp32 weights in the sparse 3D encoder (bf16 gather tables, a two-term VFE table, fp32 bias and residual adds), two-term (hi + lo bf16) weights and activations in the SECOND backbone with fp32 sums, fp32 weights and bf16 activations in the neck, head and decoder, HiFi4 math with fp32 accumulation; no bfp8. All numbers on this card were measured in this configuration.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
Quickstart (Python)
Prerequisite: a tt-metal / ttnn environment at tt-metal 44d66500520 with patches/tt-metal-eth-dispatch.patch applied. ttnn is not on PyPI.
hf download changh95/bevfusion-p150 --exclude "image/*" --local-dir bevfusion-p150 && cd bevfusion-p150
pip install -e . # adds numpy<2, pillow, pyyaml, onnx, huggingface_hub, safetensors; ttnn and torch come from tt-metal
pip install -e ".[server,test]" # optional: the HTTP server and the tests
Run the snippet from the model repo root: code/tt_bevfusion/samples/test_pcd.npz (Autoware's test.pcd) is a path relative to it.
from tt_bevfusion import BEVFusion
with BEVFusion.from_pretrained(device_id=0) as model: # weights -> your HF cache, traces captured
out = model("code/tt_bevfusion/samples/test_pcd.npz") # path, (N, C) array, .npy/.npz/.pcd, or a PointCloud
for d in out.to_dicts()[:5]:
print(d["label"], d["score"], d["center"], d["size"], d["yaw"])
from_pretraineddownloadsAutowareFoundation/bevfusionat the pinned commite1bf164b909(tagv2.0; only the lidar-only files, 34 MB) to your HF cache, opens the chip, builds the graph and captures four metal traces (the sparse encoder for each of three capacity buckets, and the dense network). The first load compiles the kernels (195.7 s with an empty JIT cache, measured on a loaded host); later loads take about 9 s.- The traces are captured during the load, so no call compiles anything: first / second call on the shipped sample 678.7 / 659.2 ms and 792.2 / 674.9 ms in two processes on a shared host (load average 8.5-9; most of a call is host pre-processing).
- The
withblock releases the traces and closes the chip. Withoutwith, callmodel.close().
| Input | A point cloud in base_link: path (.npy / .npz / .pcd / .bin), raw bytes with fmt=, an (N, C) float array or tensor, or a PointCloud; fields x, y, z (an intensity column is accepted and ignored: use_intensity: false). A column named time_lag (s) marks an already densified cloud. Optional sweeps= (client-side densification) or stream= (server-side, Autoware's own densification with poses). |
| Options | score_threshold=0.1, circle_nms_dist_threshold=0.5, iou_nms_threshold=0.1, iou_nms_search_distance_2d=10.0, remap_classes=True, max_detections=None. from_pretrained(device_id=0, dispatch="eth", weights_dir=None, device=None). |
| Output | Detections3D: boxes float32 [N, 7] (x, y, z, length, width, height, yaw), scores [N] (Autoware's existence_probability), label_ids [N] (ObjectClassification values), labels, meta (voxel and per-level sparse counts, the capacity bucket, overflow flags, per-stage counts, labels before the remapper), timing_ms. Sorted by score. |
| Methods | out.to_dict() gives the /predict JSON. out.to_dicts() gives the detection list. tt_bevfusion.viz.render_bev(points, out.to_dicts()) draws a bird's-eye view. |
- The API gives the same output as the HTTP server
/predict: both share the decoders, the device traces and the host post-processing (checked on the device bytest_api_equals_server). - Every option is a host-side knob with Autoware's value as the default; none changes a device shape. A bad value is refused before the device runs.
- One model uses one chip; calls from several threads are serialised.
- Full reference:
code/PYTHON.md. Runnable example:examples/quickstart.py(also writesquickstart_bev.png, the points and the boxes from above).
Serving (HTTP)
tt-model pull changh95/bevfusion-p150 --with-weights
tt-model serve changh95/bevfusion-p150 # or with tt-cli: tt serve changh95/bevfusion-p150
python3 code/tt_bevfusion/server/client.py --points code/tt_bevfusion/samples/test_pcd.npz --out req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/bevfusion-p150
- The image does not contain the weights.
--with-weightsputs them in your HF cache. - The server uses port 20000 (or the next free port). It is ready when the log shows
Application startup complete. client.pybuilds the request with the standard library only; add--url http://127.0.0.1:20000to send it.POST /predict:points(base64.npy/.npz/.pcd/ rawbin, base_link, fields x, y, z (intensity is accepted and ignored) or a pre-densified cloud with atime_lagcolumn); optionalsweeps(past sweeps withtime_lag_sandT_current_from_sweep) orstream(id,timestamp_s,T_world_from_ego: server-side 2-sweep densification as Autoware),params(every option above),output_format. AlsoGET /health,GET /info,GET /v1/models(stub). Contract:SERVING.mdsection 3.
{
"model": "bevfusion-p150",
"frame_id": "base_link",
"num_detections": 5,
"detections": [
{"label": "TRUCK", "label_id": 2, "score": 0.7584, "center": [24.184, 11.192, 1.535], "size": [8.107, 2.519, 2.849], "yaw": 3.1413},
{"label": "TRUCK", "label_id": 2, "score": 0.5795, "center": [20.807, -0.494, 1.61], "size": [6.096, 2.339, 3.109], "yaw": -0.0158},
...
],
"meta": {"variant": "lidar", "num_voxels": 21359, "voxel_overflow": false, "sparse": {"counts": {"l1": 21359, "l2": 27551, "l3": 15823, "l4": 7037, "out": 5393}, "bucket": "s", "overflow": false}, "counts": {"proposals": 500, "after_threshold": 5, "after_circle_nms": 5, "after_iou_nms": 5}},
"timing_ms": {"preprocess": 589.116, "device": 84.441, "postprocess": 2.36, "total": 695.67, "decode": 19.397, "model_call": 675.96}
}
center/size([length, width, height]) /yaware inbase_linkmetres / radians, Autoware's convention for this node (z is the box centre, yaw counter-clockwise from +x);label_idis theautoware_perception_msgsObjectClassificationvalue; detections are sorted by score.- No velocity is returned (Autoware decodes the network's velocity rows but does not publish them). TRAILER and remapped TRUCK labels come from Autoware's area-based class remapper (
remap_classes=falseturns it off;meta.label_before_remapkeeps the network's class).
Demo
The p150 output on the shipped sample (Autoware's test.pcd, one sweep, Apache-2.0): the p150 boxes in colour over the fp32 CPU reference's boxes (light outlines); a box without a partner would be ringed and tagged.
On public driving datasets (p150 outputs; the frames themselves are not in this repository; dataset boxes in thin grey; the model sees only the LiDAR, the camera images are for orientation):
Agreement with the fp32 CPU reference on these frames (same label, centres < 0.2 m): PandaSet 019 f40 128 of 128 reference detections; PandaSet 090 f40 127 of 129, 130 on the p150 (two low-score cars 0.4 m apart where circle NMS kept a different proposal, and one extra at score 0.101); nuScenes 0103 key-frame 20 29 of 29; the 20 frames of the animation 2,555 of 2,569.
PandaSet renders: contains data from PandaSet (Scale AI and Hesai), https://pandaset.org, licensed under CC BY 4.0 and the PandaSet Dataset Terms; converted to bird's-eye views and resized camera images (licence plates and heads of nearby road users blurred) with drawn boxes; Scale AI and Hesai do not endorse this work. The nuScenes render is non-commercial (CC BY-NC-SA 4.0): rendered from the nuScenes dataset, © Motional AD Inc., nuScenes Terms of Use; Motional does not endorse this work. Sources, changes and the full attributions: media/ATTRIBUTION.md.
Demo & Performances
Warm, batch 1, 2026-10-09. Latency: the stage bench of OPT_BASELINE.md (code/scripts/bench.py, 100 / 50 iterations) on the shipped sample test.pcd (48,048 points, one sweep, 21,359 voxels, sparse capacity bucket s) and on a PandaSet frame (213,007 points, two sweeps, 95,109 voxels, bucket m; not shipped), re-checked on the release code (30 / 10 iterations: back-to-back 53.09 / 112.01 ms, model() 658.2 / 2,781.7 ms); the served rows from uvicorn on the host (the app the container runs) and a loopback client, 30 requests of the shipped sample. The host is shared with other jobs, so the host stages move with its load; the device rows repeat to 0.01 ms. Accuracy: the p150 output against the fp32 CPU reference of the same network (same weights, same pre- and post-processing, identical voxel and rulebook input), matched by label and BEV centre distance (< 0.2 m).
| Metric | Performance |
|---|---|
Agreement with the fp32 CPU reference, shipped sample test.pcd |
5 / 5 reference detections matched (same label, centres < 0.2 m; max |Δscore| 0.0043) |
| Agreement, PandaSet 019 f40 (the frozen sample gate; not shipped) | 128 / 128 (max |Δscore| 0.0081) |
| Agreement, 7 development frames (test.pcd, PandaSet ×2, nuScenes mini_val ×3, Autoware's sample rosbag) | 355 / 359 reference detections matched (recall 0.989, precision 0.986); max |Δscore| 0.0086 |
| Agreement, 10 held-out frames never used for any precision choice (PandaSet ×4, nuScenes ×6) | 680 / 680 (max |Δscore| 0.014) |
| Agreement, PandaSet 019 and 090 frames 30-49 (40 frames; 34 of them in no gate set) | 5,174 / 5,204 (recall 0.9942, precision 0.9952); 11 of 5,174 matched pairs differ in score by more than 0.02 (max 0.068) |
| Module PCC vs the fp32 reference (replay outputs; worst of PandaSet 019 f40, nuScenes 0103 k20, rosbag #144) | sparse BEV ≥ 0.999989 · SECOND blocks ≥ 0.999985 · heat ≥ 0.999981 · top-500 overlap ≥ 0.994 · worst teacher-forced head 0.999461 |
Python model() call, shipped sample test.pcd (host pre-processing, H2D, traces, D2H, host post-processing) |
657.1 ms p50 (1.52 calls/s) |
Python model() call, PandaSet frame (213,007 points, 2 sweeps; not shipped) |
2,786.0 ms p50 (0.36 calls/s) |
| … of which the host sparse rulebook build (numpy) | 580.8 ms (test.pcd) · 2,388.4 ms (PandaSet): 86-88 % of the call |
Served /predict timing_ms.total (uvicorn on the host, the shipped sample) |
697.1 ms median |
| Served client round trip, loopback (709 kB request) | 705.1 ms median |
| Device traces, one blocking frame (sparse + dense) | 53.23 ms (21.7 + 31.5) · PandaSet 112.15 ms (80.7 + 31.5) |
| Back-to-back frames, device time per frame | 53.09 ms (18.8 frames/s) · PandaSet 112.02 ms (8.9 frames/s) |
Input packing + H2D · D2H · host post-processing (test.pcd) |
18.7 + 6.6 · 0.33 · 2.5 ms (PandaSet: 85.7 + 20.5 · 0.32 · 23.6 ms) |
from_pretrained load: empty JIT cache (loaded host) / warm cache |
195.7 s / 9.3 s (capture of the four traces 4.0 s) |
All numbers in this table were measured with dispatch on the ETH cores, 1 command queue and a 12×10 compute grid on one p150, with the published precision. No paper-style metric (mAP / NDS) is claimed: these weights were trained on TIER IV's internal data only, so a score on nuScenes or PandaSet would measure the domain gap, not the port; the accuracy rows are agreement with the CPU reference. Details: VERIFICATION_2026-10-09.md, OPT_BASELINE.md, OPT_REPORT.md.
No GPU comparison: no GPU was available on the host where this port was built and measured, so this card makes no GPU speed claim. The reference rows are the port's own fp32 CPU reference on the same host (a correctness baseline, not a speed target). Autoware runs this network as TensorRT engines with spconv plugins and GPU voxelization; this port was validated against the fp32 reference, not against a TensorRT engine. p150 power per inference was not measured (the stage bench only logged the board's telemetry, median 66-70 W over whole runs including the host stages), so no efficiency comparison is made.
Caveats
- First release: baseline port, optimization pending. End to end is host-bound: the numpy sparse rulebook build (C13: four sub-manifold and four strided neighbour maps) takes 581 ms of the 657 ms call on
test.pcdand 2,388 ms of 2,786 ms on a dense PandaSet frame. It is the first optimization target (the container has no C compiler, so: vectorised lookups or a device rulebook builder;OPT_REPORT.md). On the device the sparse encoder's im2col gathers (40.9 ms of the PandaSet frame) and its DRAM-bound matmuls (22.2 ms) come next. - Deployment status in Autoware: opt-in, not the default. Autoware's launch defaults the LiDAR detector to CenterPoint; BEVFusion runs with
lidar_detection_model_type:=bevfusion, which selects this lidar-only model (model_namebevfusion_lidar). This bundle is not a ROS 2 node (Python API and HTTP) and not a certified Autoware component; do not use it for safety-critical driving decisions. - Lidar-only release. The package's second model, the camera-lidar profile (
bevfusion_camera_lidar.onnx+ the Swin-Tbevfusion_image_backbone.onnx), is not ported yet; it is planned for a later release of this repository. Later AWML BEVFusion-L releases (7 classes, intensity, Fourier voxel features) are not on the Autoware model repo and are not ported. - Capacity buckets: the sparse encoder runs in one of three trace shapes chosen per frame from the active voxel counts (
s/m/full; the trace cost follows the bucket, not the frame). A frame whose sparse level counts exceed the largest bucket (262,144 / 262,144 / 131,072 / 65,536 / 65,536 rows) is refused with an explicit error (InputError, HTTP 400) before the device runs; nothing is clipped. The densest of the 122 measured frames (98,334 voxels) is at about 37 % of it, but with the level ratios measured on PandaSet a frame from about 180,000 voxels on can hit it, below Autoware's 256,000-voxel cap. - Documented deviations from the node:
- deterministic voxelization: the first 10 points of each voxel in input order (the training-time rule), where the deployed node keeps whichever 10 its GPU atomics leave; above Autoware's 256,000-voxel cap the first voxels in input order are kept (
meta.voxel_overflow), where Autoware's choice depends on its hash order; - points at exactly (0, 0, 0) (the missing returns of organized clouds) and non-finite points are dropped on every input path;
- the voxelization, the VFE mean, the sparse rulebooks, the decode, both NMS and the class remapper run on the host, bit-identical to the CPU reference; so does the last step of the head,
score = σ(heatmap) · query_heatandcenter + bev_posin fp32. The network (sparse 3D encoder, SECOND, SECONDFPN, heatmap, local max, top-500, the transformer decoder layer, the box heads) runs on the device; - inputs are arrays or files, not
PointCloud2messages; Autoware's 2-sweep densification is atime_lagcolumn, client-sidesweepsor the server-sidestreamstate machine (checked on the device against the CPU reference).
- deterministic voxelization: the first 10 points of each voxel in input order (the training-time rule), where the deployed node keeps whichever 10 its GPU atomics leave; above Autoware's 256,000-voxel cap the first voxels in input order are kept (
- Precision policy of this release: HiFi4 math with fp32 accumulation everywhere; fp32 weights in the sparse encoder (its gather tables are bf16,
ttnn.embeddingtakes bf16 only; the first table is two-term); the 12 SECOND convs use two bf16 weight terms and two-term activations; the neck, head and decoder use fp32 weights; no bfp8. The two-term SECOND costs 16.8 ms of the 31.4 ms dense trace and was taken over from the TransFusion port without its own A/B on the frozen gates; relaxing it is an optimization item. - Decision thresholds turn small differences into different detections: the score threshold 0.1, the yaw-norm gate (a car, truck, bus or bicycle proposal is dropped when its rotation-head norm is below 0.3), circle NMS (0.5 m), IoU-BEV NMS and the class remapper (by box area: CAR to TRUCK at 12.1 m², to TRAILER at 36 m²). A proposal within the device's error of one of them can go either way. The gates pool their frames: on Autoware's sample rosbag frame 18 of 20 reference detections match (two TRUCK / TRAILER flips at the remapper's 36 m² edge on boxes that agree to 0.1 m, 35.77 vs 36.12 m²), on PandaSet 090 f40 127 of 129 (two low-score cars 0.4 m apart where circle NMS kept the other one). Score differences are below 0.02 on every gate frame; on 34 PandaSet frames new to the gates, 11 of about 4,400 matched pairs differ by 0.021-0.068 (ten cars and one pedestrian, 12-82 m away; the cause was not analysed). The agreement rows above include all of these.
- Domain: the weights were trained on TIER IV's internal data; on another LiDAR setup accuracy can drop without fine-tuning, as the upstream card says. The fp32 reference itself shows it on public data (indicative, CPU reference, not a benchmark): on nuScenes (32-beam) car recall falls from 0.96 below 10 m to 0.24 at 40-50 m, while on PandaSet (Pandar64) it stays at 0.80-1.00 up to 50 m; BICYCLE finds ridden bicycles but not parked ones; buses are the weakest class. On PandaSet the
base_linkx origin is the pose origin under the roof rig, not the rear axle (offset unknown, not corrected). The p150 agrees with the reference on these frames (rows above). - Validation scope: agreement with the fp32 CPU reference of the same network on public driving data (PandaSet, nuScenes v1.0-mini), Autoware's sample rosbag and
test.pcd; the knobs were validated at their Autoware defaults and at one non-default set. - Does not scale to multiple p150 in a mesh configuration. The build uses a 12×10 compute grid of Tensix cores: the dispatch functions move from one Tensix column to the ETH cores (
patches/tt-metal-eth-dispatch.patch), so this build assumes that you do not need chip-to-chip ethernet communication. dispatch="worker"(server:BEVFUSION_DISPATCH=worker) is an A/B opt-in. On a p150 it gives an 11×10 grid (54.43 ms per back-to-back frame instead of 53.09 ms ontest.pcd, all in the dense trace); if ETH dispatch is not available (tt-metal without the patch), the model falls back to it with a warning. The numbers on this card do not apply to that mode.- Batch 1, one frame per request; requests are serialised on the chip.
- Not an OpenAI-compatible API;
GET /v1/modelsis a stub so the tt-model ready card does not 404. - p150 power per inference was not measured (only board telemetry during the stage bench), so no efficiency comparison is made.
Licensing
- Weights: AutowareFoundation/bevfusion at tag
v2.0(commite1bf164b909d3c6c2642012f9b5ee4d827753220), Apache-2.0 per its model card. Not redistributed here: the package only points to them. The upstream card's training data: TIER IV's internal database (not public); no public-dataset license applies to the weights. - Pre- and post-processing ported from autoware_universe
perception/autoware_bevfusion(Apache-2.0); the thresholds, the class names and the remapper matrices are read from the weights repo's yaml files at load time. - Port and serving code (
code/): Apache-2.0.patches/tt-metal-eth-dispatch.patchmodifies tt-metal (Apache-2.0). - Sample data:
code/tt_bevfusion/samples/test_pcd.npzis derived from autoware_universe'sperception/autoware_ground_segmentation/test/data/test.pcd(Apache-2.0): the non-zero returns moved tobase_link. Only this redistributable sample ships; the public-dataset frames of the accuracy rows are not in this repository. - Demo media (
media/, sources and changes inmedia/ATTRIBUTION.md):- PandaSet renders: Contains data from PandaSet (Scale AI and Hesai), https://pandaset.org, licensed under CC BY 4.0 and the PandaSet Dataset Terms. Changes: converted to bird's-eye-view renders and resized camera images (licence plates and heads of nearby road users blurred) with drawn boxes. Scale AI and Hesai do not endorse this work. Cite: P. Xiao et al., PandaSet: Advanced Sensor Suite Dataset for Autonomous Driving, ITSC 2021.
- nuScenes render (
media/bevfusion_nuscenes_0103_k20_tt_NC.jpg), non-commercial, CC BY-NC-SA 4.0: Rendered from the nuScenes dataset, © Motional AD Inc., CC BY-NC-SA 4.0 and the nuScenes Terms of Use (https://www.nuscenes.org/terms-of-use). Non-commercial use only; adaptations under the same license. Motional does not endorse this work. Cite: H. Caesar et al., nuScenes: A Multimodal Dataset for Autonomous Driving, CVPR 2020. bevfusion_test_pcd_tt_vs_cpu.png: Apache-2.0.- The agreement rows use nuScenes v1.0-mini (CC BY-NC-SA 4.0), PandaSet (CC BY 4.0 + Dataset Terms) and Autoware's sample rosbag (license unstated) as inputs; only the resulting numbers are on this card.
Provenance
These are the exact sources the container image was built from:
| component | built from |
|---|---|
| tt-metal | 44d66500520fda9f2c7060c0f6b41ec48f7ab37e + patches/tt-metal-eth-dispatch.patch (sha256 08d0ddf6…; dirty tree: the image includes the patch) |
| weights | AutowareFoundation/bevfusion@e1bf164b909d3c6c2642012f9b5ee4d827753220 (tag v2.0), files bevfusion_lidar.onnx, ml_package_bevfusion_lidar.param.yaml, detection_class_remapper.param.yaml, deploy_metadata.yaml (ONNX sha256 5c290879…) |
| Autoware reference | autoware_universe 9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd (perception/autoware_bevfusion, package 0.53.0) |
| shared package | ttaw 0.23.0, vendored as code/tt_bevfusion/ttaw from the Autoware ports' shared common repository at commit cf8d069 (code/tt_bevfusion/ttaw/VENDORED.json: version, commit and per-file sha256) |
code/ digest (image) |
d9b76fee3d5aac3c (sha256, first 16 hex digits; built.code_sha256 of tt_kernel_manifest.json) |
| image | tt-model/bevfusion-p150:99f5a9441e7e (sha256:99f5a9441e7e5b11f0c94b6cd4e153ec3714369b1f5e99fd5e923bde7756d44b) |
| base images | build stage ghcr.io/tenstorrent/tt-metal/tt-metalium/ubuntu-22.04-dev-amd64:latest @ sha256:df9d279c7f85c17c6fad982d196802682d669cca1b7ced9cbaad8181339cd5fc; runtime stage docker.io/library/ubuntu:22.04 @ sha256:5ec03bb3441e8b0bf3b4f9cd4629a1ae763010dc3035bb8da3ae6cf026486401 (tt-model's FROM tags float; these are the digests this build resolved, see build_info.json) |
| built | 2026-10-09T21:18:31+00:00 by tt-model 0.1.0 |
Model tree for changh95/bevfusion-p150
Base model
AutowareFoundation/bevfusion




