bevfusion-p150

BEVFusion-L (Autoware autoware_bevfusion, lidar-only): the LiDAR-only BEVFusion network that Autoware's autoware_bevfusion node runs by default (bevfusion_lidar.onnx of release v2.0), ported to one Tenstorrent Blackhole p150 with tt-nn. One LiDAR point cloud in base_link in (x, y, z; one sweep, or Autoware's 2-sweep densified cloud with a time_lag column; intensity is not used); Autoware DetectedObjects-equivalent 3D boxes of CAR, TRUCK, BUS, TRAILER, BICYCLE and PEDESTRIAN out, with the node's pre- and post-processing. In Autoware it is an opt-in alternative: CenterPoint is the default LiDAR detector. The package's camera-lidar profile is not in this release yet. Weights: AutowareFoundation/bevfusion v2.0 · Paper: arXiv:2205.13542 · Autoware package: autoware_bevfusion · Training code: tier4/AWML (projects/BEVFusion; the lidar-only release behind v2.0 is not published) · Port: code/

Runs on p150 (mesh P150). Configuration: dispatch on the ETH cores, 1 command queue, 12×10 compute grid. Precision: fp32 weights in the sparse 3D encoder (bf16 gather tables, a two-term VFE table, fp32 bias and residual adds), two-term (hi + lo bf16) weights and activations in the SECOND backbone with fp32 sums, fp32 weights and bf16 activations in the neck, head and decoder, HiFi4 math with fp32 accumulation; no bfp8. All numbers on this card were measured in this configuration.

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart (Python)

Prerequisite: a tt-metal / ttnn environment at tt-metal 44d66500520 with patches/tt-metal-eth-dispatch.patch applied. ttnn is not on PyPI.

hf download changh95/bevfusion-p150 --exclude "image/*" --local-dir bevfusion-p150 && cd bevfusion-p150
pip install -e .                        # adds numpy<2, pillow, pyyaml, onnx, huggingface_hub, safetensors; ttnn and torch come from tt-metal
pip install -e ".[server,test]"         # optional: the HTTP server and the tests

Run the snippet from the model repo root: code/tt_bevfusion/samples/test_pcd.npz (Autoware's test.pcd) is a path relative to it.

from tt_bevfusion import BEVFusion

with BEVFusion.from_pretrained(device_id=0) as model:          # weights -> your HF cache, traces captured
    out = model("code/tt_bevfusion/samples/test_pcd.npz")      # path, (N, C) array, .npy/.npz/.pcd, or a PointCloud

for d in out.to_dicts()[:5]:
    print(d["label"], d["score"], d["center"], d["size"], d["yaw"])
  • from_pretrained downloads AutowareFoundation/bevfusion at the pinned commit e1bf164b909 (tag v2.0; only the lidar-only files, 34 MB) to your HF cache, opens the chip, builds the graph and captures four metal traces (the sparse encoder for each of three capacity buckets, and the dense network). The first load compiles the kernels (195.7 s with an empty JIT cache, measured on a loaded host); later loads take about 9 s.
  • The traces are captured during the load, so no call compiles anything: first / second call on the shipped sample 678.7 / 659.2 ms and 792.2 / 674.9 ms in two processes on a shared host (load average 8.5-9; most of a call is host pre-processing).
  • The with block releases the traces and closes the chip. Without with, call model.close().
Input A point cloud in base_link: path (.npy / .npz / .pcd / .bin), raw bytes with fmt=, an (N, C) float array or tensor, or a PointCloud; fields x, y, z (an intensity column is accepted and ignored: use_intensity: false). A column named time_lag (s) marks an already densified cloud. Optional sweeps= (client-side densification) or stream= (server-side, Autoware's own densification with poses).
Options score_threshold=0.1, circle_nms_dist_threshold=0.5, iou_nms_threshold=0.1, iou_nms_search_distance_2d=10.0, remap_classes=True, max_detections=None. from_pretrained(device_id=0, dispatch="eth", weights_dir=None, device=None).
Output Detections3D: boxes float32 [N, 7] (x, y, z, length, width, height, yaw), scores [N] (Autoware's existence_probability), label_ids [N] (ObjectClassification values), labels, meta (voxel and per-level sparse counts, the capacity bucket, overflow flags, per-stage counts, labels before the remapper), timing_ms. Sorted by score.
Methods out.to_dict() gives the /predict JSON. out.to_dicts() gives the detection list. tt_bevfusion.viz.render_bev(points, out.to_dicts()) draws a bird's-eye view.
  • The API gives the same output as the HTTP server /predict: both share the decoders, the device traces and the host post-processing (checked on the device by test_api_equals_server).
  • Every option is a host-side knob with Autoware's value as the default; none changes a device shape. A bad value is refused before the device runs.
  • One model uses one chip; calls from several threads are serialised.
  • Full reference: code/PYTHON.md. Runnable example: examples/quickstart.py (also writes quickstart_bev.png, the points and the boxes from above).

Serving (HTTP)

tt-model pull  changh95/bevfusion-p150 --with-weights
tt-model serve changh95/bevfusion-p150       # or with tt-cli: tt serve changh95/bevfusion-p150
python3 code/tt_bevfusion/server/client.py --points code/tt_bevfusion/samples/test_pcd.npz --out req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/bevfusion-p150
  • The image does not contain the weights. --with-weights puts them in your HF cache.
  • The server uses port 20000 (or the next free port). It is ready when the log shows Application startup complete.
  • client.py builds the request with the standard library only; add --url http://127.0.0.1:20000 to send it.
  • POST /predict: points (base64 .npy / .npz / .pcd / raw bin, base_link, fields x, y, z (intensity is accepted and ignored) or a pre-densified cloud with a time_lag column); optional sweeps (past sweeps with time_lag_s and T_current_from_sweep) or stream (id, timestamp_s, T_world_from_ego: server-side 2-sweep densification as Autoware), params (every option above), output_format. Also GET /health, GET /info, GET /v1/models (stub). Contract: SERVING.md section 3.
{
  "model": "bevfusion-p150",
  "frame_id": "base_link",
  "num_detections": 5,
  "detections": [
    {"label": "TRUCK", "label_id": 2, "score": 0.7584, "center": [24.184, 11.192, 1.535], "size": [8.107, 2.519, 2.849], "yaw": 3.1413},
    {"label": "TRUCK", "label_id": 2, "score": 0.5795, "center": [20.807, -0.494, 1.61], "size": [6.096, 2.339, 3.109], "yaw": -0.0158},
    ...
  ],
  "meta": {"variant": "lidar", "num_voxels": 21359, "voxel_overflow": false, "sparse": {"counts": {"l1": 21359, "l2": 27551, "l3": 15823, "l4": 7037, "out": 5393}, "bucket": "s", "overflow": false}, "counts": {"proposals": 500, "after_threshold": 5, "after_circle_nms": 5, "after_iou_nms": 5}},
  "timing_ms": {"preprocess": 589.116, "device": 84.441, "postprocess": 2.36, "total": 695.67, "decode": 19.397, "model_call": 675.96}
}
  • center / size ([length, width, height]) / yaw are in base_link metres / radians, Autoware's convention for this node (z is the box centre, yaw counter-clockwise from +x); label_id is the autoware_perception_msgs ObjectClassification value; detections are sorted by score.
  • No velocity is returned (Autoware decodes the network's velocity rows but does not publish them). TRAILER and remapped TRUCK labels come from Autoware's area-based class remapper (remap_classes=false turns it off; meta.label_before_remap keeps the network's class).

Demo

The p150 output on the shipped sample (Autoware's test.pcd, one sweep, Apache-2.0): the p150 boxes in colour over the fp32 CPU reference's boxes (light outlines); a box without a partner would be ringed and tagged.

BEVFusion-L on p150 vs the CPU reference, test.pcd

On public driving datasets (p150 outputs; the frames themselves are not in this repository; dataset boxes in thin grey; the model sees only the LiDAR, the camera images are for orientation):

PandaSet 019 (SF Embarcadero), frame 40: six cameras and the view from above PandaSet 019, frame 40: p150 over the CPU reference
PandaSet 090 (El Camino Real), frame 40: ±122.4 m and the front camera PandaSet 090, frames 30-49 (2 s at 10 Hz), ±122.4 m
nuScenes v1.0-mini scene-0103, key-frame 20: non-commercial, CC BY-NC-SA 4.0

Agreement with the fp32 CPU reference on these frames (same label, centres < 0.2 m): PandaSet 019 f40 128 of 128 reference detections; PandaSet 090 f40 127 of 129, 130 on the p150 (two low-score cars 0.4 m apart where circle NMS kept a different proposal, and one extra at score 0.101); nuScenes 0103 key-frame 20 29 of 29; the 20 frames of the animation 2,555 of 2,569.

PandaSet renders: contains data from PandaSet (Scale AI and Hesai), https://pandaset.org, licensed under CC BY 4.0 and the PandaSet Dataset Terms; converted to bird's-eye views and resized camera images (licence plates and heads of nearby road users blurred) with drawn boxes; Scale AI and Hesai do not endorse this work. The nuScenes render is non-commercial (CC BY-NC-SA 4.0): rendered from the nuScenes dataset, © Motional AD Inc., nuScenes Terms of Use; Motional does not endorse this work. Sources, changes and the full attributions: media/ATTRIBUTION.md.

Demo & Performances

Warm, batch 1, 2026-10-09. Latency: the stage bench of OPT_BASELINE.md (code/scripts/bench.py, 100 / 50 iterations) on the shipped sample test.pcd (48,048 points, one sweep, 21,359 voxels, sparse capacity bucket s) and on a PandaSet frame (213,007 points, two sweeps, 95,109 voxels, bucket m; not shipped), re-checked on the release code (30 / 10 iterations: back-to-back 53.09 / 112.01 ms, model() 658.2 / 2,781.7 ms); the served rows from uvicorn on the host (the app the container runs) and a loopback client, 30 requests of the shipped sample. The host is shared with other jobs, so the host stages move with its load; the device rows repeat to 0.01 ms. Accuracy: the p150 output against the fp32 CPU reference of the same network (same weights, same pre- and post-processing, identical voxel and rulebook input), matched by label and BEV centre distance (< 0.2 m).

Metric Performance
Agreement with the fp32 CPU reference, shipped sample test.pcd 5 / 5 reference detections matched (same label, centres < 0.2 m; max |Δscore| 0.0043)
Agreement, PandaSet 019 f40 (the frozen sample gate; not shipped) 128 / 128 (max |Δscore| 0.0081)
Agreement, 7 development frames (test.pcd, PandaSet ×2, nuScenes mini_val ×3, Autoware's sample rosbag) 355 / 359 reference detections matched (recall 0.989, precision 0.986); max |Δscore| 0.0086
Agreement, 10 held-out frames never used for any precision choice (PandaSet ×4, nuScenes ×6) 680 / 680 (max |Δscore| 0.014)
Agreement, PandaSet 019 and 090 frames 30-49 (40 frames; 34 of them in no gate set) 5,174 / 5,204 (recall 0.9942, precision 0.9952); 11 of 5,174 matched pairs differ in score by more than 0.02 (max 0.068)
Module PCC vs the fp32 reference (replay outputs; worst of PandaSet 019 f40, nuScenes 0103 k20, rosbag #144) sparse BEV ≥ 0.999989 · SECOND blocks ≥ 0.999985 · heat ≥ 0.999981 · top-500 overlap ≥ 0.994 · worst teacher-forced head 0.999461
Python model() call, shipped sample test.pcd (host pre-processing, H2D, traces, D2H, host post-processing) 657.1 ms p50 (1.52 calls/s)
Python model() call, PandaSet frame (213,007 points, 2 sweeps; not shipped) 2,786.0 ms p50 (0.36 calls/s)
… of which the host sparse rulebook build (numpy) 580.8 ms (test.pcd) · 2,388.4 ms (PandaSet): 86-88 % of the call
Served /predict timing_ms.total (uvicorn on the host, the shipped sample) 697.1 ms median
Served client round trip, loopback (709 kB request) 705.1 ms median
Device traces, one blocking frame (sparse + dense) 53.23 ms (21.7 + 31.5) · PandaSet 112.15 ms (80.7 + 31.5)
Back-to-back frames, device time per frame 53.09 ms (18.8 frames/s) · PandaSet 112.02 ms (8.9 frames/s)
Input packing + H2D · D2H · host post-processing (test.pcd) 18.7 + 6.6 · 0.33 · 2.5 ms (PandaSet: 85.7 + 20.5 · 0.32 · 23.6 ms)
from_pretrained load: empty JIT cache (loaded host) / warm cache 195.7 s / 9.3 s (capture of the four traces 4.0 s)

All numbers in this table were measured with dispatch on the ETH cores, 1 command queue and a 12×10 compute grid on one p150, with the published precision. No paper-style metric (mAP / NDS) is claimed: these weights were trained on TIER IV's internal data only, so a score on nuScenes or PandaSet would measure the domain gap, not the port; the accuracy rows are agreement with the CPU reference. Details: VERIFICATION_2026-10-09.md, OPT_BASELINE.md, OPT_REPORT.md.

No GPU comparison: no GPU was available on the host where this port was built and measured, so this card makes no GPU speed claim. The reference rows are the port's own fp32 CPU reference on the same host (a correctness baseline, not a speed target). Autoware runs this network as TensorRT engines with spconv plugins and GPU voxelization; this port was validated against the fp32 reference, not against a TensorRT engine. p150 power per inference was not measured (the stage bench only logged the board's telemetry, median 66-70 W over whole runs including the host stages), so no efficiency comparison is made.

Caveats

  • First release: baseline port, optimization pending. End to end is host-bound: the numpy sparse rulebook build (C13: four sub-manifold and four strided neighbour maps) takes 581 ms of the 657 ms call on test.pcd and 2,388 ms of 2,786 ms on a dense PandaSet frame. It is the first optimization target (the container has no C compiler, so: vectorised lookups or a device rulebook builder; OPT_REPORT.md). On the device the sparse encoder's im2col gathers (40.9 ms of the PandaSet frame) and its DRAM-bound matmuls (22.2 ms) come next.
  • Deployment status in Autoware: opt-in, not the default. Autoware's launch defaults the LiDAR detector to CenterPoint; BEVFusion runs with lidar_detection_model_type:=bevfusion, which selects this lidar-only model (model_name bevfusion_lidar). This bundle is not a ROS 2 node (Python API and HTTP) and not a certified Autoware component; do not use it for safety-critical driving decisions.
  • Lidar-only release. The package's second model, the camera-lidar profile (bevfusion_camera_lidar.onnx + the Swin-T bevfusion_image_backbone.onnx), is not ported yet; it is planned for a later release of this repository. Later AWML BEVFusion-L releases (7 classes, intensity, Fourier voxel features) are not on the Autoware model repo and are not ported.
  • Capacity buckets: the sparse encoder runs in one of three trace shapes chosen per frame from the active voxel counts (s / m / full; the trace cost follows the bucket, not the frame). A frame whose sparse level counts exceed the largest bucket (262,144 / 262,144 / 131,072 / 65,536 / 65,536 rows) is refused with an explicit error (InputError, HTTP 400) before the device runs; nothing is clipped. The densest of the 122 measured frames (98,334 voxels) is at about 37 % of it, but with the level ratios measured on PandaSet a frame from about 180,000 voxels on can hit it, below Autoware's 256,000-voxel cap.
  • Documented deviations from the node:
    • deterministic voxelization: the first 10 points of each voxel in input order (the training-time rule), where the deployed node keeps whichever 10 its GPU atomics leave; above Autoware's 256,000-voxel cap the first voxels in input order are kept (meta.voxel_overflow), where Autoware's choice depends on its hash order;
    • points at exactly (0, 0, 0) (the missing returns of organized clouds) and non-finite points are dropped on every input path;
    • the voxelization, the VFE mean, the sparse rulebooks, the decode, both NMS and the class remapper run on the host, bit-identical to the CPU reference; so does the last step of the head, score = σ(heatmap) · query_heat and center + bev_pos in fp32. The network (sparse 3D encoder, SECOND, SECONDFPN, heatmap, local max, top-500, the transformer decoder layer, the box heads) runs on the device;
    • inputs are arrays or files, not PointCloud2 messages; Autoware's 2-sweep densification is a time_lag column, client-side sweeps or the server-side stream state machine (checked on the device against the CPU reference).
  • Precision policy of this release: HiFi4 math with fp32 accumulation everywhere; fp32 weights in the sparse encoder (its gather tables are bf16, ttnn.embedding takes bf16 only; the first table is two-term); the 12 SECOND convs use two bf16 weight terms and two-term activations; the neck, head and decoder use fp32 weights; no bfp8. The two-term SECOND costs 16.8 ms of the 31.4 ms dense trace and was taken over from the TransFusion port without its own A/B on the frozen gates; relaxing it is an optimization item.
  • Decision thresholds turn small differences into different detections: the score threshold 0.1, the yaw-norm gate (a car, truck, bus or bicycle proposal is dropped when its rotation-head norm is below 0.3), circle NMS (0.5 m), IoU-BEV NMS and the class remapper (by box area: CAR to TRUCK at 12.1 m², to TRAILER at 36 m²). A proposal within the device's error of one of them can go either way. The gates pool their frames: on Autoware's sample rosbag frame 18 of 20 reference detections match (two TRUCK / TRAILER flips at the remapper's 36 m² edge on boxes that agree to 0.1 m, 35.77 vs 36.12 m²), on PandaSet 090 f40 127 of 129 (two low-score cars 0.4 m apart where circle NMS kept the other one). Score differences are below 0.02 on every gate frame; on 34 PandaSet frames new to the gates, 11 of about 4,400 matched pairs differ by 0.021-0.068 (ten cars and one pedestrian, 12-82 m away; the cause was not analysed). The agreement rows above include all of these.
  • Domain: the weights were trained on TIER IV's internal data; on another LiDAR setup accuracy can drop without fine-tuning, as the upstream card says. The fp32 reference itself shows it on public data (indicative, CPU reference, not a benchmark): on nuScenes (32-beam) car recall falls from 0.96 below 10 m to 0.24 at 40-50 m, while on PandaSet (Pandar64) it stays at 0.80-1.00 up to 50 m; BICYCLE finds ridden bicycles but not parked ones; buses are the weakest class. On PandaSet the base_link x origin is the pose origin under the roof rig, not the rear axle (offset unknown, not corrected). The p150 agrees with the reference on these frames (rows above).
  • Validation scope: agreement with the fp32 CPU reference of the same network on public driving data (PandaSet, nuScenes v1.0-mini), Autoware's sample rosbag and test.pcd; the knobs were validated at their Autoware defaults and at one non-default set.
  • Does not scale to multiple p150 in a mesh configuration. The build uses a 12×10 compute grid of Tensix cores: the dispatch functions move from one Tensix column to the ETH cores (patches/tt-metal-eth-dispatch.patch), so this build assumes that you do not need chip-to-chip ethernet communication.
  • dispatch="worker" (server: BEVFUSION_DISPATCH=worker) is an A/B opt-in. On a p150 it gives an 11×10 grid (54.43 ms per back-to-back frame instead of 53.09 ms on test.pcd, all in the dense trace); if ETH dispatch is not available (tt-metal without the patch), the model falls back to it with a warning. The numbers on this card do not apply to that mode.
  • Batch 1, one frame per request; requests are serialised on the chip.
  • Not an OpenAI-compatible API; GET /v1/models is a stub so the tt-model ready card does not 404.
  • p150 power per inference was not measured (only board telemetry during the stage bench), so no efficiency comparison is made.

Licensing

  • Weights: AutowareFoundation/bevfusion at tag v2.0 (commit e1bf164b909d3c6c2642012f9b5ee4d827753220), Apache-2.0 per its model card. Not redistributed here: the package only points to them. The upstream card's training data: TIER IV's internal database (not public); no public-dataset license applies to the weights.
  • Pre- and post-processing ported from autoware_universe perception/autoware_bevfusion (Apache-2.0); the thresholds, the class names and the remapper matrices are read from the weights repo's yaml files at load time.
  • Port and serving code (code/): Apache-2.0. patches/tt-metal-eth-dispatch.patch modifies tt-metal (Apache-2.0).
  • Sample data: code/tt_bevfusion/samples/test_pcd.npz is derived from autoware_universe's perception/autoware_ground_segmentation/test/data/test.pcd (Apache-2.0): the non-zero returns moved to base_link. Only this redistributable sample ships; the public-dataset frames of the accuracy rows are not in this repository.
  • Demo media (media/, sources and changes in media/ATTRIBUTION.md):
    • PandaSet renders: Contains data from PandaSet (Scale AI and Hesai), https://pandaset.org, licensed under CC BY 4.0 and the PandaSet Dataset Terms. Changes: converted to bird's-eye-view renders and resized camera images (licence plates and heads of nearby road users blurred) with drawn boxes. Scale AI and Hesai do not endorse this work. Cite: P. Xiao et al., PandaSet: Advanced Sensor Suite Dataset for Autonomous Driving, ITSC 2021.
    • nuScenes render (media/bevfusion_nuscenes_0103_k20_tt_NC.jpg), non-commercial, CC BY-NC-SA 4.0: Rendered from the nuScenes dataset, © Motional AD Inc., CC BY-NC-SA 4.0 and the nuScenes Terms of Use (https://www.nuscenes.org/terms-of-use). Non-commercial use only; adaptations under the same license. Motional does not endorse this work. Cite: H. Caesar et al., nuScenes: A Multimodal Dataset for Autonomous Driving, CVPR 2020.
    • bevfusion_test_pcd_tt_vs_cpu.png: Apache-2.0.
    • The agreement rows use nuScenes v1.0-mini (CC BY-NC-SA 4.0), PandaSet (CC BY 4.0 + Dataset Terms) and Autoware's sample rosbag (license unstated) as inputs; only the resulting numbers are on this card.

Provenance

These are the exact sources the container image was built from:

component built from
tt-metal 44d66500520fda9f2c7060c0f6b41ec48f7ab37e + patches/tt-metal-eth-dispatch.patch (sha256 08d0ddf6…; dirty tree: the image includes the patch)
weights AutowareFoundation/bevfusion@e1bf164b909d3c6c2642012f9b5ee4d827753220 (tag v2.0), files bevfusion_lidar.onnx, ml_package_bevfusion_lidar.param.yaml, detection_class_remapper.param.yaml, deploy_metadata.yaml (ONNX sha256 5c290879…)
Autoware reference autoware_universe 9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd (perception/autoware_bevfusion, package 0.53.0)
shared package ttaw 0.23.0, vendored as code/tt_bevfusion/ttaw from the Autoware ports' shared common repository at commit cf8d069 (code/tt_bevfusion/ttaw/VENDORED.json: version, commit and per-file sha256)
code/ digest (image) d9b76fee3d5aac3c (sha256, first 16 hex digits; built.code_sha256 of tt_kernel_manifest.json)
image tt-model/bevfusion-p150:99f5a9441e7e (sha256:99f5a9441e7e5b11f0c94b6cd4e153ec3714369b1f5e99fd5e923bde7756d44b)
base images build stage ghcr.io/tenstorrent/tt-metal/tt-metalium/ubuntu-22.04-dev-amd64:latest @ sha256:df9d279c7f85c17c6fad982d196802682d669cca1b7ced9cbaad8181339cd5fc; runtime stage docker.io/library/ubuntu:22.04 @ sha256:5ec03bb3441e8b0bf3b4f9cd4629a1ae763010dc3035bb8da3ae6cf026486401 (tt-model's FROM tags float; these are the digests this build resolved, see build_info.json)
built 2026-10-09T21:18:31+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for changh95/bevfusion-p150

Finetuned
(1)
this model

Collection including changh95/bevfusion-p150

Paper for changh95/bevfusion-p150