BEVFormer-small weights (mirror)
A safetensors copy of the official BEVFormer-small checkpoint bevformer_small_epoch_24.pth, released by the
BEVFormer authors. Only the model state_dict is included. The optimizer state and training metadata are dropped,
and every tensor is bit-identical to the original.
This mirror provides a pinned Hugging Face source for the Tenstorrent Blackhole p150 port of Autoware's
autoware_tensorrt_bevformer
node (BEVFormer-small), published as the tt-model bundle changh95/bevformer-p150. Autoware's own
bevformer_small.onnx was exported from this checkpoint with
DerryHub/BEVFormer_tensorrt, and that ONNX file is no longer
downloadable.
Files
| File | Bytes | sha256 |
|---|---|---|
bevformer_small_epoch_24.safetensors |
238,389,932 | 51ba31289d85df5b90da32126bea21f284b03f08521323f6765459f1663c95d7 |
provenance.json |
source URL, source sha256, tensor counts |
- Source: https://github.com/zhiqi-li/storage/releases/download/v1.0/bevformer_small_epoch_24.pth
(linked from the BEVFormer model zoo),
712,087,430 bytes, sha256
4ebc5810201ca1e609c29452f07c08bc4d9c9ebfda5d7f8ee1034dd446274bc5. - Contents: 1,001 tensors: 897 float32 parameter and buffer tensors (59,568,963 elements) and 104 int64
num_batches_trackedcounters. Keys are unchanged from the mmcv checkpoint.
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
path = hf_hub_download("changh95/bevformer-small-weights", "bevformer_small_epoch_24.safetensors")
state_dict = load_file(path) # same keys as torch.load(...)["state_dict"]
Model
BEVFormer-small (config projects/configs/bevformer/bevformer_small.py):
- ResNet-101 backbone (caffe style, frozen BN) with DCNv2 in stages 3–4
- FPN on C5
- 150×150 BEV grid covering ±51.2 m
- 3 encoder layers (temporal self-attention + spatial cross-attention)
- 6 decoder layers with 900 queries and 10 nuScenes classes
It was trained on nuScenes v1.0-trainval for 24 epochs. The epoch-24 validation score in the released training log is NDS 0.4787 / mAP 0.3700.
License and training data
- This mirror is labelled
apache-2.0. That is the license of the BEVFormer code and of BEVFormer_tensorrt, which distribute and link these weights. The authors published the checkpoint as a release asset without a separate weight license. - Training data notice: the weights were trained on nuScenes, which is licensed under CC BY-NC-SA 4.0. Non-commercial terms may apply to uses of the weights. Commercial use of nuScenes requires a license from Motional. Check the dataset terms for your use case; this is not legal advice.
- All credit for the model and the weights goes to the BEVFormer authors.
Citation
@article{li2022bevformer,
title = {BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers},
author = {Li, Zhiqi and Wang, Wenhai and Li, Hongyang and Xie, Enze and Sima, Chonghao and Lu, Tong and Qiao, Yu and Dai, Jifeng},
journal = {arXiv preprint arXiv:2203.17270},
year = {2022}
}