--- license: apache-2.0 pipeline_tag: robotics tags: - autoware - ros2 - autonomous-driving - e2e - planning - vad - carla - simulation - tensorrt - onnx --- # VAD for Autoware (`tensorrt_vad`) End-to-end autonomous driving model (variant `carla_tiny`) used by the [`autoware_tensorrt_vad`](https://github.com/autowarefoundation/autoware_universe/tree/main/e2e/autoware_tensorrt_vad) node in [Autoware](https://github.com/autowarefoundation/autoware). The model follows the **VAD (Vectorized Autonomous Driving)** [1] architecture and was optimized for TensorRT deployment with NVIDIA's [DL4AGX](https://github.com/NVIDIA/DL4AGX) workflow. It replaces the traditional localization, perception, and planning modules with a single neural network: 6 surround-view camera images go in, an ego trajectory comes out, with predicted objects and map elements as auxiliary outputs. > **Simulation only.** This model is trained exclusively on CARLA simulation data from the > [Bench2Drive](https://github.com/Thinklab-SJTU/Bench2Drive) benchmark [2]. It is intended for use with the > CARLA simulator and is **not suitable for real-vehicle deployment**. ## Model overview | | | | --- | --- | | Task | End-to-end driving: camera images to ego trajectory, with predicted objects and map elements as auxiliary outputs | | Architecture | VAD (vectorized scene representation), split into an image backbone and planning heads | | Variant | `carla_tiny` (trained on Bench2Drive CARLA data) | | Object classes | `car`, `van`, `truck`, `bicycle`, `traffic_sign`, `traffic_cone`, `traffic_light`, `pedestrian`, `others` | | Map classes | `Broken`, `Solid`, `SolidSolid`, `Center`, `TrafficLight`, `StopSign` | | Runtime | TensorRT (FP16 backbone, FP32 heads by default) via the `autoware_tensorrt_vad` ROS 2 node | | Format | ONNX (Autoware builds the TensorRT engines locally on first launch) | | Compatibility | `autoware_tensorrt_vad` >= 0.1.0, TensorRT >= 8.6, CUDA >= 11.4 | | License | Apache-2.0 | The network is split into three ONNX sub-models, matching how the node consumes them: 1. **Backbone** (`vad-carla-tiny_backbone.onnx`): image feature extraction from the 6 camera images, run in FP16 by default. 2. **Head** (`vad-carla-tiny_head.onnx`): temporal planning head that consumes the backbone features (`mlvl_feats.0`) together with previous-frame history, run in FP32 by default. 3. **Head, no history** (`vad-carla-tiny_head_no_prev.onnx`): planning head used for the first frame, when no temporal history exists yet, run in FP32 by default. Key architecture parameters (from `vad-carla-tiny.param.json`): 640x384 model input per camera with ImageNet RGB normalization, 106x200 (height x width) BEV grid with 256 feature channels, 3 transformer decoder layers, 300 object queries (9 classes, 6 predicted trajectory modes over 6 timesteps), 100 map queries (6 classes, 20 points per polyline), and 6 ego planning commands over 6 timesteps. ## Files | File | Description | | --- | --- | | `vad-carla-tiny_backbone.onnx` | Image feature extraction backbone | | `vad-carla-tiny_head.onnx` | Temporal planning head (uses previous-frame history) | | `vad-carla-tiny_head_no_prev.onnx` | Planning head for the first frame (no history) | | `vad-carla-tiny.param.json` | Model package parameters: architecture, normalization, class definitions | | `deploy_metadata.yaml` | Deployment metadata recording the artifact version of this repository | > **TensorRT engines are not distributed here.** TensorRT engines are specific to the GPU architecture and > TensorRT version they are built on and are not portable, so Autoware builds them locally from the ONNX files > on first launch and caches them next to the ONNX files. ## Inputs and outputs (as used by the node) **Inputs** | Topic | Message type | Description | | --- | --- | --- | | `~/input/image0..5` | `sensor_msgs/msg/Image` (raw or compressed) | Camera images in training order: FRONT, BACK, FRONT_LEFT, BACK_LEFT, FRONT_RIGHT, BACK_RIGHT | | `~/input/camera_info0..5` | `sensor_msgs/msg/CameraInfo` | Camera calibration for cameras 0-5 | | `~/input/kinematic_state` | `nav_msgs/msg/Odometry` | Vehicle odometry | | `~/input/acceleration` | `geometry_msgs/msg/AccelWithCovarianceStamped` | Vehicle acceleration | **Outputs** | Topic | Message type | Description | | --- | --- | --- | | `~/output/trajectory` | `autoware_planning_msgs/msg/Trajectory` | Selected ego trajectory | | `~/output/trajectories` | `autoware_internal_planning_msgs/msg/CandidateTrajectories` | All 6 candidate trajectories | | `~/output/objects` | `autoware_perception_msgs/msg/PredictedObjects` | Predicted objects with future trajectories | | `~/output/map` | `visualization_msgs/msg/MarkerArray` | Predicted map elements | The camera order is fixed by the training data and must not be changed. The node remaps the CARLA object classes to Autoware classes via `object_class_remapper_carla_tiny.param.yaml` (`car`/`van` to `CAR`, `truck` to `TRUCK`, `bicycle` to `BICYCLE`, `pedestrian` to `PEDESTRIAN`, static classes and `others` to `UNKNOWN`). ## Usage in Autoware Autoware downloads these artifacts to `~/autoware_data/ml_models/vad/` (note: the local directory is named `vad`, not `tensorrt_vad`), and the node loads them from `$(var model_path)/v0.1/` as configured in `config/vad_carla_tiny.param.yaml`. After setting up CARLA following the `autoware_carla_interface` instructions, launch the end-to-end VAD system with: ```bash ros2 launch autoware_launch e2e_simulator.launch.xml \ map_path:=$HOME/autoware_data/maps/Town01 \ vehicle_model:=sample_vehicle \ sensor_model:=carla_sensor_kit \ simulator_type:=carla \ use_e2e_planning:=true ``` The node itself is started via `vad_carla_tiny.launch.xml`, which defaults `model_path` to `$HOME/autoware_data/ml_models/vad`. On first run it builds the TensorRT engines from the ONNX files, optimized for the local GPU, and caches them for subsequent runs. See the [package README](https://github.com/autowarefoundation/autoware_universe/tree/main/e2e/autoware_tensorrt_vad) for the full parameter reference. ## Training The model was trained on CARLA simulation data from the Bench2Drive benchmark and deployed with NVIDIA's DL4AGX TensorRT workflow: - Architecture: [VAD](https://github.com/hustvl/VAD) (Jiang et al., arXiv:2303.12077) - Training data / benchmark: [Bench2Drive](https://github.com/Thinklab-SJTU/Bench2Drive) (Jia et al., arXiv:2406.03877) - Deployment / TensorRT optimization: [DL4AGX](https://github.com/NVIDIA/DL4AGX) Initial release (v0.1) date: 2025-11-04. Further training details (epochs, exact frame counts) are not publicly documented. ## Limitations - **Simulation-only training**: the model is trained exclusively on CARLA simulator data, which does not capture the full complexity and variability of real-world driving. Do not deploy it on a real vehicle. - **No dynamic mission control**: the current implementation lacks a high-level command interface, so the model cannot switch driving behaviors (for example from "follow lane" to "turn right at the next intersection") at runtime. The default command is `LANE_FOLLOW`. ## Provenance | | | | --- | --- | | Original source | `https://awf.ml.dev.web.auto/planning/models/tensorrt_vad/carla_tiny/v0.1/` | | This repository | `AutowareFoundation/tensorrt_vad`, tag `v0.1` | ## Citation ```bibtex @inproceedings{jiang2023vad, title = {VAD: Vectorized Scene Representation for Efficient Autonomous Driving}, author = {Jiang, Bo and Chen, Shaoyu and Xu, Qing and Liao, Bencheng and Chen, Jiajie and Zhou, Helong and Zhang, Qian and Liu, Wenyu and Huang, Chang and Wang, Xinggang}, booktitle = {ICCV}, year = {2023} } @article{jia2024bench2drive, title = {Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving}, author = {Jia, Xiaosong and Yang, Zhenjie and Li, Qifeng and Zhang, Zhiyuan and Yan, Junchi}, journal = {arXiv preprint arXiv:2406.03877}, year = {2024} } ``` ## References - [1] Jiang et al., "VAD: Vectorized Scene Representation for Efficient Autonomous Driving", arXiv:2303.12077, 2023. Code: - [2] Jia et al., "Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving", arXiv:2406.03877, 2024. Code: - [3] NVIDIA DL4AGX, TensorRT optimization for autonomous driving workloads: