nord-base / README.md
shubh-gupta-ai's picture
Update README.md
e25eb59 verified
|
Raw
History Blame Contribute Delete
2.76 kB
---
license: cc-by-nc-sa-4.0
base_model: Qwen/Qwen2.5-VL-3B-Instruct
pipeline_tag: image-text-to-text
library_name: transformers
tags:
- autonomous-driving
- vision-language-action
- trajectory-prediction
- navsim
- vllm
---
# NoRD-Base: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning (SFT only)
**CVPR 2026** | [arXiv](https://arxiv.org/abs/2602.21172) | [Project Page](https://nord-vla-ai.github.io/) | [GitHub](https://github.com/Applied-Intuition-Open-Source/nord)
*Ishaan Rawal · Shubh Gupta · Yihan Hu · Wei Zhan*
This is the **SFT-only ablation baseline** from the paper: [Qwen2.5-VL-3B](https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct)
supervised fine-tuned to directly predict driving trajectories as discrete tokens without
chain-of-thought reasoning, using 3× fewer tokens than reasoning-based VLA baselines. This
checkpoint is also the initialization used for the Dr. GRPO stage that produces the paper's
main result.
| Model | Training | NAVSIM PDMS |
|---|---|---|
| [`nord`](https://huggingface.co/AppliedIntuitionResearch/nord) | SFT + Dr. GRPO | 0.8626 |
| **`nord-base`** (this repo) | SFT only | **0.7273** |
## Usage
Install the [`nord`](https://github.com/Applied-Intuition-Open-Source/nord) client and serve with vLLM:
```bash
vllm serve AppliedIntuitionResearch/nord-base --served-model-name qwen --dtype bfloat16 --port 8000
```
```python
from PIL import Image
import nord
agent = nord.NordAgent.from_pretrained("nord-base")
output = agent.predict(nord.NordInput(
cameras=[Image.open("fl.jpg"), Image.open("front.jpg"), Image.open("fr.jpg")],
ego_velocity_ms=(8.3, 0.0),
driving_command="straight",
))
print(output.trajectory.shape) # (40, 3) — x, y, heading at 10 Hz, 4 seconds
```
Full inference and NAVSIM evaluation instructions: [github.com/Applied-Intuition-Open-Source/nord](https://github.com/Applied-Intuition-Open-Source/nord).
This repo bundles the K-Disc trajectory tokenizer vocab (`vocab.pkl`, 2048 clusters) used to
decode the model's output tokens into `(x, y, heading)` trajectories.
## Citation
```bibtex
@inproceedings{rawal2026nord,
title={NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning},
author={Rawal, Ishaan and Gupta, Shubh and Hu, Yihan and Zhan, Wei},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year={2026}
}
```
## License
This checkpoint is released under [CC BY-NC-SA 4.0](https://creativecommons.org/licenses/by-nc-sa/4.0/)
(non-commercial). The `nord` inference client code is separately licensed under Apache 2.0 —
see [github.com/Applied-Intuition-Open-Source/nord](https://github.com/Applied-Intuition-Open-Source/nord).