--- base_model: - Qwen/Qwen3-VL-8B-Instruct --- # LookStep [简体中文](README_ZH.md) ## Model details | Field | Value | | -------------------- | -------------------------------------------------------------------------- | | Base model | `Qwen/Qwen3-VL-8B-Instruct` | | Architecture | `Qwen3VLForConditionalGeneration` | | Model type | `qwen3_vl` | | Parameters | 8,767,123,696 | | Checkpoint format | safetensors, 4 shards, 750 tensors | | Indexed tensor bytes | 17,534,247,392 bytes | | Fine-tuning method | Full-parameter SFT (`tuner_type=full`) | | Final optimizer step | 18,888 | | Training epoch | 1.0 | | Training precision | BF16 | | Training max length | 8,192 tokens | | Input modality | Navigation instruction plus front-facing RGB observations | | Output | Structured LookStep state, candidate outcomes, memory decision, and action | The online policy receives the instruction, up to six long-term event-memory frames, up to two recent frames, and the current RGB frame. It generates: ```xml ... ... keep|drop ... ... ... ... ... MOVE_FORWARD|TURN_LEFT|TURN_RIGHT|STOP ``` Use this checkpoint with the LookStep simulation code to reproduce the paper's R2R-CE and RxR-CE Val-Unseen main results. It is intended for research in embodied vision-language navigation under the published Habitat configuration. It is not a general-purpose chatbot, a standalone image captioner, a safety controller, or a validated controller for physical robots. ## Training procedure | Hyperparameter | Value | | --------------------------- | --------------------- | | GPUs | 8 × NVIDIA A100 80 GB | | Epochs | 1 | | Per-device train batch size | 2 | | Gradient accumulation | 8 | | Global batch size | 128 | | Optimizer steps | 18,888 | | Optimizer | `adamw_torch_fused` | | Learning rate | `2e-5` | | Scheduler | cosine | | Warmup ratio | 0.03 | | Weight decay | 0.01 | | Adam betas / epsilon | 0.9, 0.95 / `1e-8` | | Max gradient norm | 1.0 | | Distributed training | DeepSpeed ZeRO-2 | | Vision encoder | frozen | | Visual aligner | frozen | | LLM | trainable | | Model/data seeds | 42 / 42 | ## Reproduce with LookStep Create the pinned environment and validate the downloaded model first: ```bash conda env create -f LookStep/simulation/environment.yml conda activate lookstep-simulation MODEL_PATH=/path/to/downloaded/checkpoint-18888 \ PROCESSOR_PATH=/path/to/Qwen3-VL-8B-Instruct \ DATA_ROOT=/path/to/data \ bash LookStep/reproduce_paper.sh check-sim ``` Run a two-episode smoke test, followed by both complete main benchmarks: ```bash MODEL_PATH=/path/to/downloaded/checkpoint-18888 \ PROCESSOR_PATH=/path/to/Qwen3-VL-8B-Instruct \ DATA_ROOT=/path/to/data \ bash LookStep/reproduce_paper.sh smoke-r2r MODEL_PATH=/path/to/downloaded/checkpoint-18888 \ PROCESSOR_PATH=/path/to/Qwen3-VL-8B-Instruct \ DATA_ROOT=/path/to/data \ CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \ bash LookStep/reproduce_paper.sh eval-all bash LookStep/reproduce_paper.sh verify ``` ## Citation ``` @inproceedings{ lookstep, title={LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory}, author={Kun-Yang Yu, Yingzhe Li, Hongyu Xu, Shi-Yu Tian, Zhi Zhou, Yang Chen, Ming Yang, Sheng Wang, Qing Yu, Lan-Zhe Guo, Yu-Feng Li}, booktitle={The 2026 Conference on Empirical Methods in Natural Language Processing}, year={2026} } ``` If you have any question, please email to [yuky@lamda.nju.edu.cn](mailto:yuky@lamda.nju.edu.cn) (Kun-Yang Yu)