--- base_model: Qwen/Qwen3.5-4B tags: - vision-language-navigation - simplememvln - dual-lane - r2r - rxr --- # SimpleMemVLN Window8 Text Dual-Lane — from base Completed two-epoch joint R2R + English-guide RxR 15-degree training run. `epoch-1` contains update 3,852; `epoch-2` contains the final update 7,704. Epoch 1 is a mid-schedule snapshot; epoch 2 completes the planned schedule. No navigation evaluation result is claimed for these weights. ## Training recipe Original Qwen3.5-4B base revision `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`. Training code: `a7e442cbd3c43cf3ec238eaa527e9c67f5ce2408` on https://github.com/anhdao69/SimpleMemVLN/tree/streaming_text_dual . 30,815 episodes; 3,128,624 action steps per dataset pass before distributed tail padding. Two epochs, 7,704 updates; 232 warmup updates; seed 429. Four H100 GPUs, one complete episode per rank per microstep, gradient accumulation 2, effective global batch 8. BF16, ZeRO-2, frozen vision/merger, trainable text backbone and output head; native LR 5e-6, lane LR 1e-4. Full-episode BPTT with activation checkpointing and long-episode CPU activation offload. Text output uses canonical action-text history and supervised assistant EOS. Window8 native attention retains instruction prefix plus eight observation/action groups. Native recurrent state persists. Four step lanes at layers 16/20/24/28 add 21,054,000 parameters, write only on each completed group's final separator, and read at token rate. Attention window batching is 16. ## Loading This is a custom SimpleMemVLN wrapper state dictionary, not a directly interchangeable Transformers AutoModel checkpoint. Download this repository, then use the recorded SimpleMemVLN source/environment and the pinned original Qwen base snapshot: ```python from qwen_vl.train.vln_runtime import load_checkpoint model, serializer = load_checkpoint('/path/to/download/epoch-2', '/path/to/Qwen3.5-4B-base') ``` `navigation.json` contains the saved model/serializer/memory contract. `SHA256SUMS.json` records artifact checksums. Optimizer states, RNG recovery files, training images and credentials are excluded. The full recovery checkpoint remains on the training server. Both epochs completed successfully. Epoch-2 action-weighted training loss was 0.04583318; training loss is not a navigation success metric.