SimpleMemVLN Window8 Text Dual-Lane โ from base
Completed two-epoch joint R2R + English-guide RxR 15-degree training run.
epoch-1 contains update 3,852; epoch-2 contains the final update 7,704.
Epoch 1 is a mid-schedule snapshot; epoch 2 completes the planned schedule.
No navigation evaluation result is claimed for these weights.
Training recipe
Original Qwen3.5-4B base revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a.
Training code: a7e442cbd3c43cf3ec238eaa527e9c67f5ce2408 on
https://github.com/anhdao69/SimpleMemVLN/tree/streaming_text_dual .
30,815 episodes; 3,128,624 action steps per dataset pass before distributed tail padding.
Two epochs, 7,704 updates; 232 warmup updates; seed 429.
Four H100 GPUs, one complete episode per rank per microstep, gradient accumulation 2,
effective global batch 8. BF16, ZeRO-2, frozen vision/merger, trainable text backbone
and output head; native LR 5e-6, lane LR 1e-4. Full-episode BPTT with activation
checkpointing and long-episode CPU activation offload.
Text output uses canonical action-text history and supervised assistant EOS. Window8 native attention retains instruction prefix plus eight observation/action groups. Native recurrent state persists. Four step lanes at layers 16/20/24/28 add 21,054,000 parameters, write only on each completed group's final separator, and read at token rate. Attention window batching is 16.
Loading
This is a custom SimpleMemVLN wrapper state dictionary, not a directly interchangeable Transformers AutoModel checkpoint. Download this repository, then use the recorded SimpleMemVLN source/environment and the pinned original Qwen base snapshot:
from qwen_vl.train.vln_runtime import load_checkpoint
model, serializer = load_checkpoint('/path/to/download/epoch-2', '/path/to/Qwen3.5-4B-base')
navigation.json contains the saved model/serializer/memory contract.
SHA256SUMS.json records artifact checksums. Optimizer states, RNG recovery files,
training images and credentials are excluded. The full recovery checkpoint remains
on the training server. Both epochs completed successfully. Epoch-2 action-weighted training loss was
0.04583318; training loss is not a navigation success metric.