anhdao69's picture
Publish completed Window8 text dual-lane epoch 2 (step 7704)
fa24de8 verified
|
Raw History Blame Contribute Delete
2.33 kB
---
base_model: Qwen/Qwen3.5-4B
tags:
- vision-language-navigation
- simplememvln
- dual-lane
- r2r
- rxr
---
# SimpleMemVLN Window8 Text Dual-Lane — from base
Completed two-epoch joint R2R + English-guide RxR 15-degree training run.
`epoch-1` contains update 3,852; `epoch-2` contains the final update 7,704.
Epoch 1 is a mid-schedule snapshot; epoch 2 completes the planned schedule.
No navigation evaluation result is claimed for these weights.
## Training recipe
Original Qwen3.5-4B base revision `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`.
Training code: `a7e442cbd3c43cf3ec238eaa527e9c67f5ce2408` on
https://github.com/anhdao69/SimpleMemVLN/tree/streaming_text_dual .
30,815 episodes; 3,128,624 action steps per dataset pass before distributed tail padding.
Two epochs, 7,704 updates; 232 warmup updates; seed 429.
Four H100 GPUs, one complete episode per rank per microstep, gradient accumulation 2,
effective global batch 8. BF16, ZeRO-2, frozen vision/merger, trainable text backbone
and output head; native LR 5e-6, lane LR 1e-4. Full-episode BPTT with activation
checkpointing and long-episode CPU activation offload.
Text output uses canonical action-text history and supervised assistant EOS.
Window8 native attention retains instruction prefix plus eight observation/action
groups. Native recurrent state persists. Four step lanes at layers 16/20/24/28
add 21,054,000 parameters, write only on each completed group's final separator,
and read at token rate. Attention window batching is 16.
## Loading
This is a custom SimpleMemVLN wrapper state dictionary, not a directly interchangeable
Transformers AutoModel checkpoint. Download this repository, then use the recorded
SimpleMemVLN source/environment and the pinned original Qwen base snapshot:
```python
from qwen_vl.train.vln_runtime import load_checkpoint
model, serializer = load_checkpoint('/path/to/download/epoch-2', '/path/to/Qwen3.5-4B-base')
```
`navigation.json` contains the saved model/serializer/memory contract.
`SHA256SUMS.json` records artifact checksums. Optimizer states, RNG recovery files,
training images and credentials are excluded. The full recovery checkpoint remains
on the training server. Both epochs completed successfully. Epoch-2 action-weighted training loss was
0.04583318; training loss is not a navigation success metric.