|
Download README.md from anhdao69/SimpleMemVLN-R2R-RxR15deg-Window8-Text-DualLane-FromBase: direct link, hf CLI and curl.
- Browser
- Download file 2.33 kB
-
https://huggingface.co/anhdao69/SimpleMemVLN-R2R-RxR15deg-Window8-Text-DualLane-FromBase/resolve/main/README.md
- Command line
-
hf download hf://anhdao69/SimpleMemVLN-R2R-RxR15deg-Window8-Text-DualLane-FromBase/README.md
-
curl -L -o README.md https://huggingface.co/anhdao69/SimpleMemVLN-R2R-RxR15deg-Window8-Text-DualLane-FromBase/resolve/main/README.md
2.33 kB
| base_model: Qwen/Qwen3.5-4B | |
| tags: | |
| - vision-language-navigation | |
| - simplememvln | |
| - dual-lane | |
| - r2r | |
| - rxr | |
| # SimpleMemVLN Window8 Text Dual-Lane — from base | |
| Completed two-epoch joint R2R + English-guide RxR 15-degree training run. | |
| `epoch-1` contains update 3,852; `epoch-2` contains the final update 7,704. | |
| Epoch 1 is a mid-schedule snapshot; epoch 2 completes the planned schedule. | |
| No navigation evaluation result is claimed for these weights. | |
| ## Training recipe | |
| Original Qwen3.5-4B base revision `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`. | |
| Training code: `a7e442cbd3c43cf3ec238eaa527e9c67f5ce2408` on | |
| https://github.com/anhdao69/SimpleMemVLN/tree/streaming_text_dual . | |
| 30,815 episodes; 3,128,624 action steps per dataset pass before distributed tail padding. | |
| Two epochs, 7,704 updates; 232 warmup updates; seed 429. | |
| Four H100 GPUs, one complete episode per rank per microstep, gradient accumulation 2, | |
| effective global batch 8. BF16, ZeRO-2, frozen vision/merger, trainable text backbone | |
| and output head; native LR 5e-6, lane LR 1e-4. Full-episode BPTT with activation | |
| checkpointing and long-episode CPU activation offload. | |
| Text output uses canonical action-text history and supervised assistant EOS. | |
| Window8 native attention retains instruction prefix plus eight observation/action | |
| groups. Native recurrent state persists. Four step lanes at layers 16/20/24/28 | |
| add 21,054,000 parameters, write only on each completed group's final separator, | |
| and read at token rate. Attention window batching is 16. | |
| ## Loading | |
| This is a custom SimpleMemVLN wrapper state dictionary, not a directly interchangeable | |
| Transformers AutoModel checkpoint. Download this repository, then use the recorded | |
| SimpleMemVLN source/environment and the pinned original Qwen base snapshot: | |
| ```python | |
| from qwen_vl.train.vln_runtime import load_checkpoint | |
| model, serializer = load_checkpoint('/path/to/download/epoch-2', '/path/to/Qwen3.5-4B-base') | |
| ``` | |
| `navigation.json` contains the saved model/serializer/memory contract. | |
| `SHA256SUMS.json` records artifact checksums. Optimizer states, RNG recovery files, | |
| training images and credentials are excluded. The full recovery checkpoint remains | |
| on the training server. Both epochs completed successfully. Epoch-2 action-weighted training loss was | |
| 0.04583318; training loss is not a navigation success metric. | |