--- license: cc-by-nc-4.0 tags: - world-model - diffusion - retrieval - navigation - embodied-ai --- # FAR checkpoints Checkpoint bundles for **FAR**, a latent-diffusion world model with a learned, action-conditioned retrieval memory, and its baselines, on the LoopNav, SoundSpaces and AI2-THOR corpora. - Paper: [Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models](https://arxiv.org/abs/2609.34677) - Code: https://github.com/sony/far - Project page: https://1202kbs.github.io/FAR-Project-Page/ ## Layout The repo mirrors the code's `models/` directory, so a download lands exactly where the configs, launchers and notebooks look: ``` //config.yaml # rebuilds the model (Hydra) //checkpoints/.pth.tar # EMA generator weights + memory state + step (inference only) oasis_500m_vit_vae.pth # ViT-VAE tokenizer for LoopNav latents agent_cell_detector.pt # latent agent-cell detector for the AI2-THOR-dyn probe bundles.json # this listing with sizes and SHA-256 ``` Bundles hold what inference needs (EMA generator, the trained or frozen memory state, the training step) and cannot resume training. The SDXL VAE used for the SoundSpaces and AI2-THOR corpora is fetched from `madebyollin/sdxl-vae-fp16-fix` automatically. ## Download ```bash python scripts/download_release.py # everything, ~13 GB python scripts/download_release.py --corpus ai2thor_v3 # one corpus python scripts/download_release.py --arm loopnav/far_multicue ``` or `huggingface-cli download 1202kbs/FAR-Checkpoints --local-dir models`. ## Bundles | Corpus | Bundle | Arm | Step | Size | |---|---|---|---|---| | AI2-THOR-dyn | `ai2thor_dyn/far_meta` | FAR -- Meta | 500k | 0.56 GB | | AI2-THOR-dyn | `ai2thor_dyn/far_multicue` | FAR -- Multi-Cue | 500k | 0.56 GB | | AI2-THOR-dyn | `ai2thor_dyn/temporal` | Temporal | 500k | 0.49 GB | | AI2-THOR-dyn | `ai2thor_dyn/worldmem` | WorldMem | 500k | 0.49 GB | | AI2-THOR-dyn | `ai2thor_dyn/retriever_object` | retriever (object cue), init of the FAR arms | 12.5k | 0.06 GB | | AI2-THOR v3 | `ai2thor_v3/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB | | AI2-THOR v3 | `ai2thor_v3/temporal` | Temporal | 700k | 0.49 GB | | AI2-THOR v3 | `ai2thor_v3/worldmem` | WorldMem | 700k | 0.49 GB | | AI2-THOR v3 | `ai2thor_v3/retriever_jepa` | retriever (visual), init of the FAR arm | 400k | 0.06 GB | | LoopNav | `loopnav/far_meta` | FAR -- Meta | 700k | 0.55 GB | | LoopNav | `loopnav/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB | | LoopNav | `loopnav/far_visual` | FAR -- Visual | 700k | 0.55 GB | | LoopNav | `loopnav/far_frozen_encoder` | ablation: the pre-trained visual retriever, frozen, no adapter | 700k | 0.55 GB | | LoopNav | `loopnav/far_ablation_z` | ablation: FAR -- Multi-Cue with a fixed cue weight (0.5) | 700k | 0.55 GB | | LoopNav | `loopnav/longlive_rag` | LongLive-RAG | 700k | 0.55 GB | | LoopNav | `loopnav/temporal` | Temporal | 700k | 0.49 GB | | LoopNav | `loopnav/worldmem` | WorldMem | 700k | 0.49 GB | | LoopNav | `loopnav/retriever_jepa` | retriever (visual), init of the FAR arms | 400k | 0.06 GB | | LoopNav | `loopnav/retriever_longlive` | retriever (content query), init of LongLive-RAG | 400k | 0.06 GB | | SoundSpaces v1 | `soundspaces_v1/far_meta` | FAR -- Meta | 700k | 0.55 GB | | SoundSpaces v1 | `soundspaces_v1/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB | | SoundSpaces v1 | `soundspaces_v1/temporal` | Temporal | 700k | 0.49 GB | | SoundSpaces v1 | `soundspaces_v1/worldmem` | WorldMem | 700k | 0.49 GB | | SoundSpaces v1 | `soundspaces_v1/retriever_audio` | retriever (audio), init of the SoundSpaces v1 and v2 FAR arms | 400k | 0.06 GB | | SoundSpaces v2 | `soundspaces_v2/far_meta` | FAR -- Meta | 700k | 0.55 GB | | SoundSpaces v2 | `soundspaces_v2/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB | | SoundSpaces v2 | `soundspaces_v2/temporal` | Temporal | 700k | 0.49 GB | | SoundSpaces v2 | `soundspaces_v2/worldmem` | WorldMem | 700k | 0.49 GB | "Temporal" and "WorldMem" are the recency and field-of-view retrieval baselines, "LongLive-RAG" the content-query retrieval baseline; the FAR arms differ in the cues the retriever fuses (metadata, visual, multi-cue); the two LoopNav ablations complete the paper's LoopNav ablation table. The retriever bundles are the contrastively pre-trained encoders the FAR launchers start from; evaluation does not need them (each FAR bundle already carries its trained retriever). ## License CC BY-NC 4.0. The generator and diffusion code these weights belong to are adapted from Navigation World Models and DiT (Meta Platforms, CC BY-NC 4.0); see the code repository's THIRD_PARTY_NOTICES.md. The SoundSpaces bundles (`soundspaces_v1/*`, `soundspaces_v2/*`) were trained on episodes rendered from Matterport3D scenes, so their use is additionally subject to the [Matterport3D Terms of Use](https://kaldir.vc.cit.tum.de/matterport/MP_TOS.pdf) (non-commercial academic research). ## Citation ```bibtex @article{kim2026far, title = {Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models}, author = {Kim, Beomsu and Lai, Chieh-Hsin and Nguyen, Bac and Bar, Amir and Ye, Jong Chul and Mitsufuji, Yuki}, journal = {arXiv preprint arXiv:2609.34677}, year = {2026}, url = {https://arxiv.org/abs/2609.34677} } ```