echo / code /doc /memory_mechanisms.md
amonshano's picture
Add Echo-Memory codebase used for this run (CC BY 4.0, JD Echo Team) (part 2)
eafbe80 verified
|
Raw
History Blame Contribute Delete
4.75 kB
# Memory Mechanisms
This note maps the paper's memory rows to the repository implementation and explains the modeling role of each family. Echo-Memory treats memory as a controlled intervention on what information from chunk 1 is stored and how chunk 2 reads it back during denoising.
## Modeling View
All rows use the same action-conditioned Wan DiT backbone and the same two-chunk training/evaluation setup:
1. **Context chunk:** clean history frames are encoded into latent/context tokens, optionally with matched camera RT actions.
2. **Target chunk:** noisy target latents are denoised while the selected memory mechanism exposes information from the context chunk.
3. **Read-out:** memory is injected through raw context concatenation, compressed context tokens, spatial memory tokens, or recurrent state-space modules attached to DiT blocks.
The ablations are designed to change only the memory pathway while keeping the backbone, action conditioning, resolution, chunk length, and training schedule aligned.
## Paper Rows
| Paper family | Paper row / repo name | What is stored or read | Main code path | Training entry |
| --- | --- | --- | --- | --- |
| Raw context | `context_k1`, `context_k5`, `context_k20` | Uncompressed retrieved context frames. `K=1` is the anchor/I2V floor; `K=5/20` are context-learning capacity rows. | `diffsynth/pipelines/wan_video_new.py` context latent path | `train/context_learning/run_pre_qkv_ctx{1,5,20}.sh` |
| Compression | `framepack_weight` | Context tokens are kept at the same length but temporally reweighted. | `diffsynth/models/memory/framepack_weight.py` | `train/memory_baselines_basic/run_ablation_framepack_weight_two_chunk.sh` |
| Compression | `framepack_len_r2`, `framepack_len_r4` | Context latents and matched RT actions are pooled along time. | `diffsynth/models/memory/framepack_length.py` | `train/memory_baselines_basic/run_ablation_framepack_len_r{2,4}_two_chunk.sh` |
| Compression | `framepack_hybrid_r2`, `framepack_hybrid_r4` | Length compression plus token reweighting. | `wan_video_new.py` + FramePack helpers | `train/memory_baselines_basic/run_ablation_framepack_hybrid_r*_weight_two_chunk.sh` |
| Token-grid | `spatial_mem` | Context tokens are time-averaged and summarized into learned grid tokens. This is the implementation behind the currently reported `spatial_mem` row; it does **not** reconstruct depth or 3D geometry. | `diffsynth/models/memory/spatial_grid_memory.py` | `train/memory_baselines_basic/run_spatial_memory_baseline.sh` |
| Token-grid | `spatial_inject_none`, `spatial_concat_text`, `spatial_cross_attn_readout` | Same token-grid storage, different read-out: withheld, text-KV concat, or dedicated cross-attention. | `spatial_grid_memory.py` read-out helpers | matching `run_ablation_spatial_*_two_chunk.sh` scripts |
| Geometry-grounded spatial | `geometry_spatial_mem` | A static scene is reconstructed outside the DiT using depth, intrinsics, extrinsics, and TSDF fusion. The fused point cloud is rendered along the target trajectory, VAE-encoded, and converted into conditioning tokens. | `diffsynth/models/memory/geometry_spatial_memory.py` | `train/memory_baselines_basic/run_geometry_spatial_memory_baseline.sh` |
| State-space | `block_wise_ssm` | Paper-aligned recurrent state attached to selected DiT blocks. Checkpoint keys contain `block_wise_ssm.*`. | `diffsynth/models/memory/block_wise_ssm.py` | `train/memory_baselines_basic/run_ablation_block_wise_ssm_two_chunk.sh` |
| State-space | `videossm_hybrid` | Legacy VideoSSM hybrid baseline: depthwise temporal-conv state-space-like module. Checkpoint keys contain `videossm_hybrid.*`. | `diffsynth/models/memory/videossm_hybrid.py` | `train/memory_baselines_basic/run_videossm_hybrid_baseline.sh` |
## Naming Rules
- Do not describe `SpatialGridMemory` or the existing `spatial_mem` results as the
geometry-grounded method from arXiv:2506.05284. It is a token-grid baseline.
- Use **Geometry-grounded Spatial Memory** only when the metadata supplies
rendered static geometry through `geometry_memory` (or a configured column).
The geometry extractor is the external reconstruction pipeline: depth and
cameras → TSDF-fused static point cloud → target-view renders. The model-side
encoder does not estimate depth itself.
- Use **Block-wise SSM** only for `--use_block_wise_ssm` / `BlockWiseStateSpaceMemory`.
- Use **VideoSSM hybrid** only for the legacy `--use_videossm_hybrid` / `HybridStateSpaceMemory` baseline.
- Use **Context learning** for raw-context capacity rows (`K=1/5/20`), not for compact memory modules.
- Keep checkpoint folder names stable; `env/memory_baseline_runtime.py` and `inference/unified_inference.py` infer memory profiles from those names.