echo / README.md
amonshano's picture
Add model card: attribution, licences, training config
5c44b7e verified
|
Raw
History Blame Contribute Delete
3.47 kB
---
license: cc-by-4.0
base_model: Wan-AI/Wan2.1-T2V-1.3B
pipeline_tag: text-to-video
tags:
- world-model
- video-generation
- action-conditioned
- memory
- ablation
- research-artifact
---
# echo — Echo-Memory ablation arm (`ucpe` + two-chunk memory)
> **This is an unofficial derivative artifact, not the official Echo-Memory release.**
> It contains one ablation arm re-trained from the Echo-Memory codebase, produced for
> internal study. For the official code, paper and checkpoints, use the links below.
## Upstream / attribution
| | |
|---|---|
| Original work | **Echo-Memory: A Controlled Study of Memory in Action World Models** |
| Authors | Echo Team @ Joy Future Academy, JD |
| Paper | [arXiv:2606.09803](https://arxiv.org/abs/2606.09803) |
| Official code | https://github.com/Echo-Team-Joy-Future-Academy-JD/Echo-Memory |
| Official checkpoints | https://huggingface.co/Echo-Team/Echo-Memory |
| Code license | CC BY 4.0 |
| Base model | [Wan2.1-T2V-1.3B](https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B) (Apache 2.0) |
The Echo-Memory code included here is redistributed under CC BY 4.0 with attribution to
the Echo Team. The weights are derived from Wan2.1-T2V-1.3B, which is Apache 2.0; that
license and its notices continue to apply to the derived weights.
## What this repository contains
```
code/ Echo-Memory codebase as used for this run
checkpoints/
Step-10000.safetensors 3.5 GB
Step-20000.safetensors 3.5 GB
Step-30000.safetensors 3.5 GB (final, max_train_steps = 30000)
```
## Training configuration
Single arm of a memory-mechanism ablation, run name
`memory_baselines_basic_abl_ucpe_memory_two_chunk`.
- Base weights: `Wan2.1-T2V-1.3B` (DiT + `umt5-xxl` text encoder + `Wan2.1_VAE`)
- Trainable: `dit` only, `--remove_prefix_in_ckpt pipe.dit.`
- Memory: `--enable_context_memory --context_source replay`
`--prev_chunk_frames 81 --context_memory_frames 5`
- Mechanism under test: `--use_cgla_memory --cgla_mechanism ucpe`
`--cgla_every_n_blocks 4 --cgla_aux_loss_weight 0.01`
- Also enabled: `--use_moc --moc_temperature 1.0`, `--train_cam_pose`,
`--add_action_attn`, `--use_rt_relative`, `--use_anchor_frame`, `--cfg_target_only`
- Resolution / length: 352 x 640, 81 frames
- Optimisation: lr 2e-5, 1 epoch, batch 1/device, grad-accum 1,
`--timestep_shift 15`, `--context_drop_prob 0.1`, 30,000 steps
- Hardware: 1x NVIDIA H20
## Data — not redistributed here
Training used the **Context-as-Memory** dataset (Unreal-Engine rendered environments,
~331 GB of frames). It is **not** included in this repository: it is third-party data
and the copy used here carried no redistribution licence. Obtain it from its original
source.
## Status and caveats
- This is **one arm of an ablation sweep**, uploaded as a research artifact.
Its results had not been analysed at the time of upload — no quality claim is made,
and it should not be read as the best or recommended configuration.
- It is **not** a reproduction of, or a replacement for, the official Echo-Memory
checkpoints. Comparisons against the paper's numbers are not valid without matching
the official evaluation protocol.
## Citation
Please cite the original work:
```bibtex
@article{echomemory2026,
title = {Echo-Memory: A Controlled Study of Memory in Action World Models},
author = {Echo Team, Joy Future Academy, JD},
journal= {arXiv preprint arXiv:2606.09803},
year = {2026}
}
```