--- license: apache-2.0 library_name: diffsynth pipeline_tag: text-to-video tags: - wan - wan2.1 - echo-memory - text-to-video base_model: Wan-AI/Wan2.1-T2V-1.3B --- # Echo-Memory `context_k1` — complete Wan 2.1 1.3B DiT DiffSynth-facing **full DiT** for the released Echo-Memory `context_k1` row. - Paper: [arXiv:2606.09803](https://arxiv.org/abs/2606.09803) - Code: [Echo-Team-Joy-Future-Academy-JD/Echo-Memory](https://github.com/Echo-Team-Joy-Future-Academy-JD/Echo-Memory) - Research overlay (HF): [Echo-Team/Echo-Memory](https://huggingface.co/Echo-Team/Echo-Memory) `context_k1/epoch-0.safetensors` ## What this file is `diffusion_pytorch_model-*-of-00006.safetensors` is a **complete Wan 2.1 T2V 1.3B DiT** (825 / 825 official keys), sharded for Hub upload. Load with `origin_file_pattern="diffusion_pytorch_model*.safetensors"`. It is official [Wan-AI/Wan2.1-T2V-1.3B](https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B) DiT keys with the Echo-Memory `context_k1` overlay already merged. Research extras (`action_mlp`, `self_attn_with_action`, SSM / spatial slots) are **not** included, so the file loads like a normal Wan transformer (`strict=True` on `pipe.dit`). T5 and VAE stay the official Wan files: - `Wan-AI/Wan2.1-T2V-1.3B` `models_t5_umt5-xxl-enc-bf16.pth` - `Wan-AI/Wan2.1-T2V-1.3B` `Wan2.1_VAE.pth` ## DiffSynth ```python import torch from diffsynth.pipelines.wan_video import WanVideoPipeline, ModelConfig from diffsynth.utils.data import save_video pipe = WanVideoPipeline.from_pretrained( torch_dtype=torch.bfloat16, device="cuda", model_configs=[ ModelConfig(model_id="SOTAowner/Echo-Memory-context-k1", origin_file_pattern="diffusion_pytorch_model*.safetensors"), ModelConfig(model_id="Wan-AI/Wan2.1-T2V-1.3B", origin_file_pattern="models_t5_umt5-xxl-enc-bf16.pth"), ModelConfig(model_id="Wan-AI/Wan2.1-T2V-1.3B", origin_file_pattern="Wan2.1_VAE.pth"), ], ) video = pipe(prompt="A toy bear on a table, the camera rotates around it", seed=42) save_video(video, "echo_memory_context_k1.mp4") ``` After this repo is mirrored on ModelScope, replace the DiT `model_id` with that ModelScope id (DiffSynth downloads from ModelScope by default). ## Notes - 30,000-step `epoch-0` fine-tune, 640×352, 81-frame chunks. - Camera-action / multi-chunk revisit protocol stays in the Echo-Memory repo; this file is the Wan DiT overlay only. - Apache-2.0. Please cite the Echo-Memory paper if you use it.