Temporal-NoPE / README.md
JiaqiFeng's picture
Document recovery-v3 visual evidence
443d026 verified
|
Raw
History Blame Contribute Delete
2.4 kB
---
license: apache-2.0
library_name: pytorch
tags:
- video-generation
- causal-video
- temporal-rope
- nope
base_model: Wan-AI/Wan2.1-T2V-1.3B
---
# Temporal NoPE runtime checkpoint
This public repository carries the immutable stable causal-video checkpoint used
as the common initialization for the Temporal NoPE experiments in
[Aurora-edu/NoPE](https://github.com/Aurora-edu/NoPE).
## Checkpoint identity
- File: `checkpoints/framewise/causal_cd.pt`
- Size: `5,676,220,819` bytes
- SHA-256: `c951a6b4804cd637fecfc857e9a59d54b6e4cd7846c38360baa3cf3b525f0d22`
- State-dict key: `generator_ema`
- Architecture: Wan 2.1 T2V 1.3B causal model
The portable configuration constructs the exact 30-layer, 1536-dimensional,
12-head architecture without first downloading the redundant public base-DiT
weights, then strictly loads all 825 checkpoint entries. The strict load gate
passes with no missing or unexpected keys.
## Download
```bash
hf download JiaqiFeng/Temporal-NoPE \
checkpoints/framewise/causal_cd.pt \
configs/temporal_nope/d4_source_rope_uniformscale_c010_portable.yaml \
wan_models/Wan2.1-T2V-1.3B/Wan2.1_VAE.pth \
wan_models/Wan2.1-T2V-1.3B/config.json \
prompt_cache/eval_umt5_bf16_lmdb/data.mdb \
--local-dir /path/to/NoPE
```
The bundle includes the Wan VAE and the read-only eval prompt-embedding LMDB,
so cached-prompt visual preflights do not need to load the UMT5 encoder. The
public base DiT is intentionally omitted because the portable config strictly
loads every model parameter from `causal_cd.pt`.
Always verify the SHA-256 checksum before running an experiment.
## Recovery-v3 visual evidence
The repository also preserves the complete small-gate outputs for the final
attention-scaling diagnosis under `outputs/temporal_nope/`:
- `recovery_v3_d4_source_rope_uniformscale_c010_portable_micro_visual_gate_seed0_illidan`
- `recovery_v3_d5_source_rope_uniformscale_c030_portable_micro_visual_gate_seed0_illidan`
- `recovery_v3_attndiag_a0_c000_val004_seed0_illidan`
- `recovery_v3_attndiag_d4_c010_val004_seed0_illidan`
These 44 files include all generated MP4s, resolved configs, sample metadata,
rank completion markers, and raw temporal-attention records. Both D4 and D5
failed mandatory human review; the outputs are failure evidence, not promoted
model samples. The GitHub reports contain the formal visual decisions and
mechanism analysis.