# Data 本文只描述当前使用的 v4 数据体系。早期 v1 dense-text 数据和 v2 direct-K/V 数据已经完成其历史实验, 不再作为后续训练的数据标准。 ## 1. 数据集总览 当前训练数据集名称为 `hyworldplay_predictor_prefeature_v4`,schema 为 `predictor_context_prefeature_bf16_v1`。 | 项目 | 数值 | | --- | ---: | | 图文 pair | 100 | | 每个 pair 的动作轨迹 | 4 | | 轨迹总数 | 400 | | 每个视频帧数 | 125 | | latent frames | 32 | | 每条轨迹 chunks | 8 | | 每个 chunk latent frames | 4 | | 每个 chunk 去噪步 | 4 | | 完整 manifest records | 3200 | | 可训练 records | 2800 | | 每个 record 的监督 pair | `0→1`、`1→2`、`2→3` | | 训练 pair 总数 | 8400 | | Context blocks | `[0,1,52,53]` | | 保存精度 | BF16 safetensors | | 实际 NVMe 占用 | 约 1.3 TB | `train_eval_manifest.jsonl` 排除 chunk 0,因为 v4 需要前一 chunk 的 same-timestep hidden; `predictor-disca` 为保持相同训练分布也沿用这 2800 条记录,但不会加载 previous-chunk hidden 或 Context pre-feature。 ## 2. 数据来源和图像预处理 训练图文 pair 来自: ```text /mnt/s3files/s3-us-west2-default/zoubin/cz/projects/VBench/ vbench2_beta_i2v/data/origin ``` 选择方法: 1. 按路径名排序全部图片; 2. 使用 Python `random.Random(0).sample(...)` 随机选择 100 张; 3. caption 使用源图片文件名去掉扩展名后的文本; 4. 图像先保持宽高比,用 Lanczos 放大至覆盖目标区域,再做中心裁剪; 5. 保存为 832×480 JPEG,quality 95、无 chroma subsampling。 预处理入口: ```bash python tools/prepare_predictor_prefeature_cases.py ``` 输出为 `cases.jsonl`、`dataset_config.json` 和 `input_images/case_XXXX.jpg`。 ## 3. 动作与视频配置 四种动作轨迹固定为: | action_id | 名称 | pose | | ---: | --- | --- | | 0 | `w_s` | `w-15,s-16` | | 1 | `a_d` | `a-15,d-16` | | 2 | `up_down` | `up-15,down-16` | | 3 | `left_right` | `left-15,right-16` | 生成配置为 seed 0、480×832、125 帧、32 latent frames、8 chunks、每 chunk 4 latent frames、 每个 chunk 4 个去噪步。构建训练 tensor 的同时保存每条轨迹对应的 125 帧 MP4,作为训练集 Full-DiT 评测参考视频。 ## 4. Teacher 数据语义 Teacher 使用 HY-WorldPlay action-distilled AR checkpoint: ```text /mnt/s3files/s3-us-west2-default/zoubin/cz/checkpoints/hy_worldplay/ huggingface/hub/models--tencent--HY-WorldPlay/snapshots/ f4c29235647707b571479a69b569e4166f9f5bf8/ ar_distilled_action_model/diffusion_pytorch_model.safetensors ``` 数据构建使用精确 Full-DiT 路径: - 四个去噪步都运行完整 54 层; - 保留正式 WorldPlay text K/V 和 AR history vision memory; - 关闭 denoise cache、history cache reuse、Context pooling/pruning; - 关闭 sparse attention、SageAttention 和 FP8 GEMM; - `prompt_rewrite=false`; - Transformer 与落盘 feature 使用 BF16; - offloading 只影响显存,不改变模型定义。 Context prefill 的语义是 `joint_noncausal_full_dit_prefill`:同一目标 chunk 选中的历史帧作为完整窗口 联合、非因果计算。因此深层 pre-feature 与目标窗口绑定,不能将历史 chunk 跨窗口去重。 ## 5. 文件组织 ```text hyworldplay_predictor_prefeature_v4/ ├── dataset_config.json ├── cases.jsonl ├── manifest.jsonl # 3200 records ├── train_eval_manifest.jsonl # 2800 records,排除 chunk 0 ├── input_images/ # 100 张 832×480 图像 ├── cases/ # 100 个 case tensors ├── text_kv_exact_all54_bf16/ # 100 个 Full Teacher 54 层 Text K/V ├── steps/ # 3200 个 chunk step tensors │ └── case_XXXX/action_YY/chunk_ZZ.safetensors ├── context_prefeature/ # 12800 个 pre-feature tensors │ ├── block_00/ │ ├── block_01/ │ ├── block_52/ │ └── block_53/ ├── videos/ # 400 个 Full-DiT MP4 ├── manifests/ # 8 个 worker manifests └── logs/ ``` 实际空间分布: | 目录 | 文件数 | 大小 | | --- | ---: | ---: | | `cases/` | 100 | 2.97 GB | | `text_kv_exact_all54_bf16/` | 100 | 37.20 GiB | | `steps/` | 3200 | 337.60 GB | | `context_prefeature/` | 12800 | 1022.63 GB | | `videos/` | 400 | 0.40 GB | | `input_images/` | 100 | 0.02 GB | ## 6. Tensor schema ### 6.1 Case tensor 每个 case 只保存一次图像条件和四个候选 block 的 Text K/V: ```text image_condition_latent [1,32,1,30,52] BF16 block_{00,01,52,53}_k_txt [1,16,903,128] BF16 block_{00,01,52,53}_v_txt [1,16,903,128] BF16 text_valid_mask [1,903] bool ``` Text K/V 在磁盘中补齐到 903 tokens,真实长度由 `text_valid_mask` 指示。 ### 6.2 Full Teacher Text K/V `text_kv_exact_all54_bf16/` 是离线数据集的一部分,不是 pre-feature,也不是 rollout sidecar。每个 case 一个文件: ```text case_XXXX_text_kv_exact_all54_bf16.safetensors block_{00..53}_k_txt [1,16,903,128] BF16 block_{00..53}_v_txt [1,16,903,128] BF16 text_valid_mask [1,903] bool ``` 真实有效长度为 739–751 tokens,训练读取时按 `text_valid_mask` 去掉补零区,再交给 Teacher/Predictor。 构建会将 blocks `0、1、52、53` 与原 case tensor 逐元素比对,确保它来自相同的 frozen text prefill。 ### 6.3 Step tensor 每个 chunk 保存共享动作/相机/RoPE 信息,以及四个去噪步的监督目标: ```text action_labels [1,4] int64 target_viewmats [1,4,4,4] BF16 target_Ks [1,4,3,3] BF16 rope_temporal_size [1] int64 start_rope_start_idx [1] int64 step_{0..3}_timestep [1] FP32 step_{0..3}_noisy_sample [1,32,4,30,52] BF16 step_{0..3}_frame_condition [1,4,2048] BF16 step_{0..3}_final_hidden [1,6240,2048] BF16 step_{0..3}_velocity [1,32,4,30,52] BF16 ``` 65 通道 Predictor 输入不重复落盘。Dataset 使用 noisy sample、case-level image latent 和 I2V mask 重建 `[B,65,4,30,52]` 输入。 ### 6.4 Context pre-feature 每个 record、每个 block 保存 K/V 线性投影之前的 `img_modulated`: ```text img_modulated [1,S_context,2048] BF16 context_valid_mask [1,S_context] bool selected_frame_indices [context_frames] int64 context_viewmats [1,context_frames,4,4] BF16 context_Ks [1,context_frames,3,3] BF16 rope_temporal_size [1] int64 start_rope_start_idx [1] int64 ``` `S_context = context_frames × 30 × 52`。八个 chunks 的 Context frames 依次为 `0,4,8,12,16,20,20,20`。磁盘文件保持变长且不 padding;v4 collate 按 batch 最大 Context 长度 动态 padding,并用 valid mask 屏蔽补齐部分。 训练时由冻结的 Teacher `img_attn_k`、`img_attn_v`、`img_attn_k_norm` 和相同的 RoPE/ProPE 路径 重建 Context Vision K/V。直接 K/V 与重建 K/V 的验收阈值为 relative L2 ≤ 5e-3、cosine ≥ 0.9999。 ## 7. Manifest 关系 主键是 `(case_id, action_id, chunk_id)`。每条记录包含 case、step、四个 Context 文件的相对路径, 以及 caption、源图、动作、seed、分辨率、Teacher checkpoint 和 Context frame 数。 合并 manifest 时: - 检查 3200 个主键唯一; - 验证所有引用文件存在; - chunk 1–7 增加 `previous_chunk_id` 和 `previous_step_tensor_file`; - `manifest.jsonl` 保留全部 3200 条; - `train_eval_manifest.jsonl` 保留 chunk 1–7,共 2800 条。 ## 8. 构建、恢复和持久化 NVMe 是构建和训练读取位置,S3 是持久化 source of truth: ```text NVMe: /mnt/local_nvme/zoubin/cz/hyworldplay_predictor_prefeature_v4 S3: s3://s3-us-west2-default/zoubin/cz/projects/ HY-WorldPlay-DEV-Predictor/datasets/hyworldplay_predictor_prefeature_v4 项目挂载路径: /mnt/s3files/s3-us-west2-default/zoubin/cz/projects/ HY-WorldPlay-DEV-Predictor/datasets/hyworldplay_predictor_prefeature_v4 ``` 完整构建流程: ```bash python tools/prepare_predictor_prefeature_cases.py bash tools/launch_predictor_prefeature_build.sh ``` 第二条命令使用 8 个 GPU worker,先写 NVMe,随后合并 manifest 并同步到 S3。writer 支持按 `case/action/chunk` 跳过完整文件,任务中断后可重新执行。 全 54 层精确 Text K/V 单独构建并持久化到同一数据集: ```bash bash tools/launch_predictor_text_kv_all54_build.sh ``` 新机器训练前执行: ```bash bash tools/ensure_predictor_prefeature_nvme.sh ``` 需要全层 Text K/V 的 rollout 训练再执行: ```bash bash tools/ensure_predictor_text_kv_all54_nvme.sh ``` 如果 NVMe 不完整,该脚本使用 `aws s3 sync` 从持久前缀恢复并重新检查数量。手动将合法更新同步回 S3 使用: ```bash bash tools/sync_predictor_prefeature_to_s3.sh ``` 不要把 NVMe 当作唯一副本,也不要在未核对目标前缀时使用 `aws s3 sync --delete`。 ## 9. 完整性验收 ```bash python tools/validate_predictor_prefeature_dataset.py ``` 验收条件: - 100 case tensors; - 3200 step tensors; - 12800 Context pre-feature tensors; - 400 个 125 帧、832×480 视频; - 3200 条无重复 manifest records 和 2800 条训练 records; - 所有 tensor shape、dtype、finite、mask 和 Context frame 数正确; - 四个 block 的 selected frame indices 一致; - frozen final layer 能由 final hidden 与 frame condition 复现 velocity; - 抽样重建的 65 通道输入与 Teacher 实际输入一致。 ## 10. Validation/Test holdout 独立评估集为 `hyworldplay_predictor_vbench_val25_test50`: - VBench metadata 共 355 个 pair; - 排除训练集 100 个 pair 后剩 255 个; - 使用 seed `20260718` 随机选择 75 个; - validation 25 个、test 50 个,两者互斥且均与训练集互斥; - 每个 pair 生成四种动作; - 已保存 300 个 Full-DiT 和 300 个 Reuse 视频; - Full 为 steps `[0,1,2,3]`,Reuse 为 Full `[0,3]`、复用 `[1,2]`。 入口: ```bash bash tools/launch_vbench_holdout_full_reuse.sh ``` 持久目录: ```text datasets/hyworldplay_predictor_vbench_val25_test50 ```