Instructions to use Cccccz/HY with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use Cccccz/HY with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("Cccccz/HY", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
Download data.md from Cccccz/HY: direct link, hf CLI and curl.
- Browser
- Download file 10.6 kB
-
https://huggingface.co/Cccccz/HY/resolve/main/data.md
- Command line
-
hf download hf://Cccccz/HY/data.md
-
curl -L -o data.md https://huggingface.co/Cccccz/HY/resolve/main/data.md
Data
本文只描述当前使用的 v4 数据体系。早期 v1 dense-text 数据和 v2 direct-K/V 数据已经完成其历史实验, 不再作为后续训练的数据标准。
1. 数据集总览
当前训练数据集名称为 hyworldplay_predictor_prefeature_v4,schema 为
predictor_context_prefeature_bf16_v1。
| 项目 | 数值 |
|---|---|
| 图文 pair | 100 |
| 每个 pair 的动作轨迹 | 4 |
| 轨迹总数 | 400 |
| 每个视频帧数 | 125 |
| latent frames | 32 |
| 每条轨迹 chunks | 8 |
| 每个 chunk latent frames | 4 |
| 每个 chunk 去噪步 | 4 |
| 完整 manifest records | 3200 |
| 可训练 records | 2800 |
| 每个 record 的监督 pair | 0→1、1→2、2→3 |
| 训练 pair 总数 | 8400 |
| Context blocks | [0,1,52,53] |
| 保存精度 | BF16 safetensors |
| 实际 NVMe 占用 | 约 1.3 TB |
train_eval_manifest.jsonl 排除 chunk 0,因为 v4 需要前一 chunk 的 same-timestep hidden;
predictor-disca 为保持相同训练分布也沿用这 2800 条记录,但不会加载 previous-chunk hidden 或
Context pre-feature。
2. 数据来源和图像预处理
训练图文 pair 来自:
/mnt/s3files/s3-us-west2-default/zoubin/cz/projects/VBench/
vbench2_beta_i2v/data/origin
选择方法:
- 按路径名排序全部图片;
- 使用 Python
random.Random(0).sample(...)随机选择 100 张; - caption 使用源图片文件名去掉扩展名后的文本;
- 图像先保持宽高比,用 Lanczos 放大至覆盖目标区域,再做中心裁剪;
- 保存为 832×480 JPEG,quality 95、无 chroma subsampling。
预处理入口:
python tools/prepare_predictor_prefeature_cases.py
输出为 cases.jsonl、dataset_config.json 和 input_images/case_XXXX.jpg。
3. 动作与视频配置
四种动作轨迹固定为:
| action_id | 名称 | pose |
|---|---|---|
| 0 | w_s |
w-15,s-16 |
| 1 | a_d |
a-15,d-16 |
| 2 | up_down |
up-15,down-16 |
| 3 | left_right |
left-15,right-16 |
生成配置为 seed 0、480×832、125 帧、32 latent frames、8 chunks、每 chunk 4 latent frames、 每个 chunk 4 个去噪步。构建训练 tensor 的同时保存每条轨迹对应的 125 帧 MP4,作为训练集 Full-DiT 评测参考视频。
4. Teacher 数据语义
Teacher 使用 HY-WorldPlay action-distilled AR checkpoint:
/mnt/s3files/s3-us-west2-default/zoubin/cz/checkpoints/hy_worldplay/
huggingface/hub/models--tencent--HY-WorldPlay/snapshots/
f4c29235647707b571479a69b569e4166f9f5bf8/
ar_distilled_action_model/diffusion_pytorch_model.safetensors
数据构建使用精确 Full-DiT 路径:
- 四个去噪步都运行完整 54 层;
- 保留正式 WorldPlay text K/V 和 AR history vision memory;
- 关闭 denoise cache、history cache reuse、Context pooling/pruning;
- 关闭 sparse attention、SageAttention 和 FP8 GEMM;
prompt_rewrite=false;- Transformer 与落盘 feature 使用 BF16;
- offloading 只影响显存,不改变模型定义。
Context prefill 的语义是 joint_noncausal_full_dit_prefill:同一目标 chunk 选中的历史帧作为完整窗口
联合、非因果计算。因此深层 pre-feature 与目标窗口绑定,不能将历史 chunk 跨窗口去重。
5. 文件组织
hyworldplay_predictor_prefeature_v4/
├── dataset_config.json
├── cases.jsonl
├── manifest.jsonl # 3200 records
├── train_eval_manifest.jsonl # 2800 records,排除 chunk 0
├── input_images/ # 100 张 832×480 图像
├── cases/ # 100 个 case tensors
├── text_kv_exact_all54_bf16/ # 100 个 Full Teacher 54 层 Text K/V
├── steps/ # 3200 个 chunk step tensors
│ └── case_XXXX/action_YY/chunk_ZZ.safetensors
├── context_prefeature/ # 12800 个 pre-feature tensors
│ ├── block_00/
│ ├── block_01/
│ ├── block_52/
│ └── block_53/
├── videos/ # 400 个 Full-DiT MP4
├── manifests/ # 8 个 worker manifests
└── logs/
实际空间分布:
| 目录 | 文件数 | 大小 |
|---|---|---|
cases/ |
100 | 2.97 GB |
text_kv_exact_all54_bf16/ |
100 | 37.20 GiB |
steps/ |
3200 | 337.60 GB |
context_prefeature/ |
12800 | 1022.63 GB |
videos/ |
400 | 0.40 GB |
input_images/ |
100 | 0.02 GB |
6. Tensor schema
6.1 Case tensor
每个 case 只保存一次图像条件和四个候选 block 的 Text K/V:
image_condition_latent [1,32,1,30,52] BF16
block_{00,01,52,53}_k_txt [1,16,903,128] BF16
block_{00,01,52,53}_v_txt [1,16,903,128] BF16
text_valid_mask [1,903] bool
Text K/V 在磁盘中补齐到 903 tokens,真实长度由 text_valid_mask 指示。
6.2 Full Teacher Text K/V
text_kv_exact_all54_bf16/ 是离线数据集的一部分,不是 pre-feature,也不是 rollout sidecar。每个
case 一个文件:
case_XXXX_text_kv_exact_all54_bf16.safetensors
block_{00..53}_k_txt [1,16,903,128] BF16
block_{00..53}_v_txt [1,16,903,128] BF16
text_valid_mask [1,903] bool
真实有效长度为 739–751 tokens,训练读取时按 text_valid_mask 去掉补零区,再交给 Teacher/Predictor。
构建会将 blocks 0、1、52、53 与原 case tensor 逐元素比对,确保它来自相同的 frozen text prefill。
6.3 Step tensor
每个 chunk 保存共享动作/相机/RoPE 信息,以及四个去噪步的监督目标:
action_labels [1,4] int64
target_viewmats [1,4,4,4] BF16
target_Ks [1,4,3,3] BF16
rope_temporal_size [1] int64
start_rope_start_idx [1] int64
step_{0..3}_timestep [1] FP32
step_{0..3}_noisy_sample [1,32,4,30,52] BF16
step_{0..3}_frame_condition [1,4,2048] BF16
step_{0..3}_final_hidden [1,6240,2048] BF16
step_{0..3}_velocity [1,32,4,30,52] BF16
65 通道 Predictor 输入不重复落盘。Dataset 使用 noisy sample、case-level image latent 和 I2V mask
重建 [B,65,4,30,52] 输入。
6.4 Context pre-feature
每个 record、每个 block 保存 K/V 线性投影之前的 img_modulated:
img_modulated [1,S_context,2048] BF16
context_valid_mask [1,S_context] bool
selected_frame_indices [context_frames] int64
context_viewmats [1,context_frames,4,4] BF16
context_Ks [1,context_frames,3,3] BF16
rope_temporal_size [1] int64
start_rope_start_idx [1] int64
S_context = context_frames × 30 × 52。八个 chunks 的 Context frames 依次为
0,4,8,12,16,20,20,20。磁盘文件保持变长且不 padding;v4 collate 按 batch 最大 Context 长度
动态 padding,并用 valid mask 屏蔽补齐部分。
训练时由冻结的 Teacher img_attn_k、img_attn_v、img_attn_k_norm 和相同的 RoPE/ProPE 路径
重建 Context Vision K/V。直接 K/V 与重建 K/V 的验收阈值为 relative L2 ≤ 5e-3、cosine ≥ 0.9999。
7. Manifest 关系
主键是 (case_id, action_id, chunk_id)。每条记录包含 case、step、四个 Context 文件的相对路径,
以及 caption、源图、动作、seed、分辨率、Teacher checkpoint 和 Context frame 数。
合并 manifest 时:
- 检查 3200 个主键唯一;
- 验证所有引用文件存在;
- chunk 1–7 增加
previous_chunk_id和previous_step_tensor_file; manifest.jsonl保留全部 3200 条;train_eval_manifest.jsonl保留 chunk 1–7,共 2800 条。
8. 构建、恢复和持久化
NVMe 是构建和训练读取位置,S3 是持久化 source of truth:
NVMe:
/mnt/local_nvme/zoubin/cz/hyworldplay_predictor_prefeature_v4
S3:
s3://s3-us-west2-default/zoubin/cz/projects/
HY-WorldPlay-DEV-Predictor/datasets/hyworldplay_predictor_prefeature_v4
项目挂载路径:
/mnt/s3files/s3-us-west2-default/zoubin/cz/projects/
HY-WorldPlay-DEV-Predictor/datasets/hyworldplay_predictor_prefeature_v4
完整构建流程:
python tools/prepare_predictor_prefeature_cases.py
bash tools/launch_predictor_prefeature_build.sh
第二条命令使用 8 个 GPU worker,先写 NVMe,随后合并 manifest 并同步到 S3。writer 支持按
case/action/chunk 跳过完整文件,任务中断后可重新执行。
全 54 层精确 Text K/V 单独构建并持久化到同一数据集:
bash tools/launch_predictor_text_kv_all54_build.sh
新机器训练前执行:
bash tools/ensure_predictor_prefeature_nvme.sh
需要全层 Text K/V 的 rollout 训练再执行:
bash tools/ensure_predictor_text_kv_all54_nvme.sh
如果 NVMe 不完整,该脚本使用 aws s3 sync 从持久前缀恢复并重新检查数量。手动将合法更新同步回
S3 使用:
bash tools/sync_predictor_prefeature_to_s3.sh
不要把 NVMe 当作唯一副本,也不要在未核对目标前缀时使用 aws s3 sync --delete。
9. 完整性验收
python tools/validate_predictor_prefeature_dataset.py
验收条件:
- 100 case tensors;
- 3200 step tensors;
- 12800 Context pre-feature tensors;
- 400 个 125 帧、832×480 视频;
- 3200 条无重复 manifest records 和 2800 条训练 records;
- 所有 tensor shape、dtype、finite、mask 和 Context frame 数正确;
- 四个 block 的 selected frame indices 一致;
- frozen final layer 能由 final hidden 与 frame condition 复现 velocity;
- 抽样重建的 65 通道输入与 Teacher 实际输入一致。
10. Validation/Test holdout
独立评估集为 hyworldplay_predictor_vbench_val25_test50:
- VBench metadata 共 355 个 pair;
- 排除训练集 100 个 pair 后剩 255 个;
- 使用 seed
20260718随机选择 75 个; - validation 25 个、test 50 个,两者互斥且均与训练集互斥;
- 每个 pair 生成四种动作;
- 已保存 300 个 Full-DiT 和 300 个 Reuse 视频;
- Full 为 steps
[0,1,2,3],Reuse 为 Full[0,3]、复用[1,2]。
入口:
bash tools/launch_vbench_holdout_full_reuse.sh
持久目录:
datasets/hyworldplay_predictor_vbench_val25_test50