Spaces:
Running on Zero
A newer version of the Gradio SDK is available: 6.26.0
title: JoyAI Video Edit
emoji: π¬
colorFrom: indigo
colorTo: gray
sdk: gradio
sdk_version: 6.22.0
app_file: app.py
short_description: Streaming instruction-guided video editing
python_version: '3.12'
startup_duration_timeout: 1h
license: apache-2.0
tags:
- video-editing
- video-to-video
- streaming
- diffusion
models:
- jdopensource/JoyAI-Video-Edit
- XiaomiMiMo/MiMo-VL-7B-RL-2508
JoyAI-Video-Edit β streaming video editing
Instruction-guided video editing with jdopensource/JoyAI-Video-Edit.
Describe an edit in plain language and it is applied to your clip chunk by chunk: a
16.3B chunk-causal MMDiT edits 8 frames at a time behind a rolling KV cache, so the
result streams back while it is still being generated.
How it works
| stage | model |
|---|---|
| condition encoder | XiaomiMiMo/MiMo-VL-7B-RL-2508 (instruction + anchor frame β 1024 conditioning tokens) |
| editor | 40-block chunk-causal MMDiT, 16.3B params, source-id RoPE, 2 distilled flow-matching steps per chunk |
| video codec | XVAEChunkCausal β 16Γ spatial / 8Γ temporal causal video VAE with a 1.5Γ refinement head |
Per 8-frame chunk: VAE-encode the source window β 2 DiT denoise passes over a
[chunk 0, chunk n-1, chunk n] attention window β one extra "clean KV" pass that writes
the finished chunk into the cache β VAE-decode (seeded with a pseudo-latent
re-encoded from the previous chunk's last frame) β evict stale cache entries.
Deviations from the reference deployment
Ported from jd-opensource/JoyAI-Video-Edit
(deploy/). The tensor pipeline β chunk windows, cache ids/eviction, temporal-id RoPE,
the pseudo-latent decode trick, scheduler shift, prompt template β matches the reference.
What changed, and why:
- FlashAttention-4 β SDPA. FA4/FA3 kernels do not exist for the sm_120 GPU this Space runs on; PyTorch SDPA (flash backend) is used with identical, unmasked attention.
sgl-kernelfused ops β pure PyTorch. The fused layernorm-modulate / QK-norm-RoPE / add-gate kernels are reimplemented with the same math.sgl-kernelFP8 GEMMs βtorchao. The reference already serves the DiT image stream in FP8 (JOYOMNI_FP8_IMG=1); here both streams use per-channel FP8 weights with per-token dynamic activation scaling so the whole system fits one 48 GB GPU slice. The MLLM condition encoder is int8 weight-only.torch.compiledisabled. JIT inductor compilation cannot run inside a ZeroGPU worker.- Threads β sequential. The reference runs five pipeline-parallel worker threads feeding a WebSocket downlink; a Space request gets one GPU slice, so the stages run in order and frames are streamed back per chunk instead.
- Face-restoration gating and the ONNX detector are not included.
Sources default to the reference server's settings: 720Γ1248 (or 1248Γ720 portrait), 24 fps, 2 steps, seed 42.
Showcase clips under examples/ are the source videos from the reference repository's
README, with the authors' prompts.