VEFX-Code / README.md
xiangbog's picture
Align VEFX-Bench v1.0 protocol and metadata
befd6f9 verified
|
Raw
History Blame
9.82 kB
metadata
title: VEFX-Code
emoji: 🎬
colorFrom: indigo
colorTo: pink
sdk: static
pinned: false
license: apache-2.0
short_description: VEFX-Bench reference code & inference utils

VEFX-Bench v1.0

Benchmarking Generic Video Editing and Visual Effects

Frozen release: VEFX-Bench v1.0 β€” frozen on 2026-08-03. See the immutable dataset, manifest, and checkpoint identifiers under Reproducibility.

VEFX-Bench is a comprehensive benchmark for evaluating text-driven video editing and visual effects. It includes 5,049 annotated examples spanning 9 categories and 32 subcategories, with a frozen 300-item v1.0 test set evaluated by VEFX-Reward β€” a VLM-based reward model that scores edits across three dimensions on a 1–4 scale:

Dimension What it measures
Instructional Following (IF) Does the edit accurately reflect the editing instruction?
Render Quality (RQ) Visual clarity, temporal consistency, and physical plausibility
Edit Exclusivity (EE) Were only the intended regions modified, without side-effects?

πŸ† Model Leaderboard

VEFX-Reward scores each dimension on a 1–4 scale. Overall is the primary GeoAgg metric; it is not the arithmetic average of IF, RQ, and EE. For item j, define i_j=(IF_j-1)/3, r_j=(RQ_j-1)/3, and e_j=(EE_j-1)/3, then compute g_j=1+3*(i_j^2*r_j*e_j)^(1/4). The leaderboard computes Overall = (1/300)Β·Ξ£_j g_j: GeoAgg is computed per item first, then the 300 values are averaged. Missing or failed items receive (1,1,1). The arithmetic Mean is diagnostic only and is not used for ranking.

πŸ“… Updated: August 3, 2026 β€” For the latest results & submissions, visit the live leaderboard β†’

Rank Model Type IF ↑ RQ ↑ EE ↑ GeoAgg ↑
πŸ₯‡ Kling o3 Omni Commercial 3.033 3.588 3.043 3.057
πŸ₯ˆ Kling o1 Commercial 3.040 3.534 2.976 2.985
πŸ₯‰ Runway Gen-4.5 Commercial 2.817 3.319 2.923 2.912
4 Seedance 2.0 Commercial 2.811 3.421 3.088 2.766
5 Grok Imagine Commercial 2.606 3.346 3.376 2.723
6 Luma Ray 3 Commercial 2.702 3.403 2.705 2.717
7 UniVideo Open-source 2.294 3.266 3.091 2.516
8 Wan 2.6 Commercial 2.012 3.317 2.446 2.146
9 Luma Ray 2 Commercial 2.038 2.532 1.363 1.804
10 VACE Open-source 2.027 3.172 1.180 1.775

🎬 Demo Videos

Each demo shows the original video (left) alongside the edited video (right).

Attribute Change
"Change the color of the red industrial trailer to a bright yellow while maintaining the texture and appearance of the metal surface."
Object Removal
"Remove the woman with the grey backpack walking on the right side of the frame."
Style Transfer
"Restore the natural, realistic colors to the entire scene, replacing the current black and white style with a full-color rendition."
Camera Motion
"Perform a smooth zoom in on the distant snowy mountain peaks to create a more immersive view."

πŸ“Š Benchmark at a Glance

πŸ“ 5,049 Annotated Examples 🎬 1,419 Source Videos
πŸ“‚ 9 / 32 Categories / Subcategories πŸ€– 10 Editing Systems
πŸ“ 3 Quality Dimensions (IF, RQ, EE) πŸ§ͺ 300 Benchmark Test Pairs

πŸ€— VEFX-Reward Models

Model Backbone Params HuggingFace Status
VEFX-Reward-4B Qwen3-VL-4B-Instruct 4B VEFX-Reward/VEFX-Reward-4B βœ… Available

πŸ“¦ VEFX-Bench Dataset

The canonical frozen benchmark dataset is hosted on Hugging Face at xiangbog/VEFX-Bench. Submit and inspect results on the public VEFX-Leaderboard.

🎬 300 Source Videos (720p) πŸ“ prompts.json with editing instructions
πŸ“‚ 9 Task Categories πŸ—‚οΈ benchmark_meta.json with category labels

Task Categories: Attribute Editing Β· Camera Angle Editing Β· Camera Motion Editing Β· Creative Edit Β· Instance Editing Β· Instance Motion Editing Β· Quantity Editing Β· Style Editing Β· Visual Effect Editing

πŸ”’ Reproducibility

Artifact Frozen identifier
Dataset revision 3bf997e7eb4fa0d0c2d56cce5ddfccc1dfbda235
benchmark_meta.json SHA-256 277d89f5cf23af4163fe6ee120654210f51d19c6dee2d3aa76a3248471d1f333
Reward model revision a15a8dbe1b3eb07ee0919e8de059f170436ec9ff
model.safetensors SHA-256 c3c0d03f770f0a73631206821922213de75413bb517f7d7a2fd9ab1f2c38f59d

Download and Evaluate

from huggingface_hub import snapshot_download

# Download the benchmark dataset
snapshot_download(
    repo_id="xiangbog/VEFX-Bench",
    repo_type="dataset",
    revision="3bf997e7eb4fa0d0c2d56cce5ddfccc1dfbda235",
    local_dir="./vefx_bench",
)

Evaluation workflow:

  1. Download the 300 source videos and prompts.json
  2. Apply your video editing model to each source video following its prompt
  3. Save edited videos as 0000.mp4 through 0299.mp4 (matching source index)
  4. Score with VEFX-Reward:
import json
from vefx_reward import VEFXReward

model = VEFXReward("VEFX-Reward/VEFX-Reward-4B", device="cuda")

with open("vefx_bench/prompts.json") as f:
    prompts = json.load(f)

for idx, item in enumerate(prompts):
    scores = model.score(
        original_video=f"vefx_bench/{idx:04d}.mp4",
        edited_video=f"your_edits/{idx:04d}.mp4",
        instruction=item["instruction"],
    )
    print(f"[{idx:04d}] IF={scores['IF']:.2f}  RQ={scores['RQ']:.2f}  EE={scores['EE']:.2f}")

πŸš€ Quick Start

Installation

conda create -n vefx-bench python=3.10 -y
conda activate vefx-bench

# Install PyTorch first (match your CUDA version)
# See https://pytorch.org/get-started/locally/ for the right command
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu124

# Install remaining dependencies
pip install -r requirements.txt

# Install the package
pip install -e .

Requirements: Python β‰₯ 3.10, CUDA GPU, ~10 GB VRAM (bfloat16). Make sure your PyTorch CUDA version matches your driver.

Score a Video Edit (Python API)

from vefx_reward import VEFXReward

model = VEFXReward("VEFX-Reward/VEFX-Reward-4B", device="cuda")

scores = model.score(
    original_video="examples/sample_videos/object_removal_original.mp4",
    edited_video="examples/sample_videos/object_removal_edited.mp4",
    instruction="Remove the woman with the grey backpack walking on the right side of the frame.",
)
print(scores)
# {'IF': 2.34, 'RQ': 1.93, 'EE': 1.82, 'Overall': 2.082, 'GeoAgg': 2.082, 'Mean': 2.03}

CLI Usage

python examples/quick_start.py \
    --original examples/sample_videos/object_removal_original.mp4 \
    --edited examples/sample_videos/object_removal_edited.mp4 \
    --instruction "Remove the woman with the grey backpack walking on the right side of the frame."

Score All Included Samples

The repo includes 4 sample video pairs with prompts. Score them all:

import json
from vefx_reward import VEFXReward

model = VEFXReward("VEFX-Reward/VEFX-Reward-4B", device="cuda")

with open("examples/sample_videos/prompts.json") as f:
    samples = json.load(f)

for sample in samples:
    scores = model.score(
        original_video=f"examples/sample_videos/{sample['original']}",
        edited_video=f"examples/sample_videos/{sample['edited']}",
        instruction=sample["instruction"],
    )
    print(f"[{sample['category']}] IF={scores['IF']:.2f}  RQ={scores['RQ']:.2f}  EE={scores['EE']:.2f}")

Batch Scoring

Prepare a CSV with columns original_video, edited_video, instruction:

python examples/batch_scoring.py --csv edits.csv --output results.csv

Multi-GPU Scoring

For large-scale evaluation across multiple GPUs:

python examples/multi_gpu_scoring.py --csv edits.csv --num_gpus 4 --output results.csv

πŸ“– API Reference

VEFXReward

VEFXReward(
    model_path="VEFX-Reward/VEFX-Reward-4B",  # HuggingFace ID or local path
    device="cuda",                           # "cuda", "cuda:0", "cpu"
    dtype=torch.bfloat16,                    # torch.bfloat16 or torch.float16
    fps=4.0,                                 # Video sampling rate
    max_frame_pixels=399360,                 # Max pixels per frame
)

model.score(original_video, edited_video, instruction) β†’ dict

Score a single video edit. Returns {'IF': float, 'RQ': float, 'EE': float, 'Overall': float, 'GeoAgg': float, 'Mean': float}. Overall and GeoAgg are identical item-level v1.0 scores; Mean is diagnostic only.

model.score_batch(original_videos, edited_videos, instructions) β†’ list[dict]

Score multiple edits sequentially. Each sample is processed independently to avoid OOM.