SyncReward

SyncReward scores the audio-visual synchronization of a generated video with sound. It returns a reward in [0, 2] (0 = not synchronized, 2 = well synchronized) and is trained to match human sync judgments on clips from text-to-audio-video generators.

Paper, code repository and citation: coming soon.

Models

This repo holds two checkpoints, one per subfolder. Each has model.safetensors (fp32), config.json, SHA256SUMS and its own model card.

Folder What it is Params
syncreward-stage2/ SyncReward, the full reward model (recommended) 1,658.9 M
syncreward-stage1/ Stage-1 PE-AV audio/video encoder; frozen inside syncreward-stage2 1,635.4 M
  • Stage 1 fine-tunes the facebook/pe-av-large audio and video towers with segment-level contrastive learning on real clips (AudioSet, then filtered VGGSound).
  • Stage 2 freezes that encoder and trains a 3-layer cross-modal Transformer over 18 segments of 0.5 s ([REW; V_1..18; MOD; A_1..18] -> MLP -> 2 * sigmoid) on human ratings of generated clips. The Stage-2 encoder is bit-identical to syncreward-stage1.

Results

SyncReward-Bench: 1,000 generated clips from 7 generators, target = mean of up to 3 human ratings. Metric seed 20260820, 95% bootstrap CIs over 1,000 resamples.

Model Spearman Top-1 (N=4) Pair Acc
syncreward-stage2 (SyncReward) 0.7823 [0.757, 0.806] 0.7646 [0.714, 0.819] 0.8359 [0.823, 0.848]
syncreward-stage1 (Stage 1 only) 0.6051 0.6299 0.7455
  • Top-1 (N=4) is best-of-4 selection accuracy, averaged over 1,000 random groupings.
  • Pair Acc covers all clip pairs with different targets; prediction ties count 0.5.
  • The Stage-2 numbers were reproduced with the released code from these exact weights. All 1,000 per-clip predictions are identical to the paper's.

Download and load

hf download Guan123/SyncReward --include 'syncreward-stage2/*' --local-dir checkpoints
(cd checkpoints/syncreward-stage2 && sha256sum -c SHA256SUMS)

Loading uses the SyncReward code (repository link coming soon):

from syncreward.model import build_reward_model
model = build_reward_model("facebook/pe-av-large", checkpoint="checkpoints/syncreward-stage2").cuda().eval()
python score.py --checkpoint checkpoints/syncreward-stage2 my_video.mp4

facebook/pe-av-large provides only the processor config and module definitions; every weight comes from model.safetensors. Inference uses a centered 4.75 s window (18 segments of 0.5 s, stride 0.25 s), with audio at 48 kHz.

Training data

SyncReward-Data-80K contains 80,001 human synchronization ratings from 5 annotators over 57,099 clips generated by 8 text-to-audio-video models (about 10k ratings per model).

Each rating first marks whether the audio is related to the visible scene at all. If it is, the annotator scores synchronization as 0 (out of sync), 1 (partially synchronized) or 2 (well synchronized). We keep unrelated audio in the data on purpose: it is a common failure mode of current generators, and a sync reward should penalize it rather than never see it.

Rating Count Share
Unrelated audio 10,065 12.6%
0 โ€“ out of sync 12,907 16.1%
1 โ€“ partial 37,935 47.4%
2 โ€“ well synchronized 19,094 23.9%

The Stage-2 target of a clip is the rater-bias-corrected mean of its related ratings. Clips that every rater marked unrelated (6,261 train and 172 validation clips, 11.5% of the training pool) get target 0, the same as a clear sync failure. The 1,931 clips with mixed votes use their related ratings only.

Splits are clip-disjoint by (prompt, generator): 54,463 training and 1,636 validation clips. The 1,000 SyncReward-Bench clips are held out from training.

Limitations

SyncReward scores only the temporal alignment of on-screen sound sources inside a 4.75 s window. It does not judge audio quality or semantic relevance. Its training data comes from 8 generators.

License

Apache-2.0 (see LICENSE). The PE-AV backbone facebook/pe-av-large is also Apache-2.0.

Citation

Coming soon.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Guan123/SyncReward

Finetuned
(1)
this model