SyncReward
SyncReward scores the audio-visual synchronization of a generated video with sound. It returns a reward in [0, 2] (0 = not synchronized, 2 = well synchronized) and is trained to match human sync judgments on clips from text-to-audio-video generators.
Paper, code repository and citation: coming soon.
Models
This repo holds two checkpoints, one per subfolder. Each has model.safetensors (fp32), config.json,
SHA256SUMS and its own model card.
| Folder | What it is | Params |
|---|---|---|
syncreward-stage2/ |
SyncReward, the full reward model (recommended) | 1,658.9 M |
syncreward-stage1/ |
Stage-1 PE-AV audio/video encoder; frozen inside syncreward-stage2 |
1,635.4 M |
- Stage 1 fine-tunes the
facebook/pe-av-largeaudio and video towers with segment-level contrastive learning on real clips (AudioSet, then filtered VGGSound). - Stage 2 freezes that encoder and trains a 3-layer cross-modal Transformer over 18 segments of 0.5 s
(
[REW; V_1..18; MOD; A_1..18]-> MLP ->2 * sigmoid) on human ratings of generated clips. The Stage-2 encoder is bit-identical tosyncreward-stage1.
Results
SyncReward-Bench: 1,000 generated clips from 7 generators, target = mean of up to 3 human ratings. Metric seed 20260820, 95% bootstrap CIs over 1,000 resamples.
| Model | Spearman | Top-1 (N=4) | Pair Acc |
|---|---|---|---|
syncreward-stage2 (SyncReward) |
0.7823 [0.757, 0.806] | 0.7646 [0.714, 0.819] | 0.8359 [0.823, 0.848] |
syncreward-stage1 (Stage 1 only) |
0.6051 | 0.6299 | 0.7455 |
- Top-1 (N=4) is best-of-4 selection accuracy, averaged over 1,000 random groupings.
- Pair Acc covers all clip pairs with different targets; prediction ties count 0.5.
- The Stage-2 numbers were reproduced with the released code from these exact weights. All 1,000 per-clip predictions are identical to the paper's.
Download and load
hf download Guan123/SyncReward --include 'syncreward-stage2/*' --local-dir checkpoints
(cd checkpoints/syncreward-stage2 && sha256sum -c SHA256SUMS)
Loading uses the SyncReward code (repository link coming soon):
from syncreward.model import build_reward_model
model = build_reward_model("facebook/pe-av-large", checkpoint="checkpoints/syncreward-stage2").cuda().eval()
python score.py --checkpoint checkpoints/syncreward-stage2 my_video.mp4
facebook/pe-av-large provides only the processor config and module definitions; every weight comes from
model.safetensors. Inference uses a centered 4.75 s window (18 segments of 0.5 s, stride 0.25 s),
with audio at 48 kHz.
Training data
SyncReward-Data-80K contains 80,001 human synchronization ratings from 5 annotators over 57,099 clips generated by 8 text-to-audio-video models (about 10k ratings per model).
Each rating first marks whether the audio is related to the visible scene at all. If it is, the annotator scores synchronization as 0 (out of sync), 1 (partially synchronized) or 2 (well synchronized). We keep unrelated audio in the data on purpose: it is a common failure mode of current generators, and a sync reward should penalize it rather than never see it.
| Rating | Count | Share |
|---|---|---|
| Unrelated audio | 10,065 | 12.6% |
| 0 โ out of sync | 12,907 | 16.1% |
| 1 โ partial | 37,935 | 47.4% |
| 2 โ well synchronized | 19,094 | 23.9% |
The Stage-2 target of a clip is the rater-bias-corrected mean of its related ratings. Clips that every rater marked unrelated (6,261 train and 172 validation clips, 11.5% of the training pool) get target 0, the same as a clear sync failure. The 1,931 clips with mixed votes use their related ratings only.
Splits are clip-disjoint by (prompt, generator): 54,463 training and 1,636 validation clips. The 1,000 SyncReward-Bench clips are held out from training.
Limitations
SyncReward scores only the temporal alignment of on-screen sound sources inside a 4.75 s window. It does not judge audio quality or semantic relevance. Its training data comes from 8 generators.
License
Apache-2.0 (see LICENSE). The PE-AV backbone facebook/pe-av-large is also Apache-2.0.
Citation
Coming soon.
Model tree for Guan123/SyncReward
Base model
facebook/pe-av-large