Text-to-Video

AV-GRPO

AV-GRPO is a modality-anchored online diffusion reinforcement learning framework for joint audio-video generation. It enables full-parameter or LoRA training of the 22B LTX-2.3 model on 8 A800 GPUs.

This repository contains the model introduced in AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation.

For training, inference, and usage details, please refer to the GitHub repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using Dr-Loser/AV-GRPO 1

Paper for Dr-Loser/AV-GRPO