--- license: cc-by-nc-4.0 tags: - diffusion - flow-matching - music - lora - grpo - rl - diffrhythm base_model: ASLP-lab/DiffRhythm2 library_name: peft --- # DiffRhythm2 · MixGRPO LoRA (checkpoints, MAAP) LoRA adapter checkpoints from a **block-wise dense-reward GRPO** run on top of the [ASLP-lab/DiffRhythm2](https://huggingface.co/ASLP-lab/DiffRhythm2) continuous-time flow-matching music model. Released as **research artifacts for a negative-result study** on RL fine-tuning of long-form audio flow models. ## What is here - **`adapter_model.safetensors` / `adapter_config.json`** at the repo root — the final `ckpt_final_002000` (2000 optimization steps). Load directly via `PeftModel.from_pretrained("m-a-a-p/DiffRhythm2-MixGRPO-LoRA", …)`. - **`checkpoints/step_XXXXXX/`** — intermediate LoRA adapters at every 100 training steps (Step 100, 200, …, 1900) for anyone who wants to reproduce the reward trajectory or run ablations. - **`figures/w3030_trends.png`** — 3-panel trend plot for the final sliding window (`window=[30-30]`), showing reward, KL, and clip-fraction over the 1840 steps spent in that window. - **`config/config_used.py`** — the exact `mixgrpo_lora_diffrhythm2()` config section used for this run, for reproducibility. ## Training setup Adapter is inserted into the DiT attention layers of DiffRhythm2 (dim 2048, depth 16, 16 heads, mel_dim 64). | item | value | |---|---| | base model | `ASLP-lab/DiffRhythm2` | | optimizer | AdamW, `lr = 5e-6`, `weight_decay = 1e-4`, `eps = 1e-8` | | LoRA | `r = 32`, `alpha = 64` | | RL algorithm | Block-wise MixGRPO (sliding window over 4 blocks, 30 s each) | | group size | 12 rollouts per prompt | | PPO-style clip | `clip_range = 0.1` | | KL coefficient | `0.01` | | gradient accumulation | 3 | | inner epochs | 1 | | EMA | **disabled** (OOM mitigation) | | CUDA alloc | `expandable_segments:True` | | total steps | 2000 | | hardware | 1 × RTX PRO 6000 (96 GB), Slurm job 1738062 | **Rewards** are a 50/50 blend of the audiobox aesthetics score and an ACE-Step-style Audio Alignment Score (AAS), aggregated over each 30 s block window. ## Result: no measurable improvement Across 2000 steps in the final sliding window (`[30-30]`, 1840 steps, Steps 160→1999), reward exhibits a **statistically significant regression**: - early-half mean reward: **0.18669** (n = 920) - late-half mean reward: **0.18592** (n = 920) - Δ = **−7.7 × 10⁻⁴**, t ≈ **−26.7** - linear slope: **−8.9 × 10⁻⁴ per 1000 steps** - KL(π ∥ π_ref) simultaneously *decreases* from 0.018 to 0.013 The full reward range across all 1840 steps is only **0.006** (≈ 0.3 %), i.e. the same magnitude as the per-step reward noise. Hyperparameter sweeps in the accompanying paper (across `num_generations ∈ {4, 12, 16}`, `clip_range ∈ {1e-5, 1e-4, 0.1}`, `lr ∈ {5e-6, 1e-5, 5e-5}`) do not escape this plateau. Interpretation: the log-probability signal that GRPO needs — computed here through the surrogate ratio between the current and the pre-update policy on the ODE / SDE trajectory — is **noise-dominated** at the scale of the block- wise dense rewards used here, so the policy gradient does not consistently point in the reward-increasing direction. See the accompanying paper for the full analysis (including the diagnosis of the `1 / (1−t)` score-head singularity in the underlying flow-matching formulation). ## Intended use These adapters are released as **research artifacts** to accompany the negative-result study, not as a recommended production LoRA on top of DiffRhythm2. In particular: - Do not expect audio-quality gains over the base DiffRhythm2 model. - If you use them as a *baseline* or *sanity check* for a new RL algorithm, please cite the accompanying paper. ## Quick start ```python from peft import PeftModel from safetensors.torch import load_file # Final (2000-step) adapter, loaded via peft on top of your DiffRhythm2 DiT. adapter = load_file("adapter_model.safetensors") # Intermediate: # hf_hub_download("m-a-a-p/DiffRhythm2-MixGRPO-LoRA", # "checkpoints/step_001000/adapter_model.safetensors") ``` ## License Released under **CC BY-NC 4.0**. Non-commercial use only; academic use encouraged. Base model (`ASLP-lab/DiffRhythm2`) is subject to its own license. ## Citation If you use these checkpoints or the accompanying analysis, please cite the associated paper (details to be added when the paper appears).