I made an enhanced fork of this space (ACE-Step 1.5 XL Gradio Playground)

#21
by Blursed - opened

 

ACE-Step v1.5 Playground ๐Ÿ’ก ** ๐ŸŽถ ENHANCED ๐Ÿ—ฟ **

Try it out and let me know of any issues that you find with it.


Change Log (changes from this HuggingFace Spaces repo):

Summary

This fork transforms the stock ACE-Step v1.5 Gradio Space into a substantially more capable music-generation workstation. It upgrades the language model to the 4B variant, adds a community SFT+Turbo merge model (with automatic assembly from raw safetensors), introduces two dedicated audio-captioning backends (ACE-Captioner and MOSS-Music-8B-Thinking) for the Analyze workflow, adds a pluggable sampler/scheduler system with 12 new samplers, exposes CFG and Adaptive Projected Guidance controls in the UI, and implements broadcast-standard loudness normalization with a new lossless 32-bit float WAV output format. It also hardens the stack for modern GPUs (Blackwell-safe attention selection) and HuggingFace ZeroGPU deployment.

Changes

๐Ÿ†• New Modules (3 new files)

  • acestep/flow_samplers.py -- Pluggable flow-matching sampler system (656 lines):
    • Adds 12 custom samplers on top of the 3 native ones (Euler/ODE, Euler/SDE, Heun/ODE): euler_ancestral, dpmpp_2m, dpmpp_2m_sde, dpmpp_2m_sde_heun, dpmpp_3m_sde, er_sde, res_multistep, uni_pc, uni_pc_bh2, lms, jkass_quality, and jkass_fast.
    • Custom samplers are implemented against the model's public building blocks using rectified-flow-consistent math (k-diffusion-style updates with RF ancestral noise injection).
    • jkass_quality/jkass_fast are ports of the MIT-licensed JKASS ComfyUI samplers by jeankassio.
  • acestep/captioner.py -- ACE-Step audio captioner (ACE-Step/acestep-captioner, a Qwen2.5-Omni-7B finetune):
    • Analyzes raw audio directly (no 5Hz codes needed) and produces a descriptive ACE-Step 1.5 prompt plus a [BPM]bpm/[Key]/[TimeSig] metadata tail.
    • Lazily loaded on Analyze click and unloaded after each run to conserve GPU memory.
  • acestep/moss_captioner.py -- MOSS-Music captioner (OpenMOSS-Team/MOSS-Music-8B-Thinking):
    • Qwen3-8B-based music-understanding model with a dedicated 12.5 Hz audio encoder; chain-of-thought (<think> block) is stripped from output.
    • Shares the same analyze prompt and metadata parser as the ACE captioner.
    • Runs in full bf16 by default, streamed to GPU via device_map; quantization override available via ACESTEP_MOSS_QUANT (4bit/8bit/none).

๐ŸŽ›๏ธ Models

  • LM upgraded from 1.7B to 4B (acestep-5Hz-lm-4B) as the default, with dedicated download logic (the 4B lives in its own HF repo, unlike 0.6B/1.7B).
  • Community merge model added: acestep-v15-xl-merge-sft (jeankassio's SFT+Turbo task-arithmetic 0.5 merge) is now the default second DiT model.
    • New merge-model assembly pipeline: downloads config/modeling scaffold from the SFT donor repo, fetches the raw merged weights (~19.9 GB F32), renames to model.safetensors, and removes stale shard indexes so transformers loads it cleanly.
  • Stock SFT model (acestep-v15-xl-sft) removed from the default lineup due to consistently poor output quality (re-enable via SERVICE_MODE_DIT_MODEL_2/3 env vars if desired).
  • Optional third DiT model slot (SERVICE_MODE_DIT_MODEL_3) with shared VAE/text-encoder/silence-latent to minimize memory overhead.
  • Turbo-detection heuristic fixed: names containing sft, base, or merge are now correctly classified as full CFG-capable models.

๐ŸŽš๏ธ Generation Controls (UI)

  • CFG (Guidance Scale) slider now visible in Advanced Settings (was hidden at a fixed 7.0): default 4.0 for SFT/merge, locked at 1.0 and disabled for Turbo.
  • New APG toggle -- Adaptive Projected Guidance (the SFT model's native guidance) can be switched off for plain CFG; implemented via a runtime patch of the model's apg_forward, covering both native and custom sampler paths. Disabled for Turbo.
  • Model-aware UI: switching DiT models live-adapts the steps slider (Turbo: max 20, default 8; SFT/merge: max 200, default 50), CFG slider, and APG checkbox.
  • Sampler dropdown replaces the old binary ode/sde choice, listing all 15 native + custom samplers with guidance on which pay off per model type.

๐Ÿ”Š Audio Output

  • EBU R 128 / ReplayGain 2.0 loudness normalization at -18 LKFS applied to every render (BS.1770 K-weighting via pyloudnorm, measured on the float tensor before encoding so all formats are treated identically), with graceful fallback if unavailable.
  • Peak safety limiter: attenuates (never boosts) to prevent hard clipping in mp3/flac/16-bit WAV.
  • Two new output formats: wav and wav32 (32-bit float WAV -- lossless, preserves >0 dBFS headroom and the exact -18 LKFS target with no peak limiting).
  • Fixed a FileNotFoundError where Gradio was handed a non-existent .wav32 path (the saver's normalized path is now used).
  • Progress message reports the actual chosen format instead of hardcoded "mp3".

๐Ÿ” Analyze Workflow (Cover/Repaint)

  • Single "Analyze" button replaced by two buttons: "Analyze with ACE-Captioner" and "Analyze with MOSS Audio" (with optional icons).
  • The built-in LM still runs for lyrics, duration, and language detection; the selected captioner then overrides the caption and BPM/key/time-signature with its direct-from-audio analysis, falling back to LM output on failure.
  • Captioners load on demand and unload after each run, so both can be used alternately without exhausting GPU memory; ZeroGPU duration extended to 300s for these operations.

๐Ÿ–ฅ๏ธ Hardware & Deployment Hardening

  • Blackwell-safe attention selection (both DiT handler and LM): flash-attn3 is now gated to Hopper (sm_90) and flash-attention-2 to Ampere/Ada/Hopper via compute-capability checks -- previously the model would load on Blackwell (sm_120) then crash on the first forward pass. Everything else falls back to PyTorch-native sdpa.
  • New ACESTEP_ATTN_IMPLEMENTATION env var to force a specific attention backend.
  • More robust attention fallback chains for the text encoder (sdpa -> eager).
  • pyloudnorm>=0.1.1 added to requirements.

๐ŸŽจ Cosmetic

  • App title updated to "๐ŸŽ›๏ธ ACE-Step v1.5 Playground ๐Ÿ’ก ** ๐ŸŽถ ENHANCED ๐Ÿ—ฟ **".

Conclusion

In total, the ENHANCED fork adds three new modules and meaningfully modifies ten files against the upstream v1.5 baseline. The net result is a fork that goes well beyond the stock Space: a larger and smarter LM, a curated model lineup centered on the Turbo and community-merge checkpoints, professional loudness-normalized output with a true lossless option, two complementary audio-analysis backends, and a ComfyUI-class sampler selection -- all while remaining deployable on HuggingFace ZeroGPU and resilient across GPU generations from Ampere to Blackwell. Every enhancement degrades gracefully (captioner failures fall back to the LM, missing loudness libs fall back to peak normalization, unsupported attention kernels fall back to SDPA), keeping the original workflow intact for users who don't touch the new features.

Here is my recommended starting point for using the acestep-v15-xl-merge-sft model:

DiT Inference Steps: 128
CFG: 2.2
Shift: 2.4
Sampler / Scheduler: dpmpp_2m_sde_heun
Audio Format: wav

If using Cover or Repaint, you might get better results with either captioner (ACE-Step Captioner of MOSS-Audio). Try out them both.

Blursed changed discussion title from I made an enhanced fork of the ACE-Step 1.5 XL Gradio Playground to I made an enhanced fork of the this space (ACE-Step 1.5 XL Gradio Playground)
Blursed changed discussion title from I made an enhanced fork of the this space (ACE-Step 1.5 XL Gradio Playground) to I made an enhanced fork of this space (ACE-Step 1.5 XL Gradio Playground)

Sign up or log in to comment