Spaces:
Running on Zero
I made an enhanced fork of this space (ACE-Step 1.5 XL Gradio Playground)
ACE-Step v1.5 Playground ๐ก ** ๐ถ ENHANCED ๐ฟ **
Try it out and let me know of any issues that you find with it.
Change Log (changes from this HuggingFace Spaces repo):
Summary
This fork transforms the stock ACE-Step v1.5 Gradio Space into a substantially more capable music-generation workstation. It upgrades the language model to the 4B variant, adds a community SFT+Turbo merge model (with automatic assembly from raw safetensors), introduces two dedicated audio-captioning backends (ACE-Captioner and MOSS-Music-8B-Thinking) for the Analyze workflow, adds a pluggable sampler/scheduler system with 12 new samplers, exposes CFG and Adaptive Projected Guidance controls in the UI, and implements broadcast-standard loudness normalization with a new lossless 32-bit float WAV output format. It also hardens the stack for modern GPUs (Blackwell-safe attention selection) and HuggingFace ZeroGPU deployment.
Changes
๐ New Modules (3 new files)
acestep/flow_samplers.py-- Pluggable flow-matching sampler system (656 lines):- Adds 12 custom samplers on top of the 3 native ones (Euler/ODE, Euler/SDE, Heun/ODE):
euler_ancestral,dpmpp_2m,dpmpp_2m_sde,dpmpp_2m_sde_heun,dpmpp_3m_sde,er_sde,res_multistep,uni_pc,uni_pc_bh2,lms,jkass_quality, andjkass_fast. - Custom samplers are implemented against the model's public building blocks using rectified-flow-consistent math (k-diffusion-style updates with RF ancestral noise injection).
jkass_quality/jkass_fastare ports of the MIT-licensed JKASS ComfyUI samplers by jeankassio.
- Adds 12 custom samplers on top of the 3 native ones (Euler/ODE, Euler/SDE, Heun/ODE):
acestep/captioner.py-- ACE-Step audio captioner (ACE-Step/acestep-captioner, a Qwen2.5-Omni-7B finetune):- Analyzes raw audio directly (no 5Hz codes needed) and produces a descriptive ACE-Step 1.5 prompt plus a
[BPM]bpm/[Key]/[TimeSig]metadata tail. - Lazily loaded on Analyze click and unloaded after each run to conserve GPU memory.
- Analyzes raw audio directly (no 5Hz codes needed) and produces a descriptive ACE-Step 1.5 prompt plus a
acestep/moss_captioner.py-- MOSS-Music captioner (OpenMOSS-Team/MOSS-Music-8B-Thinking):- Qwen3-8B-based music-understanding model with a dedicated 12.5 Hz audio encoder; chain-of-thought (
<think>block) is stripped from output. - Shares the same analyze prompt and metadata parser as the ACE captioner.
- Runs in full bf16 by default, streamed to GPU via
device_map; quantization override available viaACESTEP_MOSS_QUANT(4bit/8bit/none).
- Qwen3-8B-based music-understanding model with a dedicated 12.5 Hz audio encoder; chain-of-thought (
๐๏ธ Models
- LM upgraded from 1.7B to 4B (
acestep-5Hz-lm-4B) as the default, with dedicated download logic (the 4B lives in its own HF repo, unlike 0.6B/1.7B). - Community merge model added:
acestep-v15-xl-merge-sft(jeankassio's SFT+Turbo task-arithmetic 0.5 merge) is now the default second DiT model.- New merge-model assembly pipeline: downloads config/modeling scaffold from the SFT donor repo, fetches the raw merged weights (~19.9 GB F32), renames to
model.safetensors, and removes stale shard indexes sotransformersloads it cleanly.
- New merge-model assembly pipeline: downloads config/modeling scaffold from the SFT donor repo, fetches the raw merged weights (~19.9 GB F32), renames to
- Stock SFT model (
acestep-v15-xl-sft) removed from the default lineup due to consistently poor output quality (re-enable viaSERVICE_MODE_DIT_MODEL_2/3env vars if desired). - Optional third DiT model slot (
SERVICE_MODE_DIT_MODEL_3) with shared VAE/text-encoder/silence-latent to minimize memory overhead. - Turbo-detection heuristic fixed: names containing
sft,base, ormergeare now correctly classified as full CFG-capable models.
๐๏ธ Generation Controls (UI)
- CFG (Guidance Scale) slider now visible in Advanced Settings (was hidden at a fixed 7.0): default 4.0 for SFT/merge, locked at 1.0 and disabled for Turbo.
- New APG toggle -- Adaptive Projected Guidance (the SFT model's native guidance) can be switched off for plain CFG; implemented via a runtime patch of the model's
apg_forward, covering both native and custom sampler paths. Disabled for Turbo. - Model-aware UI: switching DiT models live-adapts the steps slider (Turbo: max 20, default 8; SFT/merge: max 200, default 50), CFG slider, and APG checkbox.
- Sampler dropdown replaces the old binary ode/sde choice, listing all 15 native + custom samplers with guidance on which pay off per model type.
๐ Audio Output
- EBU R 128 / ReplayGain 2.0 loudness normalization at -18 LKFS applied to every render (BS.1770 K-weighting via
pyloudnorm, measured on the float tensor before encoding so all formats are treated identically), with graceful fallback if unavailable. - Peak safety limiter: attenuates (never boosts) to prevent hard clipping in mp3/flac/16-bit WAV.
- Two new output formats:
wavandwav32(32-bit float WAV -- lossless, preserves >0 dBFS headroom and the exact -18 LKFS target with no peak limiting). - Fixed a
FileNotFoundErrorwhere Gradio was handed a non-existent.wav32path (the saver's normalized path is now used). - Progress message reports the actual chosen format instead of hardcoded "mp3".
๐ Analyze Workflow (Cover/Repaint)
- Single "Analyze" button replaced by two buttons: "Analyze with ACE-Captioner" and "Analyze with MOSS Audio" (with optional icons).
- The built-in LM still runs for lyrics, duration, and language detection; the selected captioner then overrides the caption and BPM/key/time-signature with its direct-from-audio analysis, falling back to LM output on failure.
- Captioners load on demand and unload after each run, so both can be used alternately without exhausting GPU memory; ZeroGPU duration extended to 300s for these operations.
๐ฅ๏ธ Hardware & Deployment Hardening
- Blackwell-safe attention selection (both DiT handler and LM): flash-attn3 is now gated to Hopper (sm_90) and flash-attention-2 to Ampere/Ada/Hopper via compute-capability checks -- previously the model would load on Blackwell (sm_120) then crash on the first forward pass. Everything else falls back to PyTorch-native
sdpa. - New
ACESTEP_ATTN_IMPLEMENTATIONenv var to force a specific attention backend. - More robust attention fallback chains for the text encoder (
sdpa->eager). pyloudnorm>=0.1.1added to requirements.
๐จ Cosmetic
- App title updated to "๐๏ธ ACE-Step v1.5 Playground ๐ก ** ๐ถ ENHANCED ๐ฟ **".
Conclusion
In total, the ENHANCED fork adds three new modules and meaningfully modifies ten files against the upstream v1.5 baseline. The net result is a fork that goes well beyond the stock Space: a larger and smarter LM, a curated model lineup centered on the Turbo and community-merge checkpoints, professional loudness-normalized output with a true lossless option, two complementary audio-analysis backends, and a ComfyUI-class sampler selection -- all while remaining deployable on HuggingFace ZeroGPU and resilient across GPU generations from Ampere to Blackwell. Every enhancement degrades gracefully (captioner failures fall back to the LM, missing loudness libs fall back to peak normalization, unsupported attention kernels fall back to SDPA), keeping the original workflow intact for users who don't touch the new features.
Here is my recommended starting point for using the acestep-v15-xl-merge-sft model:
DiT Inference Steps: 128
CFG: 2.2
Shift: 2.4
Sampler / Scheduler: dpmpp_2m_sde_heun
Audio Format: wav
If using Cover or Repaint, you might get better results with either captioner (ACE-Step Captioner of MOSS-Audio). Try out them both.