Spaces:
Sleeping
Sleeping
| # Pipeline timing and realtime suitability | |
| Last updated: 2026-05-12 | |
| ## Measurement scope | |
| The timing benchmark in `docs/benchmark-subprocesses.json` measures synthetic drum fixtures with: | |
| - `stem=all`, so Demucs is bypassed. | |
| - Warm process after import/model-library initialization. | |
| - Synthetic rock/funk/halftime fixtures generated by `synth_generator.py`. | |
| - `scripts/benchmark_subprocesses.py` as the benchmark driver. | |
| This isolates the sample-extraction subprocesses from source-separation noise. Demucs timing depends heavily on model, hardware, track length, first-run downloads, and CPU/GPU availability, so it is analyzed separately. | |
| ## Significant subprocesses | |
| | Subprocess | Current implementation | Timing behavior | Realtime suitability | | |
| |---|---|---|---| | |
| | Source load / stem extraction | `extract_stem`; full mix via `librosa`, stems via Demucs | Full mix is usually small; Demucs dominates full jobs | Full mix: near-realtime. Demucs: no. | | |
| | BPM detection | `detect_bpm` using onset envelope and beat tracking | Usually sub-second for short fixtures | Near-realtime with buffering; not critical path. | | |
| | Onset detection + slicing | `detect_onsets` multi-band SuperFlux-style envelope | Often the largest pure-DSP stage | Near-realtime with bounded lookahead. | | |
| | Classification | Rule-based spectral analysis per hit | Fast relative to onset/clustering | Near-realtime. | | |
| | Batch clustering | Mel fingerprints + transient NCC + agglomerative clustering | Pairwise/batch; scales poorly with many hits | Not realtime. Final-quality batch mode. | | |
| | Online clustering | Prototype assignment per hit | Scales with hit count × cluster count | Near-realtime preview path. | | |
| | Representative selection | Scores each candidate hit | Moderate for many clusters/hits | Near-realtime for moderate hit counts. | | |
| | Synthesis | Weighted aligned average per multi-hit cluster | Usually small | Near-realtime for moderate clusters. | | |
| | Export/package | WAV/MIDI/render/ZIP writes | Disk-bound; ZIP is batch finalization | Not meaningful as realtime; finalization step. | | |
| ## Current benchmark summary | |
| The checked-in benchmark files were refreshed on 2026-05-12 with synthetic 2-bar fixtures and Demucs bypassed: | |
| - `docs/benchmark-subprocesses.json`: `batch_quality` clustering. | |
| - `docs/benchmark-online-preview.json`: `online_preview` clustering. | |
| | Stage | Batch quality mean | Online preview mean | | |
| |---|---:|---:| | |
| | source load | 0.010 s | 0.010 s | | |
| | BPM detection | 0.155 s | 0.126 s | | |
| | onset detection + slicing | 1.964 s | 1.763 s | | |
| | classification | 0.042 s | 0.041 s | | |
| | clustering | 0.046 s | 0.037 s | | |
| | representative selection | 0.177 s | 0.158 s | | |
| | synthesis | 0.001 s | 0.001 s | | |
| | export/package | 0.158 s | 0.291 s | | |
| On these small fixtures, `online_preview` reduced clustering time compared with `batch_quality`, while export time increased because this pass now writes every accepted hit as a review WAV under `review/hits/`. The total run is still dominated by onset detection, so the next realtime optimization target is streaming/incremental onset analysis rather than only clustering. | |
| First cold runs can be much slower because imports and library initialization are paid up front. | |
| ## Batch quality versus online preview clustering | |
| ### `batch_quality` | |
| Current final-quality clustering path: | |
| 1. Compute mel fingerprint for each hit. | |
| 2. Compute pairwise mel cosine prefilter. | |
| 3. Compute transient NCC only for candidate pairs. | |
| 4. Build distance matrix. | |
| 5. Run agglomerative clustering. | |
| 6. Optionally merge singleton clusters. | |
| This gives better global grouping, but it is fundamentally batch-oriented because it wants the full similarity matrix. | |
| ### `online_preview` | |
| Current near-realtime-oriented clustering path: | |
| 1. Process hits in onset order. | |
| 2. Compute one mel fingerprint and one transient per hit. | |
| 3. Compare the new hit against existing cluster prototypes. | |
| 4. Assign it to the best prototype or create a new cluster until the target cap is reached. | |
| 5. Update prototype fingerprints/transients using energy-weighted rolling averages. | |
| Complexity is roughly `O(number_of_hits × number_of_clusters)`, not `O(number_of_hits²)`, and does not require future hits before producing a current assignment. It is suitable for progressive preview and fast iteration, but it is not guaranteed to match the global batch clustering result. | |
| ## What can run in or near realtime | |
| These can be performed progressively with small buffers: | |
| - Source decode for already-separated/full-mix audio. | |
| - Onset envelope computation. | |
| - Peak picking with bounded lookahead. | |
| - Hit slicing once enough tail audio is buffered. | |
| - Rule-based classification. | |
| - Mel fingerprint extraction. | |
| - Online prototype clustering. | |
| - Representative preview selection. | |
| - Basic reconstruction preview. | |
| ## What should stay offline/batch | |
| - Demucs source separation. | |
| - All-pairs transient NCC for large hit sets. | |
| - Agglomerative clustering. | |
| - Final ZIP packaging. | |
| - Full high-quality rerender/export. | |
| ## Recommended runtime strategy | |
| | Phase | Mode | Purpose | | |
| |---|---|---| | |
| | Upload / first pass | `stem=all`, `clustering_mode=online_preview` | Fast inspection and parameter tuning. | | |
| | Final extraction from full mix/stem | `stem=all`, `clustering_mode=batch_quality` | Better grouping without source separation. | | |
| | Final extraction from full song | `stem=drums`, `clustering_mode=batch_quality`, disk cache on | Best quality with offline Demucs cost paid once. | | |
| ## Disk cache impact | |
| Disk cache now stores decoded full mix or Demucs stem output under `.cache/stems/`, keyed by: | |
| - Source SHA-256. | |
| - Stem name. | |
| - Demucs model. | |
| - Demucs shifts. | |
| - Demucs overlap. | |
| - Device/decode mode. | |
| This does not make Demucs realtime, but it prevents repeated source separation work when retuning onset/clustering parameters for the same source and stem settings. | |
| ## Remaining realtime work | |
| The current `online_preview` mode is invoked by the batch job API after onset detection. To make the application genuinely realtime/progressive, add: | |
| 1. A streaming/ranged audio analysis API. | |
| 2. Incremental onset detector state. | |
| 3. Incremental hit artifact writing. | |
| 4. UI that appends hits/clusters as they arrive instead of waiting for the completed manifest. | |
| 5. Optional final `batch_quality` consolidation pass. | |