File size: 21,266 Bytes
b8fa9bf
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3703c4e
b8fa9bf
 
 
 
 
3703c4e
 
 
 
 
b8fa9bf
 
 
 
 
 
 
3703c4e
b8fa9bf
 
 
3703c4e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
03d531b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
e07820e
03d531b
 
 
e07820e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f08496f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3e9170f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
e33cc90
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0b5f0f0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8c10624
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7f1f066
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ce84147
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ab6f318
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f026127
 
 
 
 
 
 
 
fa35534
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5a90820
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
# Progress log

Last updated: 2026-05-12

## Pass 1: project review, timing, and Gradio replacement

Completed:

1. Inspected the original project structure and active Gradio entrypoints.
2. Moved previous Gradio interfaces into `legacy/`.
3. Created `pipeline_runner.py` as the timed orchestration layer.
4. Created `app.py` as a FastAPI backend.
5. Created a custom no-build browser frontend under `web/`.
6. Added stage timing to each extraction run.
7. Added synthetic benchmarking via `scripts/benchmark_subprocesses.py`.
8. Added initial docs for project review, timing/realtime, API, UI, and remaining work.

Outcome:

The application became usable without Gradio and produced per-run manifests/artifacts.

## Pass 2: feature ledger and continued development

Completed in this pass:

1. Added first-class docs for features, tasks, and progress.
2. Added `GET /api/jobs` for active/completed run listing.
3. Added run-history UI panel that indexes `.runs/*/output/manifest.json`.
4. Added disk caching for decoded full mix and Demucs stem outputs.
5. Extended cache clearing to remove both memory and disk cache.
6. Added `clustering_mode` pipeline parameter.
7. Added `online_preview` clustering using prototype assignment.
8. Added frontend controls for clustering mode and disk cache.
9. Fixed duplicate sample writes in `sample_extractor.build_archive`.
10. Updated README and docs to reflect the new state.

Outcome:

The project now has a clearer product surface: final-quality batch extraction, faster online-style preview clustering, persistent run history, and explicit docs tracking what is done versus still missing.

## Current assessment

The application is not “fully complete” as an editing workstation, but it is substantially implemented as an extraction and review workstation. The remaining gaps are concentrated around mutating corrections/editing, run comparison, and frontend engineering hardening.

## Next recommended pass

Implement the editing loop:

1. Add edit state for deleted/restored hits and shifted onsets.
2. Add label change, cluster reassignment, merge, and split.
3. Re-export without rerunning Demucs/onset detection when only grouping changes.
4. Save edit decisions into the manifest.
5. Add side-by-side run comparison for parameter tuning.

## Validation performed in this pass

- Compiled active Python files with `python3 -m py_compile app.py pipeline_runner.py sample_extractor.py scripts/*.py`.
- Ran FastAPI smoke job through `scripts/test_api_job.py`.
- Ran an online-preview API smoke job with synthetic audio.
- Verified `GET /api/jobs` history output and `POST /api/cache/clear` behavior.
- Verified SSE completion and review-hit artifact serving.
- Refreshed batch and online benchmark JSON files:
  - `docs/benchmark-subprocesses.json`
  - `docs/benchmark-online-preview.json`

## Pass 3: hit review and streaming progress

Completed in this pass:

1. Added `GET /api/jobs/{job_id}/events` as a server-sent-events progress stream.
2. Updated the frontend to consume SSE via `EventSource`, with the existing polling loop retained as fallback.
3. Added per-hit review artifact export under `review/hits/`.
4. Added a top-level `hits` array to each run manifest with onset, duration, classification, cluster label, representative flag, and file path.
5. Added API serialization for hit playback/download URLs.
6. Added selected-hit and selected-sample audio players.
7. Made waveform onset markers clickable by selecting the nearest detected hit.
8. Added hit table and sample-table audition controls.
9. Hardened artifact file serving by using resolved path containment via `Path.relative_to()`.
10. Refreshed batch and online benchmark JSON files after the review-hit export change.

Outcome:

The app now supports a real review loop for inspecting what the onset detector and clustering produced. Users can audition individual detected slices, representative samples, stem audio, and reconstruction audio from one screen. Progress updates are lower-latency and less wasteful via SSE while still remaining robust in browsers that need polling fallback.

## Pass 4: interactive supervised UX foundation

Completed in this pass:

1. Added the supplied interactive UX document set under `docs/interactive-ux/`.
2. Read and aligned the UX documents with the project as currently implemented.
3. Added `supervised_state.py` for persistent semantic state beside each completed run manifest.
4. Added `supervision_state.json` generation after each successful extraction.
5. Added state schema for hits, clusters, constraints, events, suggestions, confidence, review queue, and undo snapshots.
6. Added supervised editing endpoints:
   - `GET /api/jobs/{job_id}/state`
   - `POST /api/jobs/{job_id}/hits/{hit_id}/move`
   - `POST /api/jobs/{job_id}/hits/{hit_id}/pull-out`
   - `POST /api/jobs/{job_id}/hits/{hit_id}/suppress`
   - `POST /api/jobs/{job_id}/hits/{hit_id}/review`
   - `POST /api/jobs/{job_id}/clusters/{cluster_id}/lock`
   - `GET /api/jobs/{job_id}/suggestions`
   - `POST /api/jobs/{job_id}/suggestions/{suggestion_id}/accept`
   - `POST /api/jobs/{job_id}/suggestions/{suggestion_id}/reject`
   - `GET /api/jobs/{job_id}/explain/cluster/{cluster_id}`
   - `POST /api/jobs/{job_id}/undo`
7. Added an interactive supervision UI panel with:
   - state summary,
   - selected-hit actions,
   - target cluster picker,
   - outlier-first review queue,
   - cluster board,
   - suggestion inbox,
   - constraint/event log,
   - cluster explanation drawer.
8. Added `scripts/test_interactive_supervision.py` to verify the supervised API loop.

Outcome:

The app is now an extraction and supervised-review workstation at the semantic-state level. User corrections are persisted as constraints/events and can be inspected, suggested from, and undone. The next required step is edited-state export so these decisions affect downloadable artifacts.

## Current assessment after Pass 4

The project now satisfies the first interactive UX milestone for replayable supervised state:

```text
analyze audio
→ inspect hits/clusters
→ move/pull/suppress/favorite/lock
→ persist constraints/events
→ update confidence and review queue
→ generate/accept/reject suggestions
→ explain clusters
→ undo semantic edits
→ reload completed run with decisions intact
```

It now satisfies the first full semantic-edit loop because edited semantic state can be rendered into separate supervised sample WAVs, MIDI, reconstruction, and ZIP output.

## Next recommended pass after Pass 4

1. Add cluster merge/relabel/split workflows.
2. Add feature-vector cache for true local reclustering.
3. Add edited-vs-original run comparison.
4. Add browser-level UI tests and migrate frontend to TypeScript/Vite after UX stabilizes.


## Pass 5: supervised export, force-onset, restore, and suggestion diffs

Completed in this pass:

1. Added `supervised_export.py` to render `supervision_state.json` into edited artifacts under `supervised/`.
2. Added `POST /api/jobs/{job_id}/export` for edited sample-pack export.
3. Added `POST /api/jobs/{job_id}/hits/force-onset` to create user-forced hit slices from `stem.wav`.
4. Added add-onset waveform mode in the frontend.
5. Added `POST /api/jobs/{job_id}/hits/{hit_id}/restore` and a restore button for suppressed hits.
6. Added exact suggestion diff previews through `suggestion.diff` and a UI `Diff` action.
7. Updated supervised export to exclude suppressed hits and honor favorite/pinned representatives.
8. Added `scripts/test_supervised_export_and_force_onset.py`.
9. Added `docs/SUPERVISED_EXPORT_AND_FORCE_ONSET.md` and updated feature/API/task/progress docs.

Outcome:

The project now closes the main semantic-edit loop:

```text
analyze audio
→ inspect hits/clusters
→ move/pull/suppress/restore/favorite/lock/force-onset
→ inspect suggestions and diffs
→ export edited WAV/MIDI/reconstruction/ZIP artifacts
```

The original batch artifacts remain immutable. Edited outputs are written separately under `supervised/`.

Validation performed in this pass:

- `python3 -m py_compile app.py pipeline_runner.py sample_extractor.py supervised_state.py supervised_export.py scripts/*.py`
- `node --check web/app.js`
- `python3 scripts/test_supervised_export_and_force_onset.py`
- `python3 scripts/test_sse_and_review_hits.py`
- `python3 scripts/test_interactive_supervision.py`
- `python3 scripts/test_api_job.py`

Next recommended pass after Pass 5:

1. Add cluster merge/relabel/split workflows.
2. Add cached feature-vector local reclustering around edited hits.
3. Add edited-vs-original run comparison.
4. Add browser-level UI tests and migrate the frontend to TypeScript/Vite once the UX stops shifting.


## Pass 6: visual simplification toward the supplied reference UI

Completed in this pass:

1. Reworked `web/index.html` around a simpler first screen: top file/action bar, large waveform workspace, compact right-side extraction controls, sample cards, and lower utility panels.
2. Replaced the previous dark dashboard styling with a light, minimal, card-based visual system closer to the supplied screenshot.
3. Moved advanced extraction controls into a collapsible panel so the primary workflow exposes only stem, sensitivity, cluster count, extraction, export, and fast-mode decisions.
4. Added representative sample cards with waveform thumbnails decoded from the sample WAV URLs in the browser.
5. Updated waveform rendering to use grey filled waveforms plus colored lollipop onset markers similar to the reference image.
6. Preserved all existing IDs/API wiring so supervised editing, SSE progress, run history, force-onset, and edited export continue to work.

Validation performed in this pass:

- `python3 -m py_compile app.py pipeline_runner.py sample_extractor.py supervised_state.py supervised_export.py scripts/*.py`
- `node --check web/app.js`
- HTML parser sanity check for `web/index.html`
- `python3 scripts/test_api_job.py`

Outcome:

The app now looks and behaves more like a focused sample-extraction tool instead of a generic control dashboard while keeping the advanced review/supervision functionality available below the main workflow.


## Pass 6: minimal waveform-first UI

Completed in this pass:

1. Reworked the frontend visual language from a dense workstation dashboard into a light waveform-first interface.
2. Added the top file identity bar and primary purple extraction action.
3. Promoted the waveform and extracted sample cards as the default workflow.
4. Moved advanced controls and power-user tools into collapsible areas.

Outcome:

The app matched the general direction of the supplied reference image, but still had visible complexity: separate audio players, a waveform header, fast-mode buttons in the right card, and multiple utility panels visible in the main page flow.

## Pass 7: reference-alignment hardening

Completed in this pass:

1. Removed the visible waveform header and made the waveform card visually quiet.
2. Added a single custom transport row with play, elapsed/total time, and seek line.
3. Hid native source/stem/reconstruction/hit/sample audio elements from the default layout while preserving playback behavior.
4. Moved `Online preview mode` and `Fast full-mix mode` into the collapsed `Advanced` panel.
5. Collapsed pipeline, run history, supervision, and detailed tables into one `Review & edit` workbench below the sample cards.
6. Simplified sample cards to thumbnail waveform, play affordance, and label-first presentation.
7. Updated docs to distinguish visual-reference work from remaining interaction-depth work.

Outcome:

The default UI is now aligned with the supplied image: file/action bar, large waveform card, compact right controls, one transport row, and sample cards. Advanced extraction and semantic editing remain available but no longer dominate the first screen.


## Pass 8: fixed no-scroll workstation layout and upload repair

Completed in this pass:

1. Reworked `web/index.html` into a fixed workstation shell: top bar, left sidebar, center workspace, right sidebar, and bottom dock.
2. Locked the document viewport with `overflow: hidden` and moved long content into panel-local scroll containers.
3. Moved pipeline logs/history/selection context into the left sidebar.
4. Moved extraction, export, and advanced settings into the right sidebar.
5. Moved semantic review/edit tools and raw tables into the bottom dock.
6. Added an explicit top-bar `Upload audio` affordance.
7. Added whole-window drag/drop handling so dropping a file anywhere on the app selects it instead of opening it in the browser.
8. Added a full-screen drag overlay to make the drop target obvious.
9. Preserved existing frontend ids and backend APIs so extraction, SSE, supervision, force-onset, and edited export remain wired.

Outcome:

The UI no longer behaves like a scrollable webpage. It now behaves like a compact desktop-style sample extraction workstation with simple expandable tool panels around a central waveform/sample workspace.

## Pass 9: full-context reproduced audio and clearer parameters

Completed in this pass:

1. Added explicit source/context/reconstruction layers to the pipeline export:
   - `source.wav`
   - `stem.wav`
   - `context_bed.wav`
   - `target_reconstruction.wav`
   - `reconstruction.wav`
2. Changed `reconstruction.wav` to a full-context reproduced mix: non-target context bed plus sample-triggered target reconstruction.
3. Kept `target_reconstruction.wav` as the isolated sample-only target layer for debugging and focused listening.
4. Matched the target reconstruction RMS to the target stem before mixing it back into context.
5. Updated `sample-pack.zip` to include both the full-context reproduced mix and target-stem reconstruction.
6. Updated supervised edited export so edited packs follow the same audio-layer model under `supervised/`.
7. Added Source / Stem / Reproduced preview modes to the single transport row; completed jobs default to Reproduced.
8. Reworked the right sidebar into Common controls vs Advanced parameters.
9. Grouped advanced parameters by pipeline stage: stem separation, hit detection, grouping, export/cache.
10. Added `docs/REPRODUCED_AUDIO_AND_PARAMETERS.md`.

Outcome:

The app is easier to understand for normal use: the main right-side controls are now only the few controls users are likely to touch repeatedly, while lower-level DSP/model controls stay available but grouped by stage. Reproduced audio is now useful for musical judgment because separated-stem runs are previewed inside the rest of the mix rather than as an isolated sample-triggered stem only.

## Pass 10: clean default UI

Completed in this pass:

1. Collapsed secondary panels by default: source hints, selection context, pipeline, run history, exports, and review/edit no longer open on initial load.
2. Reduced the bottom dock to a compact tab bar unless a bottom tool is explicitly opened.
3. Shortened the top bar and reduced logo/upload/backend/action button visual weight.
4. Shrunk sidebars, panel padding, gaps, transport height, and sample-card dimensions to give the waveform and samples more room.
5. Hid common-control helper paragraphs and long field hints from the default extraction flow.
6. Preserved advanced controls, logs, history, supervision, tables, and edited export behind existing expandable panels.
7. Added `docs/CLEAN_DEFAULT_UI.md` to document the clean-default layout contract.

Outcome:

The UI remains a fixed workstation with sidebars and a bottom dock, but the first screen is now much less cluttered. The visible default workflow is upload, extract, listen, inspect waveform, and audition sample cards; debug/review/power-user tools are available only when opened.


## Pass 11: immediate waveform and real progress flow

Completed in this pass:

1. Added browser-side waveform rendering immediately after file selection/drop using Web Audio decoding.
2. Added a persistent uploaded-source overview so the waveform is visible before backend extraction starts.
3. Added a backend `progress` object to job payloads and SSE job events.
4. Extended `StageTiming` with `progress`, `work_done`, and `work_total`.
5. Added chunk-accurate Demucs stem progress reporting through a progress pool passed to `apply_model`.
6. Changed the waveform renderer to show active extraction progress as an accent-colored completed region on the left and a neutral remaining region on the right.
7. Added a `Start here` left-sidebar flow with Load → Controls → Extract → Review/export states.
8. Updated the stage list to include exact per-stage progress bars where progress is available.
9. Added `docs/IMMEDIATE_WAVEFORM_AND_REAL_PROGRESS.md`.

Outcome:

The app now starts with a clear sequence instead of a blank workstation. Uploading or dropping a file immediately produces a waveform, and extraction progress is visualized directly on that waveform using backend-reported progress only. The UI does not fake ETA or time-based progress; non-instrumented stages step only at real boundaries.

## 2026-05-12 — API validation and visible error pass

Completed:

- Fixed a real `/api/jobs` regression where the UI submitted `subdivision` as a string and the backend rejected it with `400`.
- Added backend parameter coercion before validation for integer, float, and boolean pipeline params.
- Updated frontend `collectParams()` to use `/api/config.defaults` as the type source for controls, including `<select>` elements.
- Added a global error banner so failed API requests and pipeline errors are visible without browser devtools.
- Removed a duplicate `historyList` id in the HTML.
- Added `docs/API_ERRORS_AND_PARAMETER_VALIDATION.md` and `scripts/test_param_validation_and_api_errors.py`.

Validation:

```bash
python3 scripts/test_param_validation_and_api_errors.py
node --check web/app.js
python3 -m py_compile app.py pipeline_runner.py sample_extractor.py supervised_state.py supervised_export.py scripts/*.py
```

## Pass: automatic simple card flow

Completed in this pass:

1. Changed the default interaction model from manual configuration to automatic processing after drop/upload.
2. Added `PipelineParams.auto_tune` and a new `Automatic parameter tuning` pipeline stage.
3. Added progressive `sample` events from the export stage and serialized active-job `partial_samples`.
4. Redesigned sample cards into type-grouped columns with draw/dismiss/trim/extend actions.
5. Added waveform zoom and pan while preserving click-to-select/force-onset semantics.
6. Replaced expensive representative scoring internals with a lightweight deterministic scorer.
7. Added fallback MIDI writing when `pretty_midi` is missing.
8. Hardened BPM detection against sparse/empty beat events.

Validation performed:

- Python compile checks for app, pipeline, DSP, supervision, export, and scripts.
- `node --check web/app.js`.
- DOM id/reference consistency check.
- `scripts/test_param_validation_and_api_errors.py`.
- `scripts/test_progress_contract.py`.
- `scripts/test_api_job.py`.
- `scripts/test_sse_and_review_hits.py`.
- `scripts/test_supervised_export_and_force_onset.py`.
- `scripts/test_interactive_supervision.py`.


## 2026-05-12 Reference image UI alignment

- Reworked the frontend shell to match the provided Sample Extractor reference image.
- Added centered file picker/current filename, right-aligned export actions, persistent right settings panel, waveform-first canvas, grouped sample columns, and compact bottom selection bar.
- Kept automatic drop-to-process behavior and progressive sample-card rendering.
- Moved secondary pipeline/history/supervision/tables into a compact tools drawer.


## Pass 14: selected cards and Spleeter backend

Completed in this pass:

1. Added `spleeter` as the default separation backend, with selectable `spleeter:2stems`, `spleeter:4stems`, and `spleeter:5stems` profiles.
2. Kept `demucs` as a quality/fallback backend and `none` as the full-mix preview backend.
3. Added optional `requirements-spleeter.txt` instead of forcing TensorFlow/Spleeter into the base install.
4. Added per-card checkbox state with manual select-all/clear behavior.
5. Added selected-only backend export via `POST /api/jobs/{job_id}/export-selected`.
6. Made `draw another` persist the chosen representative in semantic state.
7. Made trim/extend rewrite playable preview audio immediately under `overrides/hits/`.
8. Added `scripts/test_selected_export_card_actions.py`.

Outcome:

The default app now behaves more like a card review tool: drop audio, let Spleeter/fallback separation run, review grouped cards, select/dismiss/draw/trim, and export only the selected pack.

## Upload/runtime fallback update (2026-05-12)

- Added a visible top-bar `Choose audio` affordance in addition to whole-app drag/drop.
- Fixed the default hidden state of the error banner so placeholder errors are not shown on page load.
- API errors now surface request path/status/detail in the visible banner and pipeline logs.
- `/api/config` now includes runtime diagnostics for optional separation backends.
- If Spleeter is unavailable, the simple UI keeps the app usable by switching to full-mix mode; backend fallback also uses full-mix rather than silently launching Demucs.