Spaces:
Sleeping
Sleeping
| # Feature feasibility matrix | |
| ## Scoring convention | |
| All scores are between `0.0` and `1.0`. | |
| For positive dimensions, higher is better: | |
| - **Feasibility**: likelihood this can be implemented robustly with the current project direction. | |
| - **Value UX**: how much nicer/faster the application feels. | |
| - **Value quality**: expected improvement in extraction correctness. | |
| - **Realtime fit**: how well the interaction can update from cached features without full reruns. | |
| - **MVP fit**: whether it belongs in the first serious interactive version. | |
| For cost dimensions, higher is harder/more expensive: | |
| - **Complexity**: algorithmic/product complexity. | |
| - **Effort**: implementation effort. | |
| ## Matrix | |
| | Interaction | Feasibility | Complexity | Effort | Value UX | Value quality | Realtime fit | MVP fit | Verdict | | |
| |---|---:|---:|---:|---:|---:|---:|---:|---| | |
| | Move sample to cluster β auto-tune clustering | 0.90 | 0.55 | 0.45 | 0.95 | 0.90 | 0.85 | 0.95 | Build early | | |
| | Pull sample out β protect distinction | 0.90 | 0.55 | 0.45 | 0.90 | 0.90 | 0.85 | 0.95 | Build early | | |
| | Lock confirmed cluster identity | 0.95 | 0.25 | 0.20 | 0.80 | 0.75 | 0.95 | 1.00 | Build immediately | | |
| | Outlier-first review queue | 0.95 | 0.30 | 0.25 | 0.90 | 0.80 | 0.95 | 1.00 | Build immediately | | |
| | Low-confidence visual emphasis | 0.95 | 0.25 | 0.20 | 0.85 | 0.65 | 0.95 | 0.95 | Build immediately | | |
| | Bleed brush/suppression | 0.85 | 0.55 | 0.50 | 0.90 | 0.85 | 0.80 | 0.85 | Build early | | |
| | Click-to-add missed onset | 0.90 | 0.45 | 0.40 | 0.90 | 0.80 | 0.90 | 0.90 | Build early | | |
| | Cluster naming as reusable semantic hint | 0.75 | 0.65 | 0.55 | 0.75 | 0.65 | 0.70 | 0.40 | Useful, not first | | |
| | Star/favorite sample β optimize around it | 0.95 | 0.35 | 0.30 | 0.75 | 0.70 | 0.90 | 0.75 | Build early | | |
| | Explain this cluster | 0.85 | 0.50 | 0.45 | 0.80 | 0.65 | 0.85 | 0.80 | Build early | | |
| | Live counterfactual parameter previews | 0.80 | 0.70 | 0.65 | 0.90 | 0.75 | 0.75 | 0.55 | High value, later | | |
| | Temporal pattern supervision | 0.75 | 0.70 | 0.65 | 0.80 | 0.85 | 0.70 | 0.45 | Later | | |
| | Reconstruction-error-driven correction | 0.70 | 0.80 | 0.75 | 0.85 | 0.90 | 0.55 | 0.35 | Later, powerful | | |
| | Multi-resolution semantic clustering | 0.80 | 0.70 | 0.65 | 0.85 | 0.80 | 0.75 | 0.50 | Later | | |
| | Auto-clean this family | 0.80 | 0.60 | 0.55 | 0.75 | 0.80 | 0.75 | 0.50 | Later | | |
| | Drag clusters in semantic space | 0.70 | 0.75 | 0.75 | 0.90 | 0.65 | 0.65 | 0.30 | Cool, not first | | |
| | Cluster gravity / physics metaphor | 0.60 | 0.80 | 0.80 | 0.80 | 0.50 | 0.60 | 0.20 | Risky novelty | | |
| | Context-aware classification | 0.65 | 0.80 | 0.75 | 0.65 | 0.80 | 0.55 | 0.25 | Researchy | | |
| | Teach mode across songs | 0.70 | 0.85 | 0.80 | 0.85 | 0.85 | 0.60 | 0.25 | Later platform feature | | |
| | Predictive batch questions | 0.85 | 0.55 | 0.50 | 0.90 | 0.80 | 0.80 | 0.70 | Build after uncertainty scoring | | |
| ## Highest ROI set | |
| | Rank | Feature | Why | | |
| |---:|---|---| | |
| | 1 | Lock confirmed cluster identity | Easy and prevents frustrating reclustering drift | | |
| | 2 | Outlier-first review queue | Huge UX gain from simple uncertainty ranking | | |
| | 3 | Move sample to cluster as supervision | Core differentiator; directly improves results | | |
| | 4 | Pull sample out / protect distinction | Required counterpart to positive supervision | | |
| | 5 | Click-to-add missed onset | Direct correction beats indirect threshold tweaking | | |
| | 6 | Bleed brush | Removes many false positives quickly | | |
| | 7 | Explain cluster | Makes the system debuggable and trustworthy | | |
| | 8 | Predictive batch questions | Multiplies the effect of each correction | | |
| | 9 | Counterfactual previews | Makes advanced tuning understandable | | |
| | 10 | Reconstruction-error correction | Very powerful, but architecturally heavier | | |
| ## Technical conclusion | |
| The most feasible "magic" is not heavy ML. It is constraint-aware clustering, cached feature vectors, uncertainty scoring, and local recomputation. | |
| That foundation should be implemented before adding higher-risk semantic-space or personalized-model features. | |
| ## Implementation alignment as of 2026-05-12 | |
| | Interaction | Current status | Notes | | |
| |---|---|---| | |
| | Move sample to cluster β auto-tune clustering | partial | Implemented as semantic state mutation with `force-cluster`/`must-link` constraints and heuristic move suggestions. True constrained local reclustering is still open. | | |
| | Pull sample out β protect distinction | partial | Implemented as a new user cluster plus `cannot-link` and split suggestions. True constrained reclustering is still open. | | |
| | Lock confirmed cluster identity | done | Lock/unlock persists in `supervision_state.json` and appears in the UI. Replay into future reruns is still open. | | |
| | Outlier-first review queue | done | Implemented with heuristic confidence/priority. Feature-margin and reconstruction-impact ranking remain open. | | |
| | Low-confidence visual emphasis | done | Low-confidence and suppressed hits are visually distinguished in the hit table and review queue. | | |
| | Bleed brush/suppression | partial | Suppress selected hit and similar suppression suggestions are implemented. Region brush and restore are still open. | | |
| | Click-to-add missed onset | todo | Waveform click currently auditions nearest existing hit only. | | |
| | Cluster naming as reusable semantic hint | todo | User clusters receive generated labels; explicit rename/semantic hinting is not implemented. | | |
| | Star/favorite sample β optimize around it | partial | Favorite pins representative in semantic state; artifact re-export does not yet honor it. | | |
| | Explain this cluster | done | Explanation endpoint and UI drawer are implemented. | | |
| | Predictive batch questions | partial | Suggestion inbox exists; exact diff previews and richer question phrasing are open. | | |
| | Live counterfactual parameter previews | todo | Not implemented. | | |
| | Reconstruction-error-driven correction | todo | Not implemented. | | |
| | Multi-resolution semantic clustering | todo | Not implemented. | | |
| | Auto-clean this family | todo | Not implemented. | | |
| | Drag clusters in semantic space | backlog | Not implemented. | | |
| | Cluster gravity / physics metaphor | backlog | Not implemented. | | |
| | Context-aware classification | backlog | Not implemented. | | |
| | Teach mode across songs | backlog | Not implemented. | | |
| ## Revised highest-ROI next set | |
| | Rank | Feature | Why now | | |
| |---:|---|---| | |
| | 1 | Supervised re-export | Makes current semantic edits affect the downloadable sample pack. | | |
| | 2 | Force-onset from waveform | Adds the missing direct correction primitive for missed hits. | | |
| | 3 | Suppression restore | Required safety counterpart to suppression. | | |
| | 4 | Cached feature refs | Unlocks real local reclustering and better confidence. | | |
| | 5 | Diff preview for suggestions | Makes batch suggestions safer and more trustworthy. | | |
| | 6 | Constraint violation detection | Prevents silent conflicts once constraints become richer. | | |
| | 7 | Browser tests | Protects the increasingly stateful UI from regressions. | | |
| ## Pass 5 implementation status | |
| Implemented after initial alignment: | |
| - supervised edited-state export under `supervised/`, | |
| - add-onset waveform mode backed by `POST /api/jobs/{job_id}/hits/force-onset`, | |
| - suppressed-hit restore backed by `POST /api/jobs/{job_id}/hits/{hit_id}/restore`, | |
| - exact suggestion diff previews in API state and UI, | |
| - validation via `scripts/test_supervised_export_and_force_onset.py`. | |
| Still open: | |
| - cluster merge/relabel/split workflows, | |
| - cached feature-vector local reclustering, | |
| - edited-vs-original comparison, | |
| - browser-level UI tests. | |