Diptych: Scoped, AI-Interpreted Comparison for Reference Listening in Music Production
Abstract
Reference listening is a common strategy in music production, but current comparison tools often obscure a key human judgment: deciding what should be compared. We present Diptych, an AI-assisted system that lets users define comparison scope across whole tracks or independently selected segments, while inspecting structured audio features and scope-specific AI interpretations. We evaluated Diptych in a within-participants study with 12 musicians, complemented by source-blinded ratings from four expert listeners. Participants used the system to surface additional differences, nine of ten of which received at least partial expert support, and reported good usability and greater clarity about possible next steps. These findings suggest that AI support for creative comparison should prioritize user-defined scope, inspectable evidence, and actionable guidance, while avoiding authoritative judgments that exceed what the evidence can support.
Community
Reference listening is a common strategy in music production, but current comparison tools often obscure a key human judgment: deciding what should be compared. We present DipTych, an AI-assisted system that lets users define comparison scope across whole tracks or independently selected segments, while inspecting structured audio features and scope-specific AI interpretations. We evaluated DipTych in a within-participants study with 12 musicians, complemented by source-blinded ratings from four expert listeners. Participants used the system to surface additional differences, nine of ten of which received at least partial expert support, and reported good usability and greater clarity about possible next steps. These findings suggest that AI support for creative comparison should prioritize user-defined scope, inspectable evidence, and actionable guidance, while avoiding authoritative judgments that exceed what the evidence can support.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Louder, Longer, Livelier: Acoustic Shortcuts and Underspecified Rationales in Speech LLM Judges (2026)
- CommSketch: How Speaking while Sketching Steers Human--AI Design Ideation (2026)
- MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation (2026)
- SpeechCritic: Learning a Diagnostic Speech Judge from Limited Human Preferences (2026)
- VoxReason: Auditing Source-Grounded Speech Plans Before Synthesis (2026)
- H2H Music Improv: A Communication Model and Audio-Visual Dataset for Music Improvisation (2026)
- CoLMbo-SV: A Grounded Language Model for Explainable Speaker Verification (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.39963 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper