File size: 9,569 Bytes
b2e4883 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 | # 08 β Build Plan
## Timeline overview
```
NOW (3 weeks before) Pre-hack prep
HACK DAY 1 (Sat Aug 22) Build core features
HACK DAY 2 (Sun Aug 23) Polish, integrate, demo
POST-HACK Publish, ship, follow up
```
## Pre-hackathon (you have ~3 weeks)
Goal: arrive at the event with a working skeleton, training data ready, and most patterns written. This compresses day 1 to "polish core, add stretch features" instead of "build from zero."
### Week 1: foundations
**By end of week 1:**
- [ ] Get Stable Audio 3 API access / model weights downloaded
- [ ] HuggingFace account created
- [ ] Request gated access to `stabilityai/stable-audio-3-small`
- [ ] Get inference running locally on Vega (or document why cloud-only)
- [ ] Time a generation, establish baseline performance
- [ ] Set up RunPod account (or chosen cloud GPU provider)
- [ ] Test deploying a simple inference container
- [ ] Verify you can ssh/exec into it
- [ ] Stand up the repo
- [ ] Next.js + TypeScript + Tailwind initialized
- [ ] Folder structure per [`02-architecture.md`](02-architecture.md#folder-structure)
- [ ] CI with axe-core for accessibility (even on empty pages, it's a foundation)
- [ ] Deploy "coming soon" placeholder to Vercel
- [ ] Set up audio service skeleton
- [ ] FastAPI + SA3 inference wrapper
- [ ] Single `/generate` endpoint, no LoRA yet
- [ ] Install NVDA on Windows, run through Reaper once to baseline familiarity
- [ ] Set up Reaper with a drum sampler VST for testing MIDI export
### Week 2: data and templates
**By end of week 2:**
- [ ] Curate training data for brutal-drum LoRA
- [ ] 20β40 one-shots (recordings or sourced, all licensed)
- [ ] 10β20 short loops
- [ ] Manifest in `data/training/manifest.yaml`
- [ ] All preprocessed to 44100 Hz mono
- [ ] Write the 6 demo-critical pattern templates
- [ ] `d-beat.json`
- [ ] `blast-traditional.json`
- [ ] `skank.json`
- [ ] `half-time-metal.json`
- [ ] `djent-polyrhythm.json`
- [ ] `punk-rock.json`
- [ ] Write the onomatopoeia mapping table with at least 10 entries
- [ ] Implement prompt parser MVP (handles the 6 patterns + 10 onomatopoeias)
- [ ] Implement pattern engine MVP (loads templates, expands to bars, applies tempo)
- [ ] Implement MIDI generation (downloads work, drags into Reaper)
### Week 3: integration and rehearsal
**By end of week 3:**
- [ ] Train brutal-drum LoRA on cloud GPU
- [ ] First attempt: 500 steps, see how it shapes up
- [ ] Iterate: adjust learning rate, dataset, prompt format
- [ ] Final: ship-ready LoRA + publish to HuggingFace (draft, don't publish yet)
- [ ] Implement Reaper Web Control client
- [ ] Connect, parse state, extract tempo
- [ ] Tempo resolution priority implemented
- [ ] Implement Web Speech API integration
- [ ] STT works in Chrome
- [ ] TTS works
- [ ] Voice state machine (idle β listening β parsing β generating β ready)
- [ ] Build basic UI shell
- [ ] Voice button, transcript display, status panel
- [ ] Action buttons (play, download, regenerate)
- [ ] Visual grid component (basic version)
- [ ] Accessibility pass 1
- [ ] Tab through every element
- [ ] NVDA run-through of the voice flow
- [ ] ARIA labels on every interactive element
- [ ] Skip links
- [ ] High contrast check
- [ ] End-to-end rehearsal
- [ ] Generate a pattern in PatternTalk
- [ ] Drag MIDI into Reaper
- [ ] Play it
- [ ] Generate sample
- [ ] Play sample
- [ ] Switch on NVDA, redo the flow
- [ ] Write the team-pitch and post to Music Hackspace Discord
## Hackathon Day 1 (Saturday Aug 22)
### Morning (9:00 AM β 12:30 PM): setup + team formation
- [ ] Arrive at PHI Centre, get settled
- [ ] Finalize team composition (if recruiting on-site)
- [ ] Confirm Reaper + PatternTalk working environment
- [ ] Connect to venue WiFi, verify cloud GPU access
### Early afternoon (12:30 PM β 3:00 PM): core demo path
**Goal:** End-to-end voice β pattern β MIDI β Reaper working.
- [ ] Finalize pattern engine: tempo, bars, time signature, cymbal overrides
- [ ] Finalize MIDI export, verify it sounds right in Reaper
- [ ] Wire up Web Speech API end-to-end
- [ ] Wire up sample generation from cloud GPU
- [ ] First smoke test: voice prompt β MIDI + sample β download β Reaper
### Late afternoon (3:00 PM β 6:00 PM): variations and audio context
- [ ] Implement 4-variation engine (humanize velocities, micro-timing, fill variants)
- [ ] Implement uploaded audio BPM detection (Meyda)
- [ ] Wire up Reaper tempo sync (Web Control client)
- [ ] Pre-generate demo variations for the showcase patterns
### Evening (6:00 PM β 9:00 PM): polish + accessibility
- [ ] Visual grid component (sighted users)
- [ ] Pattern library page (skeleton)
- [ ] Accessibility pass with NVDA
- [ ] Tab through everything
- [ ] Voice announcements work
- [ ] Error states announced
- [ ] Demo rehearsal (round 1)
- [ ] Evening check-in / informal demos at the venue
**Day 1 deliverable:** A working PatternTalk that takes voice prompts, generates MIDI, downloads it, and plays in Reaper. Plus 4-variation picker. Plus screenreader-tested.
## Hackathon Day 2 (Sunday Aug 23)
### Morning (9:00 AM β 12:00 PM): stretch features
Priority order (cut from the bottom if time runs short):
1. [ ] **Sample preview player** with progress indicator
2. [ ] **Pattern library** with search/filter and shareable URLs
3. [ ] **User-defined onomatopoeias** (localStorage)
4. [ ] **CLAP plugin wrap** of the web UI (if a teammate can take it)
5. [ ] **Real-time audio captioning** (stretch, probably cut)
### Midday (12:00 PM β 2:00 PM): polish
- [ ] Final visual grid polish (animations, color tuning)
- [ ] Voice response tuning (speak rate, prompts)
- [ ] Pre-generate all demo cache samples
- [ ] Pre-stage the LoRA weights download for live demo
- [ ] Test with cloud GPU off (Vega-only path) as fallback
### Afternoon (2:00 PM β 4:00 PM): demo rehearsal
- [ ] Full demo run-through, timed (target: 3β4 minutes)
- [ ] NVDA demo run-through, timed
- [ ] VoiceOver demo run-through, timed
- [ ] Rehearse the pitch (3 sentences max)
- [ ] Backup plan: video recording of working demo, in case of network issues
### Late afternoon (4:00 PM β 6:00 PM): present
- [ ] Public demos
- [ ] Jury presentation
- [ ] Feedback session
### Evening (6:00 PM onwards): wind down
- [ ] Final cleanup
- [ ] Publish LoRA to HuggingFace (if not already)
- [ ] Push code to GitHub, write a proper README
- [ ] Write blog post / social announcement
## Checkpoints
Use these to catch problems early. If a checkpoint fails, stop and address it before moving on.
| Checkpoint | When | What "passing" means |
|---|---|---|
| **C1: Inference works** | Pre-hack week 1 | SA3 small generates a sample on Vega in < 10 min OR cloud GPU works |
| **C2: Training data ready** | Pre-hack week 2 | 30+ samples in `data/training/` with manifest, all licensed |
| **C3: 6 patterns authored** | Pre-hack week 2 | All 6 demo patterns load, expand, generate MIDI |
| **C4: End-to-end MIDI** | Pre-hack week 3 | Voice prompt β MIDI β Reaper plays correctly |
| **C5: NVDA reads the UI** | Pre-hack week 3 | Screenreader announces every state change |
| **C6: Demo path works** | Hack day 1 | Full voice β MIDI β sample β Reaper in < 60 seconds |
| **C7: LoRA published** | Hack day 1 evening | Weights on HuggingFace with model card |
| **C8: Variations work** | Hack day 2 morning | 4 variations show in UI, picker works |
| **C9: Demo polished** | Hack day 2 afternoon | Rehearsal runs in 3β4 minutes without mistakes |
## What to cut if time runs out
In priority order (drop from the bottom):
1. β Real-time audio captioning
2. β Haptic metronome PWA
3. β Reduced-physical-load mode
4. β CLAP plugin wrap
5. β User-defined onomatopoeias
6. β Pattern library page (if a minimal in-app picker works)
7. β Audio context BPM detection (fall back to prompt/Reaper-only)
8. β Visual grid polish (fall back to functional ugly grid)
9. β Dark mode (fall back to light only)
The must-haves are: voice prompt, pattern generation, MIDI export, sample generation, Reaper tempo sync. If we have those five, we have a demo that lands.
## What to add if time is plentiful
In priority order:
1. β
CLAP plugin wrap
2. β
Audio context BPM detection
3. β
Pattern library page with shareable URLs
4. β
User-defined onomatopoeias
5. β
Visual grid polish with animations
6. β
Dark mode + theme support
7. β
Multiple LoRA adapters (one per genre: rock, jazz, funk)
8. β
Real-time audio captioning
## Communication cadence
- **Standup** (15 min) at start of each day
- **Mid-day check-in** (5 min, async) β what's blocked, what's next
- **End-of-day retro** (15 min) β what shipped, what's deferred
- **Demo** (3β4 min) at end of each day for informal feedback
## Post-hackathon
**Within 1 week:**
- [ ] Publish LoRA weights to HuggingFace
- [ ] Publish repo on GitHub with full README
- [ ] Blog post / dev.to article about the project
- [ ] Submit to Music Hackspace showcase if invited
**Within 1 month:**
- [ ] Add more pattern templates (target: 30+)
- [ ] Train LoRAs for additional genres
- [ ] Add CLAP plugin wrap (if not done at hackathon)
- [ ] Integrate with Narwall for continuous accessibility testing
**Within 3 months:**
- [ ] Public launch of the web app
- [ ] Community pattern library with moderation
- [ ] "Train your own LoRA" UI
- [ ] Mobile companion app for haptic metronome
|