File size: 7,483 Bytes
b2e4883 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 | # HANDOFF β stable-jam resume guide
> This file exists so a fresh session can `cd D:/CODE/stable-jam` and pick up
> exactly where the previous one left off β no loss of context. Read this first,
> then the linked docs.
**Product:** Jam Buddy (consolidated here as **stable-jam**). "You start playing,
it joins in." A call-and-response AI music companion using Stable Audio 3,
built for the Stability AI Challenge at Music Hackspace Montreal (Aug 22β23 2026).
---
## What this repo is (and is not)
`stable-jam/` is the **source + docs + tests** for the Jam Buddy product. It does
NOT contain the heavy vendored model repos β those stay in the previous
location and are referenced by env vars (below). Don't expect to `git clone` and
get a GPU model; the SA3 weights are a separate download.
**Copied from** `D:/CODE/unstable-drums/` (the previous working dir, still intact
if you need to recover anything). This dir was created to give the hackathon a
clean, product-accurate home.
## Product state (what actually works, verified this session)
- **Web app** (`apps/web`): hardware-sampler rack UI. Instrument pads (bass/lead/
rhythm/synth/drums), Genre/Mood/Tempo knobs, a "Your take" section with BOTH a
MIDI file input and an audio file input, JOIN IN + PLAY BOTH buttons, status
panel (aria-live) showing "Buddy tempo: N BPM", and an audio player. Every
response is saved to `generations/` at the repo root.
- **SA3 pipeline** (`tools/jam_buddy.py`): the CLI that does the generation.
Two input modes:
- `--midi take.mid` β detects tempo (exact) + response length from the take,
generates a complementary melodic part at that tempo. **Ignores the groove**
(SA3 can't parse MIDI notes).
- `--wav take.wav` β TRUE audio-to-audio via `init_audio` + `init_noise_level`
(default 0.4). The buddy actually HEARS the groove and responds rhythmically.
- `--bpm N` / `--model small-music|small-sfx` / `--genre` / `--duration`.
- Drums force `small-sfx` (clean isolated hits); everything else `small-music`.
- **API route** (`apps/web/app/api/jambuddy/route.ts`): POST /api/jambuddy with
`{knobs, bpm?, midi?, audio?, duration?}`, shells to the python, returns the
WAV + `X-Jam-Buddy-BPM` header. Writes output to `<repo>/generations/`.
- **Prompt builder** (`apps/web/lib/jambuddy/prompt.ts`): pure function building
the SA3 prompt from the knobs per the official SA3 prompting guide
(TrackType: Instrument, instrument, genre, mood, BPM, studio recording).
Negative prompt steers away from a full mix. `MODEL_FOR_INSTRUMENT` maps
drumsβsmall-sfx.
- **Player** (`apps/web/lib/jambuddy/player.ts`): browser Web Audio. `playTogether`
renders the MIDI take as oscillators and plays it with the buddy WAV on the
same clock. **Note: this was "banked" β the user reported PLAY BOTH only
played the generated part, not the MIDI synth. It's parked, not debugged.**
## Verified this session (proof)
- Web: `pnpm run test` -> 31/31 pass (prompt, player, engine, conversation);
`tsc --noEmit` clean; page serves the rack.
- CLI: `jam_buddy.py --midi <file>` -> detected 158 BPM, 24.3s response matching
the take. Audio-to-audio on a 5s take -> ~10s wall, valid WAV (peak 0.43,
rms 0.08). A ~378s take times out on CPU β keep audio takes short.
- The generations dir fills with timestamped WAVs.
## Key environment / paths
SA3 (the engine) now lives IN this repo at `stable-audio-3/`. The venv's
editable-install `.pth` was repointed to this location (was `D:/CODE/unstable-drums/...`).
The `/api/jambuddy` route auto-resolves it via `join(repoRoot, "stable-audio-3", ...)`.
| Var | Value |
|---|---|
| `JAM_BUDDY_ROOT` | `D:/CODE/stable-jam` (repo root; route walks up to find `tools/jam_buddy.py`) |
| `JAM_BUDDY_PYTHON` | `D:/CODE/stable-jam/stable-audio-3/.venv/Scripts/python.exe` (the SA3 venv) |
| SA3 weights | cached in `stable-audio-3/.venv` + HF cache on G:/AI/models/huggingface |
| `HF_TOKEN` | in old repo's `.env` (needed for gated SA3 model access) β copy if regenerating weights |
The SA3 venv was extended with `mido` and `librosa`
(`uv pip install --python .../stable-audio-3/.venv/Scripts/python.exe mido librosa`).
A fresh session must use that venv or recreate it.
## How to run (fresh session)
```bash
cd D:/CODE/stable-jam
# 1. install web deps (node_modules was NOT copied)
cd apps/web && pnpm install && cd ../..
# 2. run tests + typecheck
cd apps/web && npx vitest run && npx tsc --noEmit
# 3. dev server
cd apps/web && pnpm dev # -> http://localhost:3000
# 4. python pipeline (SA3 venv is in THIS repo)
./stable-audio-3/.venv/Scripts/python.exe tools/jam_buddy.py --midi take.mid --instrument bass --out out.wav
```
## Docs (all in `docs/`)
- `01-vision.md` β Jam Buddy vision (rewritten from the old PatternTalk framing).
- `02-architecture.md` β note: may still describe the old service layout; the
real arch is: Next.js web + `tools/jam_buddy.py` + SA3 venv.
- `05-accessibility.md` β the a11y strategy (still relevant; the product is
screenreader-compatible by design).
- `06-stable-audio-integration.md` β SA3 setup, audio-to-audio, Vega notes.
- `HARNESS.md` β agent harness commands + test contract (may still say
"PatternTalk" in places; the real product is Jam Buddy).
## What's deliberately NOT in this repo (and why)
- `text2midi/` (2.8G vendored model) β not part of the Jam Buddy product; it
stays in the previous working dir if needed.
- The SA3 weights / `.venv` are inside `stable-audio-3/` here but **gitignored**
(large, regenerable, license-gated). See `.gitignore`.
- `node_modules/`, `.next/` β regenerable, not committed.
- `generations/*.wav` β gitignored (regenerable output).
- Source song/GP files + scratch WAVs in the old `tools/` β not product source.
## Pending / next steps (from the last session)
1. **PLAY BOTH** β unbank / fix: the MIDI synth part wasn't audible. Likely the
object URL or the oscillator gain. That's the "co-play" wow moment.
2. **Metadata sidecar** β add a `<file>.json` next to each `generations/*.wav`
with prompt / negative / model / BPM / duration / source. Answers "what did it
go off on."
3. **Re-render the docs** `02-architecture.md`, `03-data-model.md`,
`04-ux-voice-first.md`, `07-reaper-integration.md` to match Jam Buddy (they
still describe the old PatternTalk drum-machine in places).
4. **LoRA / underfit** (style trainer) is a real path but needs a GPU β not shipped.
5. **Demo script** β the judge's-eye review (session bg_161745) recommended a
3-min demo script; see `docs/01-vision.md` "What success looks like".
## The judge's review (session `bg_161745` β hackathon judge design review)
Key finding: **docs and product were telling different stories** (PatternTalk vs
Jam Buddy). That's why this repo was consolidated to `stable-jam`. Scoring
rubric + recs live in that session. The #1 fix was "kill the PatternTalk framing"
β largely done in README + 01-vision; the deeper docs still need the same pass.
## The SA3 tempo/time-signature reality (don't re-learn it)
Stable Audio 3 has NO tempo or time-signature conditioning channel. "BPM" is a
weak semantic hint in the prompt. So "make SA3 follow a time signature" is not
possible with the stock model. The answer is MIDI-first (timing from MIDI/audio
detection) + SA3 as the timbre generator. Full reasoning is in the skill ref
`sa3-no-tempo-control.md`. The LoRA style path is separate.
|