# audio-brief API — audio in → SA3 prompt out The Space exposes a headless endpoint that runs the analysis pipeline and returns a ready-to-use SA3 prompt plus the measured JSON. **No wallet, no LLM** — the prompt is the deterministic structural prompt built from measurements alone. - **Endpoint:** `/audio_to_prompt` - **Base URL:** `https://kalamishere-audio-brief.hf.space/` - **Auth:** optional shared token (see *Gating* below). ## Parameters (positional) | # | Name | Type | Default | Notes | |---|---|---|---|---| | 1 | `audio` | file | — | wav/mp3/flac/ogg + mp4/mov/m4a/webm/… (video & aac-family are ffmpeg-extracted) | | 2 | `token` | str | `""` | required only when the Space secret `AUDIO_BRIEF_API_TOKEN` is set | | 3 | `fast` | bool | `false` | skip demucs stems + bass-MIDI → ~2 s instead of ~30 s; prompt omits stem/bassline detail | | 4 | `bpm_prior` | str | `"default"` | genre slug or numeric string to seed the beat tracker | | 5 | `llm` | bool | `false` | run the Pollinations LLM to write a natural-language prompt (analyse → brief → lens). Needs `POLLINATIONS_API_KEY` on the Space. **Falls back to the deterministic prompt if the LLM errors** — the call never fails because of the LLM | | 6 | `lens` | str | `"loose"` | LLM constraint: `match` (tightest, keeps section timings), `loose` (feel + BPM/key, free arrangement), `free` (mood + BPM only). Ignored when `llm=false` | | 7 | `original_prompt` | str | `""` | the caller's current gen prompt; non-empty switches both paths to anchored steering (measured tempo/key locked, user vocabulary authoritative) | | 8 | `embedding` | bool | `false` | also run the CLAP similarity stage and return the L2-normalized vector inline as `embedding` (float list) + `embedding_dim`. Runs as a **subprocess island** (torch/checkpoint never resident next to a gen); off by default — heavy, CPU-bound | ## Two prompt modes - **Deterministic (default, `llm=false`)** — the prompt is measurements formatted as structured text. No wallet, no cost, works for any caller. - **LLM-polished (`llm=true`)** — analyse → LLM writes a brief (infers genre/mood from the numbers) → LLM compresses it into a natural-language SA3 prompt. Costs Pollinations pollen, billed to the Space's `POLLINATIONS_API_KEY`. If that call fails (no key / out of pollen / timeout) the response degrades to the deterministic prompt and says so in `prompt_source`. ## Returns ```json { "prompt": "Sparse ambient micro-cue in A minor · BPM …", "prompt_source": "llm:claude-sonnet-4-6", "brief": "This is an extremely brief, near-static cue in A minor …", "match_style_prompt": "…", "embedding": [0.031, -0.052, …], // 512 floats, L2-normalized; null unless embedding=true "embedding_dim": 512, // null unless embedding=true "embedding_ckpt": "default-630k", // which CLAP checkpoint made it; don't compare across differing values "measured": { "bpm": 128.0, "key": "A", "key_mode": "minor", "key_confidence": 0.82, "duration_s": 30.0, "lufs_i": -9.1, "lufs_lra": 4.2, "true_peak_db": -0.3, "sections": [...], "stems": ["bass","drums","other","vocals"], "stem_stats": {...}, "voiceover_present": false, "bass_midi": {...} }, "fast_mode": false, "stages_ok": ["bpm_key","sections","loudness","stems","bass_midi"], "errors": [] } ``` `prompt_source` is `"llm:"` on the LLM path, `"deterministic"` when `llm=false`, or `"deterministic (llm failed: …)"` when an LLM run fell back. `brief` is `null` unless `llm=true`. ## Python (recommended) ```python from gradio_client import Client, handle_file c = Client("kalamishere/audio-brief") # or the full .hf.space URL # deterministic (no cost) r = c.predict( handle_file("track.wav"), "YOUR_TOKEN", # "" if the Space isn't gated True, # fast=True "default", False, # llm=False "loose", api_name="/audio_to_prompt", ) print(r["prompt"]) # LLM-polished (Space pays pollen) r = c.predict( handle_file("track.wav"), "YOUR_TOKEN", False, "default", True, # llm=True "loose", # lens: match | loose | free api_name="/audio_to_prompt", ) print(r["prompt"], "\n\nvia", r["prompt_source"]) ``` ## Enabling the LLM path Set a second Space secret (Settings → Variables and secrets) so the LLM calls have a wallet to bill: ``` POLLINATIONS_API_KEY = ``` Without it, `llm=true` calls still succeed — they just fall back to the deterministic prompt (you'll see the reason in `prompt_source`). ## Gating (protect the CPU-heavy endpoint) The public Space runs demucs on every full call, so leave it gated. Set a Space secret in **Settings → Variables and secrets**: ``` AUDIO_BRIEF_API_TOKEN = ``` When set, calls must pass a matching `token` (constant-time compared) or get `invalid or missing API token`. When unset, the endpoint is open — fine for local runs. ## ⚠️ The LLM path is slow — use `submit()`, not `predict()` `llm=true` makes **two** Pollinations calls (brief + prompt), so a full request runs ~40 s. `gradio_client`'s synchronous `predict()` has a short HTTP read timeout and will raise `httpx.ReadTimeout` on it. Use the job/poll API instead: ```python job = c.submit(handle_file("track.wav"), "YOUR_TOKEN", True, "default", True, "loose", api_name="/audio_to_prompt") while not job.done(): time.sleep(2) r = job.result() # blocks until the queue finishes, no read timeout print(r["prompt"], "via", r["prompt_source"]) ``` Deterministic calls (`llm=false`, `fast=true`) finish in ~2 s and are fine with plain `predict()`. If you want the LLM path to fit a single `predict()` call, it can be collapsed to one Pollinations request (skip the separate brief) — ~20 s and half the pollen, slightly less mood nuance. ## Similarity embedding (CLAP) Pass `embedding=true` to get a **512-float, L2-normalized CLAP vector** back in `embedding`. Because it's L2-normalized, cosine similarity between two tracks is just the dot product — that's the distance measurement for "how far this track sits from a trending cluster." - **Deterministic on full-length audio.** laion-clap (fusion off) randomly truncates anything longer than its ~10 s window, so a naive whole-track call returns a *different* vector every time (cos ~0.4-0.9 between repeats of the same file). The worker instead tiles the audio into fixed 10 s windows, embeds each, and means them — same file → same vector (repeat cos 1.0). A self-test repeat-embed guard fails the run if that ever regresses. - **`embedding_ckpt`** names the checkpoint that produced the vector. Vectors from different checkpoints are NOT comparable — gate any distance computation on matching `embedding_ckpt`. - **Stream-safe by construction.** CLAP is loaded only in `clap_worker.py`, a short-lived subprocess that embeds one file and exits — torch and the checkpoint are never resident in the API process, so an embed can't OOM or crash an in-flight SA3 generation. A failure lands as an `errors` entry and `embedding: null`; the rest of the response still stands. - **Poisoned-checkpoint guard.** On load the worker embeds sine / white-noise / silence and asserts they land far apart. A broken checkpoint (everything collapses to one vector, top-k still *looks* plausible) fails the call loudly instead of returning a silently-wrong vector. Threshold: `CLAP_SELFTEST_MAXCOS` (default 0.90). - **Checkpoint.** Defaults to the auto-downloaded general 630k-audioset checkpoint (HTSAT-tiny). Set `LAION_CLAP_MUSIC_CKPT` to the `music_audioset_epoch_15_esc_90.14.pt` path to use the music-specialized checkpoint (HTSAT-base) — the worker auto-matches the architecture. - **Cost.** CPU-bound, ~10-15 s on top of the analysis. Leave it off for runs that only need BPM/key/sections. ## Notes / limits - **Speed:** full analysis is demucs-bound (~30 s on HF free CPU). Use `fast=true` for interactive/high-volume use, or upgrade the Space tier. - **Concurrency:** the queue caps concurrency at 4 (see `demo.queue`). Heavy parallel traffic on the free tier will swap — see the low-RAM ceiling note in the pipeline docs. - **YouTube isn't an API input** — pass a file. (The UI's YouTube box yt-dlp's to a file first; that flow is UI-only.)