HOA7 Spatial Field Decoder (hoa64 v0.5.0): 7th-order Ambisonics encode/decode, Wigner-D rotation, DOA analysis, vision fuse, diffusion conditioning
570b87b verified | language: | |
| - en | |
| library_name: other | |
| pipeline_tag: other | |
| license: other | |
| tags: | |
| - ambisonics | |
| - spatial-audio | |
| - hoa | |
| - spherical-harmonics | |
| - audio-analysis | |
| - tool | |
| - agi-tool | |
| - wigner-d | |
| # HOA7 Spatial Field Decoder (hoa64) | |
| A **formula-driven spatial field decoder** for **7th-order Ambisonics** — **64 channels** (Ambix **ACN + SN3D**). It encodes/decodes spherical fields, rotates them with **Wigner-D** matrices, analyzes direction-of-arrival / energy, fuses vision detections onto the sphere, and emits **spatial conditioning** for generative pipelines. Pure NumPy; **not an LLM and not a learned model** — deterministic spherical-harmonic geometry, no weights. | |
| | Property | Choice | | |
| |----------|--------| | |
| | Version | `0.5.0` | | |
| | Geometry | Ambix **ACN + SN3D** | | |
| | Channels | 64 = (7+1)², order 7 | | |
| | Rotation | **Wigner-D** (~150× vs dense) | | |
| | Audio | WAV / live mic (ffmpeg·arecord) | | |
| | Vision | Boxes / YOLO labels / optional torchvision | | |
| | Agents | HTTP `:8765` + `spatial-report` CLI | | |
| | Diffusion | Conditioning JSON + optional ComfyUI submit | | |
| | Language | Python 3, NumPy only (optional torch/torchvision) | | |
| ## What it does | |
| - **Encode** point sources / plane waves / scene mixes → HOA-7 coefficients (`encode_points`, `encode_plane_waves`, `encode_scene`). | |
| - **Decode** coefficients → samples on the sphere, product-grid render, or a single beamform readout (`decode_directions`, `decode_grid`, `beamform`). | |
| - **Rotate** the field in the listener frame via Wigner-D / zyz (`hoa_rotation_matrix`, `apply_hoa_rotation`, `rotate_yaw_pitch_roll`). | |
| - **Analyze** DOA (intensity vector + peak search), directional power, field energy, per-STFT-band and per-frame reports (`analysis.py`, `report.py`). | |
| - **Fuse vision** — object boxes / rays / YOLO detections projected onto the sphere alongside audio (`vision.py`, `detector.py`, `fuse_reports`). | |
| - **Condition** — spatial reports → plain-text control lines for T2I/T2V prompts, structured JSON for ControlNet-style nodes, or a ComfyUI API payload (`conditioning.py`). | |
| - **Serve** — HTTP API and iterative agent state (`server.py`, `rnn_stub.py`). | |
| ## Coordinates (always) | |
| - **+X** front, **+Y** left, **+Z** up (Ambix listener frame). | |
| - Azimuth 0° = front, **+90° = left**, −90° = right. | |
| - Elevation 0° = horizon, +90° = zenith. | |
| - W (omnidirectional HOA channel) maps to field size / POV: low W → tight / subject-focused / narrow FOV; high W → wide / environmental / immersive FOV. | |
| ## Install / run | |
| ```bash | |
| cd spatial-hoa # this repo root | |
| export PYTHONPATH="$PWD${PYTHONPATH:+:$PYTHONPATH}" | |
| python3 -m hoa64 --help | |
| # or install the CLI on PATH: | |
| ln -s "$PWD/scripts/spatial-report" ~/.local/bin/spatial-report | |
| ``` | |
| **Deps:** NumPy (system `ffmpeg`/`arecord` for live capture; optional torch/torchvision for the real object detector). | |
| ## CLI map (`spatial-report`) | |
| | Command | Purpose | | |
| |---------|---------| | |
| | `analyze` | Ambix / mono WAV → spatial JSON | | |
| | `demo-scene` | Synthetic multi-source audio | | |
| | `vision` | Raw sphere boxes → report | | |
| | `detect` | Image / YOLO / demo → report | | |
| | `live` | Mic capture → report | | |
| | `condition` | Report → diffusion prompt + control vector | | |
| | `serve` | HTTP API `:8765` | | |
| ## HTTP API (`POST /v1/spatial/analyze`) | |
| Modes: `demo_scene` · `ambix_file` · `mono_file` · `vision` · `fuse` · `detect` · `live` · `condition` | |
| ```bash | |
| curl -s -X POST http://127.0.0.1:8765/v1/spatial/analyze \ | |
| -H 'Content-Type: application/json' \ | |
| -d '{"mode":"demo_scene","order":3}' | |
| ``` | |
| An OpenAI-compatible function schema is provided at `tools/spatial_analyze.openai.json`. | |
| ## Agent integration (Qwythos / Pi) | |
| The HOA-7 calculator is **not** part of any LLM's weights. Agents should call it as a tool and treat the returned `one_liner` / `doa_*` / fuse fields as ground-truth geometry: | |
| - Run the CLI: `spatial-report analyze /path/to.wav --ambix -o /tmp/spatial.json && cat /tmp/spatial.json` | |
| - Or HTTP: `POST http://127.0.0.1:8765/v1/spatial/analyze` | |
| - See `integrations/qwythos_system_snippet.md` for the exact agent prompt block. | |
| ## Quick test pack | |
| ```bash | |
| python3 examples/demo_e2e_testpack.py # artifacts in /tmp/spatial_hoa_e2e/ | |
| spatial-report analyze /tmp/spatial_hoa_e2e/scene_ambix4.wav --ambix -o /tmp/a.json | |
| spatial-report detect --demo-image /tmp/frame.png -o /tmp/v.json | |
| spatial-report condition /tmp/spatial_hoa_e2e/fuse_report.json --prompt 'cinematic interior' -o /tmp/c.json | |
| ``` | |
| ## Tests | |
| ```bash | |
| python3 tests/test_basis.py | |
| python3 tests/test_encode_decode.py | |
| python3 tests/test_rotate_rnn.py | |
| python3 tests/test_phase1_audio.py | |
| python3 tests/test_phase2_wigner.py | |
| python3 tests/test_phase3_vision.py | |
| python3 tests/test_integration_extras.py | |
| ``` | |
| ## Limitations | |
| - **Deterministic geometry, not a learned model.** There are no trained weights; accuracy is bounded by the exact spherical-harmonic basis (Farina / Ambix), not by data. The `rnn_stub.py` integrator is an **explicit Euler pose/rotation loop** — learned field dynamics are not implemented yet. | |
| - **No audio synthesis.** This decodes/analyzes spatial fields and produces conditioning; it does not generate audio content. | |
| - **Live capture** requires system `ffmpeg`/`arecord`. | |
| - **Object detection** is optional: the demo path uses synthetic boxes; the real detector downloads torchvision weights on first use. | |
| - **Coordinate convention is Ambix ACN/SN3D.** Interop with other conventions (FuMa, N3D) requires explicit conversion. | |
| ## License | |
| This repository carries no explicit license file. The code is provided as-is by the author; contact `woodfireind` for usage terms. Third-party optional deps (torch/torchvision) keep their own licenses. | |