--- language: - en library_name: other pipeline_tag: other license: other tags: - ambisonics - spatial-audio - hoa - spherical-harmonics - audio-analysis - tool - agi-tool - wigner-d --- # HOA7 Spatial Field Decoder (hoa64) A **formula-driven spatial field decoder** for **7th-order Ambisonics** — **64 channels** (Ambix **ACN + SN3D**). It encodes/decodes spherical fields, rotates them with **Wigner-D** matrices, analyzes direction-of-arrival / energy, fuses vision detections onto the sphere, and emits **spatial conditioning** for generative pipelines. Pure NumPy; **not an LLM and not a learned model** — deterministic spherical-harmonic geometry, no weights. | Property | Choice | |----------|--------| | Version | `0.5.0` | | Geometry | Ambix **ACN + SN3D** | | Channels | 64 = (7+1)², order 7 | | Rotation | **Wigner-D** (~150× vs dense) | | Audio | WAV / live mic (ffmpeg·arecord) | | Vision | Boxes / YOLO labels / optional torchvision | | Agents | HTTP `:8765` + `spatial-report` CLI | | Diffusion | Conditioning JSON + optional ComfyUI submit | | Language | Python 3, NumPy only (optional torch/torchvision) | ## What it does - **Encode** point sources / plane waves / scene mixes → HOA-7 coefficients (`encode_points`, `encode_plane_waves`, `encode_scene`). - **Decode** coefficients → samples on the sphere, product-grid render, or a single beamform readout (`decode_directions`, `decode_grid`, `beamform`). - **Rotate** the field in the listener frame via Wigner-D / zyz (`hoa_rotation_matrix`, `apply_hoa_rotation`, `rotate_yaw_pitch_roll`). - **Analyze** DOA (intensity vector + peak search), directional power, field energy, per-STFT-band and per-frame reports (`analysis.py`, `report.py`). - **Fuse vision** — object boxes / rays / YOLO detections projected onto the sphere alongside audio (`vision.py`, `detector.py`, `fuse_reports`). - **Condition** — spatial reports → plain-text control lines for T2I/T2V prompts, structured JSON for ControlNet-style nodes, or a ComfyUI API payload (`conditioning.py`). - **Serve** — HTTP API and iterative agent state (`server.py`, `rnn_stub.py`). ## Coordinates (always) - **+X** front, **+Y** left, **+Z** up (Ambix listener frame). - Azimuth 0° = front, **+90° = left**, −90° = right. - Elevation 0° = horizon, +90° = zenith. - W (omnidirectional HOA channel) maps to field size / POV: low W → tight / subject-focused / narrow FOV; high W → wide / environmental / immersive FOV. ## Install / run ```bash cd spatial-hoa # this repo root export PYTHONPATH="$PWD${PYTHONPATH:+:$PYTHONPATH}" python3 -m hoa64 --help # or install the CLI on PATH: ln -s "$PWD/scripts/spatial-report" ~/.local/bin/spatial-report ``` **Deps:** NumPy (system `ffmpeg`/`arecord` for live capture; optional torch/torchvision for the real object detector). ## CLI map (`spatial-report`) | Command | Purpose | |---------|---------| | `analyze` | Ambix / mono WAV → spatial JSON | | `demo-scene` | Synthetic multi-source audio | | `vision` | Raw sphere boxes → report | | `detect` | Image / YOLO / demo → report | | `live` | Mic capture → report | | `condition` | Report → diffusion prompt + control vector | | `serve` | HTTP API `:8765` | ## HTTP API (`POST /v1/spatial/analyze`) Modes: `demo_scene` · `ambix_file` · `mono_file` · `vision` · `fuse` · `detect` · `live` · `condition` ```bash curl -s -X POST http://127.0.0.1:8765/v1/spatial/analyze \ -H 'Content-Type: application/json' \ -d '{"mode":"demo_scene","order":3}' ``` An OpenAI-compatible function schema is provided at `tools/spatial_analyze.openai.json`. ## Agent integration (Qwythos / Pi) The HOA-7 calculator is **not** part of any LLM's weights. Agents should call it as a tool and treat the returned `one_liner` / `doa_*` / fuse fields as ground-truth geometry: - Run the CLI: `spatial-report analyze /path/to.wav --ambix -o /tmp/spatial.json && cat /tmp/spatial.json` - Or HTTP: `POST http://127.0.0.1:8765/v1/spatial/analyze` - See `integrations/qwythos_system_snippet.md` for the exact agent prompt block. ## Quick test pack ```bash python3 examples/demo_e2e_testpack.py # artifacts in /tmp/spatial_hoa_e2e/ spatial-report analyze /tmp/spatial_hoa_e2e/scene_ambix4.wav --ambix -o /tmp/a.json spatial-report detect --demo-image /tmp/frame.png -o /tmp/v.json spatial-report condition /tmp/spatial_hoa_e2e/fuse_report.json --prompt 'cinematic interior' -o /tmp/c.json ``` ## Tests ```bash python3 tests/test_basis.py python3 tests/test_encode_decode.py python3 tests/test_rotate_rnn.py python3 tests/test_phase1_audio.py python3 tests/test_phase2_wigner.py python3 tests/test_phase3_vision.py python3 tests/test_integration_extras.py ``` ## Limitations - **Deterministic geometry, not a learned model.** There are no trained weights; accuracy is bounded by the exact spherical-harmonic basis (Farina / Ambix), not by data. The `rnn_stub.py` integrator is an **explicit Euler pose/rotation loop** — learned field dynamics are not implemented yet. - **No audio synthesis.** This decodes/analyzes spatial fields and produces conditioning; it does not generate audio content. - **Live capture** requires system `ffmpeg`/`arecord`. - **Object detection** is optional: the demo path uses synthetic boxes; the real detector downloads torchvision weights on first use. - **Coordinate convention is Ambix ACN/SN3D.** Interop with other conventions (FuMa, N3D) requires explicit conversion. ## License This repository carries no explicit license file. The code is provided as-is by the author; contact `woodfireind` for usage terms. Third-party optional deps (torch/torchvision) keep their own licenses.