PaleoSonic_Engine / README.md
livadies's picture
Update README.md
4bfb969 verified
|
Raw
History Blame Contribute Delete
2.82 kB
---
title: PaleoSonic Engine
emoji: 🚀
colorFrom: green
colorTo: indigo
sdk: gradio
sdk_version: 6.11.0
app_file: app.py
pinned: false
license: mit
short_description: Latent-to-Latent Bio-Sonic Engine
---
Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
---
title: PaleoSonic Engine
emoji: 🦴
colorFrom: yellow
colorTo: red
sdk: gradio
app_file: app.py
pinned: false
---
# 🦴 PALEO-SONIC: CHRONO-LATENT ENGINE
**Synthesizing raw acoustic frequencies directly from biological textures using Cross-Modal Latent bridging.**
Welcome to the **PaleoSonic Engine**. This is not an image-to-text-to-audio wrapper. This project performs open-heart surgery on state-of-the-art neural architectures to create a direct **Latent-to-Latent bridge** between machine vision and audio generation.
## 🧬 The Concept
Inspired by archaeoacoustics and synthetic biology, we wanted to know: *What does a fossil sound like? What is the acoustic frequency of Cretaceous period dinosaur skin or a piece of amber?*
Instead of relying on LLMs to "describe" an image and feeding that text to a music generator, we built a custom pipeline that translates the visual geometry of macro-textures directly into sound waves.
## 🛠️ The Architecture Hack (Under the Hood)
We fused the visual cortex of **Google's SigLIP** with the vocal cords of **Meta's MusicGen**:
1. **Vision Encoder:** `google/siglip-base-patch16-224` slices the image into 196 mathematical patches.
2. **The Bridge:** A custom `nn.Linear` layer maps the visual embedding dimensions to the audio conditioning dimensions.
3. **Monkey Patching the Brain:** We dynamically intercept and replace the `text_encoder.forward` method inside MusicGen. The audio decoder *thinks* it is reading 196 tokens of text, but our Trojan function actually injects the raw geometric vectors of the image directly into its cross-attention layers.
The result is pure, zero-shot synesthesia. The audio generated is the literal mathematical interpretation of the image's structure.
## 🕹️ How to Use
1. Upload a macro photograph of a biological or geological texture (e.g., reptile scales, fossilized wood, amber, meteorites).
2. Click **TRANSMUTE GEOMETRY TO SOUND**.
3. Listen to the synthesized biological resonance.
---
### 🎵 Powered by Livadies
This architectural experiment was commissioned to generate the baseline frequencies for the next generation of synthetic music.
**Powered by Livadies. The first artist to synthesize tracks from Cretaceous DNA.**
* 🟢 [Spotify](https://open.spotify.com/artist/0j8EmbhNFjiVhIJcZHdfUD)
* 🔴 [YouTube](https://music.youtube.com/channel/UCe6BJsKd0uj1kAQcdHqyXQw)
* 🟡 [Yandex](https://music.yandex.ru/artist/21918652)
🔥 Active Project Baseline: **«RUSSIAN WINTER 26»**