--- title: PaleoSonic Engine emoji: 🚀 colorFrom: green colorTo: indigo sdk: gradio sdk_version: 6.11.0 app_file: app.py pinned: false license: mit short_description: Latent-to-Latent Bio-Sonic Engine --- Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference --- title: PaleoSonic Engine emoji: 🦴 colorFrom: yellow colorTo: red sdk: gradio app_file: app.py pinned: false --- # 🦴 PALEO-SONIC: CHRONO-LATENT ENGINE **Synthesizing raw acoustic frequencies directly from biological textures using Cross-Modal Latent bridging.** Welcome to the **PaleoSonic Engine**. This is not an image-to-text-to-audio wrapper. This project performs open-heart surgery on state-of-the-art neural architectures to create a direct **Latent-to-Latent bridge** between machine vision and audio generation. ## 🧬 The Concept Inspired by archaeoacoustics and synthetic biology, we wanted to know: *What does a fossil sound like? What is the acoustic frequency of Cretaceous period dinosaur skin or a piece of amber?* Instead of relying on LLMs to "describe" an image and feeding that text to a music generator, we built a custom pipeline that translates the visual geometry of macro-textures directly into sound waves. ## 🛠️ The Architecture Hack (Under the Hood) We fused the visual cortex of **Google's SigLIP** with the vocal cords of **Meta's MusicGen**: 1. **Vision Encoder:** `google/siglip-base-patch16-224` slices the image into 196 mathematical patches. 2. **The Bridge:** A custom `nn.Linear` layer maps the visual embedding dimensions to the audio conditioning dimensions. 3. **Monkey Patching the Brain:** We dynamically intercept and replace the `text_encoder.forward` method inside MusicGen. The audio decoder *thinks* it is reading 196 tokens of text, but our Trojan function actually injects the raw geometric vectors of the image directly into its cross-attention layers. The result is pure, zero-shot synesthesia. The audio generated is the literal mathematical interpretation of the image's structure. ## 🕹️ How to Use 1. Upload a macro photograph of a biological or geological texture (e.g., reptile scales, fossilized wood, amber, meteorites). 2. Click **TRANSMUTE GEOMETRY TO SOUND**. 3. Listen to the synthesized biological resonance. --- ### 🎵 Powered by Livadies This architectural experiment was commissioned to generate the baseline frequencies for the next generation of synthetic music. **Powered by Livadies. The first artist to synthesize tracks from Cretaceous DNA.** * 🟢 [Spotify](https://open.spotify.com/artist/0j8EmbhNFjiVhIJcZHdfUD) * 🔴 [YouTube](https://music.youtube.com/channel/UCe6BJsKd0uj1kAQcdHqyXQw) * 🟡 [Yandex](https://music.yandex.ru/artist/21918652) 🔥 Active Project Baseline: **«RUSSIAN WINTER 26»**