jatshi's picture
Upload folder using huggingface_hub
d8ec09e verified
|
Raw
History Blame Contribute Delete
857 Bytes
---
base_model:
- openai/whisper-small
- Qwen/Qwen2.5-1.5B-Instruct
library_name: pytorch
pipeline_tag: feature-extraction
license: apache-2.0
tags:
- audio
- speech-enhancement
- multimodal
- projector
---
# Audio-Codec-LLM native audio projector (v3 smoke)
`audio_projector.pt` is the trainable Whisper-to-Qwen continuous prefix projector
from the v3 RTX 4090 smoke run. Whisper-small and Qwen2.5-1.5B-Instruct were loaded
as real frozen base models; two optimizer steps reduced the smoke loss from
3.31035 to 2.70892 (mean 3.00964), with 3,859.58 MiB peak VRAM.
This artifact proves the tensor path, backward pass, optimizer and serialization.
Two steps are not a convergence experiment and do not establish enhancement
quality. See `run_manifest.json` and the
[source repository](https://github.com/Jatshi/llm-guided-speech-enhancement).