jatshi's picture
Upload folder using huggingface_hub
d8ec09e verified
|
Raw
History Blame Contribute Delete
857 Bytes
metadata
base_model:
  - openai/whisper-small
  - Qwen/Qwen2.5-1.5B-Instruct
library_name: pytorch
pipeline_tag: feature-extraction
license: apache-2.0
tags:
  - audio
  - speech-enhancement
  - multimodal
  - projector

Audio-Codec-LLM native audio projector (v3 smoke)

audio_projector.pt is the trainable Whisper-to-Qwen continuous prefix projector from the v3 RTX 4090 smoke run. Whisper-small and Qwen2.5-1.5B-Instruct were loaded as real frozen base models; two optimizer steps reduced the smoke loss from 3.31035 to 2.70892 (mean 3.00964), with 3,859.58 MiB peak VRAM.

This artifact proves the tensor path, backward pass, optimizer and serialization. Two steps are not a convergence experiment and do not establish enhancement quality. See run_manifest.json and the source repository.