Qapdex's picture
|
download
raw
3.5 kB
# Audio Transcription Pipeline
A modular audio transcription pipeline with speech recognition, audience response classification, speaker diarization, meeting summarization, and ASCII visualization.
## Architecture
```
Audio Input (WAV/FLAC/MP3)
┌────────────────────┐
│ Transcriber │ ← faster-whisper (base model, CPU int8)
│ Word-level timing │ Language detection, beam search
└────────┬───────────┘
┌────────────────────┐
│ Audience Class. │ ← Rule-based + AST (MIT/ast-finetuned-audioset)
│ 10 event classes │ Applause, laughter, cheering, music, etc.
└────────┬───────────┘
┌────────────────────┐
│ Diarizer │ ← pyannote/speaker-diarization-3.1 or mock
│ Speaker labeling │ HF_TOKEN required for full diarization
└────────┬───────────┘
┌────────────────────┐
│ Summarizer │ ← Qwen2.5-0.5B-Instruct GGUF
│ Structured JSON │ Overview, decisions, action items
└────────┬───────────┘
┌────────────────────┐
│ GlyphCast ASCII │ ← Spectrogram → ASCII art
│ 7 charsets/3 modes│ Dark, light, hallow themes
└────────────────────┘
```
## Quick Start
### CLI
```bash
cd /app/audio_transcription_pipeline_1527
# Full transcription
python main.py transcribe data/jfk_speech.wav --output json
# With summary + ASCII visualization
python main.py transcribe data/jfk_speech.wav --summary --ascii
# All pipeline modules at once
python main.py all data/jfk_speech.wav --output json
# Audience response classification only
python main.py audience data/test_tone.wav --output json
# ASCII visualization only
python main.py ascii-viz data/jfk_speech.wav --mode dark
```
### Web Demo (Runs as background service)
- **Frontend (Gradio)**: Port 8080
- **Backend API (FastAPI)**: Port 8081
```bash
# Start backend
python -m uvicorn api:app --host 0.0.0.0 --port 8081 --log-level warning
# Start frontend (in another terminal)
python app.py
```
### API Endpoints
| Endpoint | Method | Description |
|----------|--------|-------------|
| `/health` | GET | Health check |
| `/transcribe` | POST | Upload audio → transcription + analysis |
## Models Used
| Model | Source | Size | Purpose |
|-------|--------|------|---------|
| faster-whisper base | Systran/faster-whisper-base | ~150 MB | Speech recognition |
| AST audioset | MIT/ast-finetuned-audioset-10-10-0.4593 | ~230 MB | Audience classification |
| Qwen2.5-0.5B-Instruct | Qwen/Qwen2.5-0.5B-Instruct-GGUF | ~350 MB | Meeting summarization |
| Whisper-large-v3-turbo | openai/whisper-large-v3-turbo | ~2 GB | Encoder for hybrid model |
| Qwen3-8B-Instruct | Qwen/Qwen3-8B-GGUF | ~5 GB | LLM for hybrid model |
## Hybrid Speech-Transcription Model
See `hybrid_model/ARCHITECTURE.md` for details on the Encoder-Projector-LLM architecture.
## Requirements
- Python 3.10+
- CPU with ~8GB RAM minimum
- Optional: HuggingFace token for pyannote diarization
## License
MIT

Xet Storage Details

Size:
3.5 kB
·
Xet hash:
1e5cdda5c5f2d50dad56d961cbaa5efdd363c1eebc9bc8dbe1fbc3462823b34d

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.