Varun2007's picture
initial clean deployment commit with compilers
8c3e275
|
Raw
History Blame Contribute Delete
539 Bytes
# Feature Spec: Local Media Explainer
## Goal
To extract highly detailed context, summaries, and structural walkthroughs from uploaded images, videos, and audio memos using offline LLMs.
## Requirements
1. **Vision Integration**: Call local `moondream`/`llava` vision APIs when images are uploaded.
2. **Audio Processing**: Call local SLM `llama3.2:1b` on audio text transcripts to formulate contextual summaries.
3. **Structured Storage**: Save outputs in raw text structures in SQLite and render them directly to the user's interface.