Spaces:
Runtime error
Runtime error
A newer version of the Gradio SDK is available: 6.22.0
metadata
title: Video Intelligence Chat
emoji: π¬
colorFrom: blue
colorTo: purple
sdk: gradio
sdk_version: 4.44.1
python_version: '3.10'
app_file: app.py
pinned: false
Video Intelligence Chat
Chat with an AI about any video you upload. The app extracts frames, transcribes speech, and analyzes content using free-tier APIs β all running on Hugging Face Spaces CPU Basic hardware.
One-Click Deploy to Hugging Face Spaces
- Go to huggingface.co/new-space
- Set SDK to
Gradioand Hardware toCPU Basic(free) - Clone your new Space and copy all project files in:
git clone https://huggingface.co/spaces/YOUR_USERNAME/YOUR_SPACE_NAME
cd YOUR_SPACE_NAME
cp -r /path/to/video-intelligence-chat/* .
git add .
git commit -m "Initial deploy"
git push
- Add your secret:
- Go to Space Settings β Variables and Secrets
- Click New Secret:
NVIDIA_API_KEY=<your key>
The Space rebuilds automatically and is live in ~2 minutes.
Getting a Free NVIDIA API Key
- Sign up at build.nvidia.com (free account)
- Navigate to any model page and click Get API Key
- Your free account includes ~1,000 API credits
Required Secret
| Secret name | Description |
|---|---|
NVIDIA_API_KEY |
From build.nvidia.com |
Hardware
- CPU Basic (free tier) is sufficient
- 16 GB RAM, 2 vCPU
faster-whisperruns locally on Space CPU β no GPU needed
Optional Environment Variables
| Variable | Default | Description |
|---|---|---|
FRAME_INTERVAL |
5 |
Seconds between extracted frames |
WHISPER_MODEL |
base |
Whisper model size: tiny, base, small |
MAX_VIDEO_DURATION |
600 |
Warn if video exceeds this many seconds |
Recommended Usage
- Video length: Under 5 minutes on the free NVIDIA tier (to conserve API credits)
- Supported formats: MP4, MOV, AVI, WebM
- Max upload size: ~100 MB (Gradio default)
Free Tier Limits
| Service | Free Allowance |
|---|---|
| NVIDIA NIM | ~1,000 API credits on signup |
| Hugging Face Spaces | Unlimited CPU Basic spaces |
| faster-whisper (local) | No limits β runs on Space hardware |
| HF Space RAM | 16 GB RAM, 2 vCPU |
Tech Stack
| Layer | Tool |
|---|---|
| Frontend / Backend | Gradio (Python) |
| Hosting | Hugging Face Spaces |
| Vision AI | NVIDIA NIM β meta/llama-3.2-11b-vision-instruct |
| Chat LLM | NVIDIA NIM β meta/llama-3.1-8b-instruct |
| Speech-to-text | faster-whisper (CPU, runs locally) |
| Video processing | ffmpeg + OpenCV |