A newer version of the Gradio SDK is available: 6.24.0
metadata
title: Multimodal Feature Extractor
emoji: 🎯
colorFrom: purple
colorTo: blue
sdk: gradio
sdk_version: 5.9.1
app_file: app.py
pinned: false
license: mit
short_description: Extract embeddings from video, image, GIF, and text
Mixpeek Multimodal Feature Extractor
Extract embeddings and features from videos, images, GIFs, and text using Mixpeek's multimodal AI pipeline.
Features
- Video Processing: Decompose videos into segments using time, scene, or silence-based splitting
- Image & GIF Support: Direct embedding generation for static and animated images
- Text Embedding: Generate embeddings from plain text
- Transcription: Whisper-powered speech-to-text (95%+ accuracy)
- OCR: Extract text from video frames and images
- AI Descriptions: Generate natural language descriptions of content
- 1408D Embeddings: Google Vertex AI multimodal vectors for semantic search
Supported Formats
| Type | Formats |
|---|---|
| Video | MP4, MOV, AVI, MKV, WebM, FLV |
| Image | JPG, PNG, WebP, BMP |
| GIF | Animated and static GIFs |
| Text | Plain text strings |
Usage
- Get your API key from mixpeek.com
- Choose your input type (file upload or text)
- Configure extraction settings
- Click "Extract Features" to process
Output
Each extracted segment contains:
- Timing data (start/end times for video segments)
- Transcription text
- 1408D multimodal embedding vector
- OCR extracted text (if enabled)
- AI-generated description (if enabled)
- Thumbnail URL (if enabled)