--- title: Multimodal Feature Extractor emoji: 🎯 colorFrom: purple colorTo: blue sdk: gradio sdk_version: 5.9.1 app_file: app.py pinned: false license: mit short_description: Extract embeddings from video, image, GIF, and text --- # Mixpeek Multimodal Feature Extractor Extract embeddings and features from **videos**, **images**, **GIFs**, and **text** using Mixpeek's multimodal AI pipeline. ## Features - **Video Processing**: Decompose videos into segments using time, scene, or silence-based splitting - **Image & GIF Support**: Direct embedding generation for static and animated images - **Text Embedding**: Generate embeddings from plain text - **Transcription**: Whisper-powered speech-to-text (95%+ accuracy) - **OCR**: Extract text from video frames and images - **AI Descriptions**: Generate natural language descriptions of content - **1408D Embeddings**: Google Vertex AI multimodal vectors for semantic search ## Supported Formats | Type | Formats | |------|---------| | Video | MP4, MOV, AVI, MKV, WebM, FLV | | Image | JPG, PNG, WebP, BMP | | GIF | Animated and static GIFs | | Text | Plain text strings | ## Usage 1. Get your API key from [mixpeek.com](https://mixpeek.com) 2. Choose your input type (file upload or text) 3. Configure extraction settings 4. Click "Extract Features" to process ## Output Each extracted segment contains: - Timing data (start/end times for video segments) - Transcription text - 1408D multimodal embedding vector - OCR extracted text (if enabled) - AI-generated description (if enabled) - Thumbnail URL (if enabled) ## Links - [Mixpeek Documentation](https://docs.mixpeek.com) - [API Reference](https://docs.mixpeek.com/api-reference) - [Get API Key](https://mixpeek.com)