HuggingFaceTB/SmolVLM-500M-Instruct
Image-Text-to-Text • 0.5B • Updated • 149k • 198
Transcribe audio or YouTube video into text
Generate videos from text prompts (optional image guidance)
Free reverse face search — find anyone by photo
Generate a 3D mesh from a single image
Generate and preview app code from a text description
flux.1-dev / flux.1-krea-dev
Import a portrait, click to move the head!
Chat with Mini-Omni 2 - powered by Gradio and WebRTC ⚡️
Add vectors to Hub datasets and do in memory vector search.
An end-to-end (e2e) Voice Language Model by Fish Audio.