Running on Zero MCP Featured 47 LTX-Best-Face-ID 🎬 47 Distilled LTX-2.3 identity video from a reference photo
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Paper • 2607.11562 • Published about 1 month ago • 78
OpenMOSS-Team/MOSS-Transcribe-Diarize Audio-Text-to-Text • 0.9B • Updated 12 days ago • 194k • 376
Running on Zero MCP Featured 1.59k FireRed Image Edit 1.0 Fast 🔥 1.59k FireRed-Image-Edit × Qwen-Image-Edit-Rapid (Transformers)
HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents Paper • 2604.07430 • Published Apr 8 • 181
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents Paper • 2604.11784 • Published Apr 13 • 143
HauhauCS/Gemma-4-E4B-Uncensored-HauhauCS-Aggressive Image-Text-to-Text • 8B • Updated Apr 6 • 921k • 980
Cephalo Collection Cephalo is a series of multimodal vision large language models (V-LLMs) designed to integrate visual and linguistic reasoning in materials science. • 17 items • Updated Apr 16, 2025 • 5