DiffusionGemma Radiology VQA
Radiology VQA & report infill with a discrete-diffusion LLM
None defined yet.
Radiology VQA & report infill with a discrete-diffusion LLM
Predict robot action chunks from an image + instruction
Chat with fable-traces, a Qwen3-4B-Instruct finetune
Khmer speech recognition with Qwen3-ASR-0.6B
Parallel region captioning with multimodal diffusion LLM
2x latent super-resolution with FlowUpscaler in Flux.2 space
Text-to-image with SeFi-Image-5B Semantic-First Diffusion
Keep identity from reference, follow lineart structure
Music understanding model for caption and analysis
Text/speech to spoken response + 3D talking-avatar video
Word-level timestamp alignment from audio + transcript
Multi-modal generation with diffusion transformers
Polish speech recognition with fine-tuned Whisper Small
Real-time zero-shot stereo disparity estimation
Phone-use GUI agent - screenshot + task to next action
GUI grounding with VISTA-9B — predict click coordinates
Multi-view visual reasoning VLM based on Qwen3-VL 4B
Object and Material Selection VLM
Document-parsing VLM (1.2B) by KoreaDeep
Vietnamese text-to-speech with Kokoro TTS
Interleaved text and image generation with SenseNova-U1
Scientific object generator (molecules, proteins, materials)
Real-time audio-visual social world model (22B)