Cosmos3-Super Image-to-Video (4-Step)
Animate a first frame with Cosmos3-Super I2V (64B, 4-step)
None defined yet.
Animate a first frame with Cosmos3-Super I2V (64B, 4-step)
Web computer-use agent β screenshot + task to next action
Embodied AI VLM for visual grounding and planning
Text-to-image with Microsoft Mage-Flow-Base (4.1B MMDiT)
Persian OCR with Bina 0.1 vision-language model
Temporal reasoning over multi-temporal satellite imagery
3D spatial reasoning VLM with video-diffusion 3D priors
Hierarchical driving world model β predict future frames
Watch & think simultaneously β streaming video reasoning
Dense 1.3B embodied video generation β T2V, I2V, T2I
NVIDIA Cosmos3-Edge β reason over images & video
Video temporal grounding with TimeLens2-2B
Extract medication/drug entities from clinical text
Clinical lab result entity extraction with Bio_ClinicalBERT
PubMedBERT clinical diagnosis entity extraction NER
Play chess against a tiny reasoning LLM that thinks first
Video temporal grounding with TimeLens2-8B
Clinical entity extraction via Bio_ClinicalBERT spaCy NER
World-state visual QA with BAAI Orca-4B
Tiny arithmetic specialist that completes integer equations
VHS / analog camcorder imagery from text via Krea-2-Raw LoRA
ShellD β lightweight DiT text-to-image (256Γ256)
Reasoning-guided region planning for image editing
YOLOv11s aerial object detection on VisDrone