Daniel0902 's Collections
nvidia/Llama-Nemotron-VLM-Dataset-v1
Viewer
• Updated • 2.86M • 4.42k
• 168
Updated • 117
• 5
xintongzhang/CoF-SFT-Data-5.4k
Preview
• Updated • 102
• 2
Viewer
• Updated • 17.6k • 5.23k
• 38
Viewer
• Updated • 112k • 67
HuggingFaceTB/SmolVLM-500M-Instruct
Image-Text-to-Text
• 0.5B • Updated • 101k
• 197
HuggingFaceTB/SmolVLM-256M-Instruct
Image-Text-to-Text
• 0.3B • Updated • 463k
• 403
HuggingFaceTB/SmolVLM2-256M-Video-Instruct
Image-Text-to-Text
• 0.3B • Updated • 55.2k
• 114
microsoft/Florence-2-base
Image-Text-to-Text
• 0.2B • Updated • 2.77M
• 397
Text Generation
• 1B • Updated • 10.1k
• 154
Image-Text-to-Text
• 1B • Updated • 358
• 112
1B • Updated • 98
0.7B • Updated • 28
• 1
Robotics
• 8B • Updated • 441k
• 254
OpenGVLab/InternVL2_5-78B
Image-Text-to-Text
• 78B • Updated • 596
• 193
sd2-community/stable-diffusion-2-1
Text-to-Image
• 0.9B • Updated • 11.2k
• 35
Efficient-Large-Model/SANA-WM_bidirectional
Image-to-Video
• Updated • 128
95.6M • Updated • 455k
• 85
Robotics
• 0.9B • Updated • 3.23k
• 23
Viewer
• Updated • 495 • 1.48k
• 5
Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with
Long-Term Memory
Paper
• 2508.09736
• Published • 58
Viewer
• Updated • 146 • 9.74k
• 39