-
Ai Youtube Assistant
🔥2Generate summarized text from YouTube video transcripts and ask questions
-
lucianosb/llama-2-7b-langchain-chat-GGUF
Text Generation • 7B • Updated • 417 • 12 -
deepseek-ai/DeepSeek-V3
Text Generation • 685B • Updated • 1.07M • • 4.18k -
deepseek-ai/deepseek-vl-1.3b-chat
Image-Text-to-Text • 2B • Updated • 7.42k • 73
Collections
Discover the best community collections!
Collections trending this week
-
LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
Paper • 2404.01331 • Published • 28 -
OmniFusion Technical Report
Paper • 2404.06212 • Published • 78 -
MoDE: CLIP Data Experts via Clustering
Paper • 2404.16030 • Published • 14 -
WildGaussians: 3D Gaussian Splatting in the Wild
Paper • 2407.08447 • Published • 9
-
LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
Paper • 2404.01331 • Published • 28 -
Data curation via joint example selection further accelerates multimodal learning
Paper • 2406.17711 • Published • 3 -
Unveiling Encoder-Free Vision-Language Models
Paper • 2406.11832 • Published • 55
-
LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
Paper • 2404.01331 • Published • 28 -
LVLM-Intrepret: An Interpretability Tool for Large Vision-Language Models
Paper • 2404.03118 • Published • 25 -
DesignQA: A Multimodal Benchmark for Evaluating Large Language Models' Understanding of Engineering Documentation
Paper • 2404.07917 • Published • 3 -
Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Paper • 2404.07973 • Published • 33
-
Ai Youtube Assistant
🔥2Generate summarized text from YouTube video transcripts and ask questions
-
lucianosb/llama-2-7b-langchain-chat-GGUF
Text Generation • 7B • Updated • 417 • 12 -
deepseek-ai/DeepSeek-V3
Text Generation • 685B • Updated • 1.07M • • 4.18k -
deepseek-ai/deepseek-vl-1.3b-chat
Image-Text-to-Text • 2B • Updated • 7.42k • 73
-
LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
Paper • 2404.01331 • Published • 28 -
OmniFusion Technical Report
Paper • 2404.06212 • Published • 78 -
MoDE: CLIP Data Experts via Clustering
Paper • 2404.16030 • Published • 14 -
WildGaussians: 3D Gaussian Splatting in the Wild
Paper • 2407.08447 • Published • 9
-
LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
Paper • 2404.01331 • Published • 28 -
Data curation via joint example selection further accelerates multimodal learning
Paper • 2406.17711 • Published • 3 -
Unveiling Encoder-Free Vision-Language Models
Paper • 2406.11832 • Published • 55
-
LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
Paper • 2404.01331 • Published • 28 -
LVLM-Intrepret: An Interpretability Tool for Large Vision-Language Models
Paper • 2404.03118 • Published • 25 -
DesignQA: A Multimodal Benchmark for Evaluating Large Language Models' Understanding of Engineering Documentation
Paper • 2404.07917 • Published • 3 -
Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Paper • 2404.07973 • Published • 33