Collections
Discover the best community collections!
Collections trending this week
-
LocalMamba: Visual State Space Model with Windowed Selective Scan
Paper β’ 2403.09338 β’ Published β’ 8 -
GiT: Towards Generalist Vision Transformer through Universal Language Interface
Paper β’ 2403.09394 β’ Published β’ 26 -
Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
Paper β’ 2402.19479 β’ Published β’ 35 -
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
Paper β’ 2405.10300 β’ Published β’ 31
-
Vision Transformer with Quadrangle Attention
Paper β’ 2303.15105 β’ Published β’ 2 -
Language Grounded QFormer for Efficient Vision Language Understanding
Paper β’ 2311.07449 β’ Published β’ 2 -
MultiBooth: Towards Generating All Your Concepts in an Image from Text
Paper β’ 2404.14239 β’ Published β’ 9
-
LocalMamba: Visual State Space Model with Windowed Selective Scan
Paper β’ 2403.09338 β’ Published β’ 8 -
GiT: Towards Generalist Vision Transformer through Universal Language Interface
Paper β’ 2403.09394 β’ Published β’ 26 -
Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
Paper β’ 2402.19479 β’ Published β’ 35 -
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
Paper β’ 2405.10300 β’ Published β’ 31
-
Vision Transformer with Quadrangle Attention
Paper β’ 2303.15105 β’ Published β’ 2 -
Language Grounded QFormer for Efficient Vision Language Understanding
Paper β’ 2311.07449 β’ Published β’ 2 -
MultiBooth: Towards Generating All Your Concepts in an Image from Text
Paper β’ 2404.14239 β’ Published β’ 9