๐๐๏ธ๐ New Research Alert - ICCV 2025 (Poster)! ๐๐๏ธ๐ ๐ Title: Is Less More? Exploring Token Condensation as Training-Free Test-Time Adaptation ๐
๐ Description: Token Condensation as Adaptation (TCA) improves the performance and efficiency of Vision Language Models in zero-shot inference by introducing domain anchor tokens.
๐ฅ Authors: Zixin Wang, Dong Gong, Sen Wang, Zi Huang, Yadan Luo
๐๐๏ธ๐ New Research Alert - ICCV 2025 (Oral)! ๐๐๏ธ๐ ๐ Title: Diving into the Fusion of Monocular Priors for Generalized Stereo Matching ๐
๐ Description: The proposed method enhances stereo matching by efficiently combining unbiased monocular priors from vision foundation models. This method addresses misalignment and local optima issues using a binary local ordering map and pixel-wise linear regression.
๐๐๐ New Research Alert - ICCV 2025 (Oral)! ๐๐ค๐ ๐ Title: Understanding Co-speech Gestures in-the-wild ๐
๐ Description: JEGAL is a tri-modal model that learns from gestures, speech and text simultaneously, enabling devices to interpret co-speech gestures in the wild.
๐ฅ Authors: @sindhuhegde, K R Prajwal, Taein Kwon, and Andrew Zisserman
๐๐ก๐ New Research Alert - ICCV 2025 (Oral)! ๐๐ช๐ ๐ Title: LoftUp: Learning a Coordinate-based Feature Upsampler for Vision Foundation Models ๐
๐ Description: LoftUp is a coordinate-based transformer that upscales the low-resolution features of VFMs (e.g. DINOv2 and CLIP) using cross-attention and self-distilled pseudo-ground truth (pseudo-GT) from SAM.
๐ฅ Authors: Haiwen Huang, Anpei Chen, Volodymyr Havrylov, Andreas Geiger, and Dan Zhang
๐๐ท๏ธ๐ New Research Alert - ICCV 2025 (Oral)! ๐๐งฉ๐ ๐ Title: Heavy Labels Out! Dataset Distillation with Label Space Lightening ๐
๐ Description: The HeLlO framework is a new corpus distillation method that removes the need for large soft labels. It uses a lightweight, online image-to-label projector based on CLIP. This projector has been adapted using LoRA-style, parameter-efficient tuning. It has also been initialized with text embeddings.
๐๐ค๐ New Research Alert - ICCV 2025 (Oral)! ๐๐ค๐ ๐ Title: Variance-based Pruning for Accelerating and Compressing Trained Networks ๐
๐ Description: The one-shot pruning method efficiently compresses networks, reducing computation and memory usage while retaining almost full performance and requiring minimal fine-tuning.
๐ฅ Authors: Uranik Berisha, Jens Mehnert, and Alexandru Paul Condurache
๐๐๏ธ๐ New Research Alert - ICCV 2025 (Oral)! ๐๐๏ธ๐ ๐ Title: Token Activation Map to Visually Explain Multimodal LLMs ๐
๐ Description: The Token Activation Map (TAM) is an advanced explainability method for multimodal LLMs. Using causal inference and a Rank Gaussian Filter, TAM reveals token-level interactions and eliminates redundant activations. The result is clearer, high-quality visualizations that enhance understanding of object localization, reasoning and multimodal alignment across models.
๐ฅ Authors: Yi Li, Hualiang Wang, Xinpeng Ding, Haonan Wang, and Xiaomeng Li