AI & ML interests

None defined yet.

Recent Activity

DmitryRyuminย 
posted an update 9 months ago
view post
Post
1385
๐Ÿš€๐Ÿ‘๏ธ๐ŸŒŸ New Research Alert - ICCV 2025 (Poster)! ๐ŸŒŸ๐Ÿ‘๏ธ๐Ÿš€
๐Ÿ“„ Title: Is Less More? Exploring Token Condensation as Training-Free Test-Time Adaptation ๐Ÿ”

๐Ÿ“ Description: Token Condensation as Adaptation (TCA) improves the performance and efficiency of Vision Language Models in zero-shot inference by introducing domain anchor tokens.

๐Ÿ‘ฅ Authors: Zixin Wang, Dong Gong, Sen Wang, Zi Huang, Yadan Luo

๐Ÿ“… Conference: ICCV, 19 โ€“ 23 Oct, 2025 | Honolulu, Hawai'i, USA ๐Ÿ‡บ๐Ÿ‡ธ

๐Ÿ“„ Paper: Is Less More? Exploring Token Condensation as Training-free Test-time Adaptation (2410.14729)

๐Ÿ“ Repository: https://github.com/Jo-wang/TCA

๐Ÿš€ ICCV-2023-25-Papers: https://github.com/DmitryRyumin/ICCV-2023-25-Papers

๐Ÿš€ Added to the Session 1: https://github.com/DmitryRyumin/ICCV-2023-25-Papers/blob/main/sections/2025/main/session-1.md

๐Ÿ“š More Papers: more cutting-edge research presented at other conferences in the DmitryRyumin/NewEraAI-Papers curated by @DmitryRyumin

๐Ÿ” Keywords: #TestTimeAdaptation #TokenCondensation #VisionLanguageModels #TrainingFreeAdaptation #ZeroShotLearning #EfficientAI #AI #ICCV2025 #ResearchHighlight
DmitryRyuminย 
posted an update 9 months ago
view post
Post
2518
๐Ÿš€๐Ÿ‘๏ธ๐ŸŒŸ New Research Alert - ICCV 2025 (Oral)! ๐ŸŒŸ๐Ÿ‘๏ธ๐Ÿš€
๐Ÿ“„ Title: Diving into the Fusion of Monocular Priors for Generalized Stereo Matching ๐Ÿ”

๐Ÿ“ Description: The proposed method enhances stereo matching by efficiently combining unbiased monocular priors from vision foundation models. This method addresses misalignment and local optima issues using a binary local ordering map and pixel-wise linear regression.

๐Ÿ‘ฅ Authors: Chengtang Yao, Lidong Yu, Zhidan Liu, Jiaxi Zeng, Yuwei Wu, and Yunde Jia

๐Ÿ“… Conference: ICCV, 19 โ€“ 23 Oct, 2025 | Honolulu, Hawai'i, USA ๐Ÿ‡บ๐Ÿ‡ธ

๐Ÿ“„ Paper: Diving into the Fusion of Monocular Priors for Generalized Stereo Matching (2505.14414)

๐Ÿ“ Repository: https://github.com/YaoChengTang/Diving-into-the-Fusion-of-Monocular-Priors-for-Generalized-Stereo-Matching

๐Ÿš€ ICCV-2023-25-Papers: https://github.com/DmitryRyumin/ICCV-2023-25-Papers

๐Ÿš€ Added to the 3D Pose Understanding Section: https://github.com/DmitryRyumin/ICCV-2023-25-Papers/blob/main/sections/2025/main/3d-pose-understanding.md

๐Ÿ“š More Papers: more cutting-edge research presented at other conferences in the DmitryRyumin/NewEraAI-Papers curated by @DmitryRyumin

๐Ÿ” Keywords: #StereoMatching #MonocularDepth #VisionFoundationModels #3DReconstruction #Generalization #AI #ICCV2025 #ResearchHighlight
DmitryRyuminย 
posted an update 9 months ago
view post
Post
2849
๐Ÿš€๐Ÿ‘Œ๐ŸŒŸ New Research Alert - ICCV 2025 (Oral)! ๐ŸŒŸ๐ŸคŒ๐Ÿš€
๐Ÿ“„ Title: Understanding Co-speech Gestures in-the-wild ๐Ÿ”

๐Ÿ“ Description: JEGAL is a tri-modal model that learns from gestures, speech and text simultaneously, enabling devices to interpret co-speech gestures in the wild.

๐Ÿ‘ฅ Authors: @sindhuhegde , K R Prajwal, Taein Kwon, and Andrew Zisserman

๐Ÿ“… Conference: ICCV, 19 โ€“ 23 Oct, 2025 | Honolulu, Hawai'i, USA ๐Ÿ‡บ๐Ÿ‡ธ

๐Ÿ“„ Paper: Understanding Co-speech Gestures in-the-wild (2503.22668)

๐ŸŒ Web Page: https://www.robots.ox.ac.uk/~vgg/research/jegal
๐Ÿ“ Repository: https://github.com/Sindhu-Hegde/jegal
๐Ÿ“บ Video: https://www.youtube.com/watch?v=TYFOLKfM-rM

๐Ÿš€ ICCV-2023-25-Papers: https://github.com/DmitryRyumin/ICCV-2023-25-Papers

๐Ÿš€ Added to the Human Modeling Section: https://github.com/DmitryRyumin/ICCV-2023-25-Papers/blob/main/sections/2025/main/human-modeling.md

๐Ÿ“š More Papers: more cutting-edge research presented at other conferences in the DmitryRyumin/NewEraAI-Papers curated by @DmitryRyumin

๐Ÿ” Keywords: #CoSpeechGestures #GestureUnderstanding #TriModalRepresentation #MultimodalLearning #AI #ICCV2025 #ResearchHighlight
DmitryRyuminย 
posted an update 9 months ago
view post
Post
3991
๐Ÿš€๐Ÿ’ก๐ŸŒŸ New Research Alert - ICCV 2025 (Oral)! ๐ŸŒŸ๐Ÿช„๐Ÿš€
๐Ÿ“„ Title: LoftUp: Learning a Coordinate-based Feature Upsampler for Vision Foundation Models ๐Ÿ”

๐Ÿ“ Description: LoftUp is a coordinate-based transformer that upscales the low-resolution features of VFMs (e.g. DINOv2 and CLIP) using cross-attention and self-distilled pseudo-ground truth (pseudo-GT) from SAM.

๐Ÿ‘ฅ Authors: Haiwen Huang, Anpei Chen, Volodymyr Havrylov, Andreas Geiger, and Dan Zhang

๐Ÿ“… Conference: ICCV, 19 โ€“ 23 Oct, 2025 | Honolulu, Hawai'i, USA ๐Ÿ‡บ๐Ÿ‡ธ

๐Ÿ“„ Paper: LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models (2504.14032)

๐ŸŒ Github Page: https://andrehuang.github.io/loftup-site
๐Ÿ“ Repository: https://github.com/andrehuang/loftup

๐Ÿš€ ICCV-2023-25-Papers: https://github.com/DmitryRyumin/ICCV-2023-25-Papers

๐Ÿš€ Added to the Foundation Models and Representation Learning Section: https://github.com/DmitryRyumin/ICCV-2023-25-Papers/blob/main/sections/2025/main/foundation-models-and-representation-learning.md

๐Ÿ“š More Papers: more cutting-edge research presented at other conferences in the DmitryRyumin/NewEraAI-Papers curated by @DmitryRyumin

๐Ÿ” Keywords: #LoftUp #VisionFoundationModels #FeatureUpsampling #Cross-AttentionTransformer #CoordinateBasedLearning #SelfDistillation #PseudoGroundTruth #RepresentationLearning #AI #ICCV2025 #ResearchHighlight
DmitryRyuminย 
posted an update 9 months ago
view post
Post
1972
๐Ÿš€๐Ÿท๏ธ๐ŸŒŸ New Research Alert - ICCV 2025 (Oral)! ๐ŸŒŸ๐Ÿงฉ๐Ÿš€
๐Ÿ“„ Title: Heavy Labels Out! Dataset Distillation with Label Space Lightening ๐Ÿ”

๐Ÿ“ Description: The HeLlO framework is a new corpus distillation method that removes the need for large soft labels. It uses a lightweight, online image-to-label projector based on CLIP. This projector has been adapted using LoRA-style, parameter-efficient tuning. It has also been initialized with text embeddings.

๐Ÿ‘ฅ Authors: @roseannelexie , @Huage001 , Zigeng Chen, Jingwen Ye, and Xinchao Wang

๐Ÿ“… Conference: ICCV, 19 โ€“ 23 Oct, 2025 | Honolulu, Hawai'i, USA ๐Ÿ‡บ๐Ÿ‡ธ

๐Ÿ“„ Paper: Heavy Labels Out! Dataset Distillation with Label Space Lightening (2408.08201)

๐Ÿ“บ Video: https://www.youtube.com/watch?v=kAyK_3wskgA

๐Ÿš€ ICCV-2023-25-Papers: https://github.com/DmitryRyumin/ICCV-2023-25-Papers

๐Ÿš€ Added to the Efficient Learning Section: https://github.com/DmitryRyumin/ICCV-2023-25-Papers/blob/main/sections/2025/main/efficient-learning.md

๐Ÿ“š More Papers: more cutting-edge research presented at other conferences in the DmitryRyumin/NewEraAI-Papers curated by @DmitryRyumin

๐Ÿ” Keywords: #DatasetDistillation #LabelCompression #CLIP #LoRA #EfficientAI #FoundationModels #AI #ICCV2025 #ResearchHighlight
  • 2 replies
ยท
DmitryRyuminย 
posted an update 9 months ago
view post
Post
4829
๐Ÿš€๐Ÿค–๐ŸŒŸ New Research Alert - ICCV 2025 (Oral)! ๐ŸŒŸ๐Ÿค–๐Ÿš€
๐Ÿ“„ Title: Variance-based Pruning for Accelerating and Compressing Trained Networks ๐Ÿ”

๐Ÿ“ Description: The one-shot pruning method efficiently compresses networks, reducing computation and memory usage while retaining almost full performance and requiring minimal fine-tuning.

๐Ÿ‘ฅ Authors: Uranik Berisha, Jens Mehnert, and Alexandru Paul Condurache

๐Ÿ“… Conference: ICCV, 19 โ€“ 23 Oct, 2025 | Honolulu, Hawai'i, USA ๐Ÿ‡บ๐Ÿ‡ธ

๐Ÿ“„ Paper: Variance-Based Pruning for Accelerating and Compressing Trained Networks (2507.12988)

๐Ÿš€ ICCV-2023-25-Papers: https://github.com/DmitryRyumin/ICCV-2023-25-Papers

๐Ÿš€ Added to the Efficient Learning Section: https://github.com/DmitryRyumin/ICCV-2023-25-Papers/blob/main/sections/2025/main/efficient-learning.md

๐Ÿ“š More Papers: more cutting-edge research presented at other conferences in the DmitryRyumin/NewEraAI-Papers curated by @DmitryRyumin

๐Ÿ” Keywords: #VarianceBasedPruning #NetworkCompression #ModelAcceleration #EfficientDeepLearning #VisionTransformers #AI #ICCV2025 #ResearchHighlight
DmitryRyuminย 
posted an update 10 months ago
view post
Post
3042
๐Ÿš€๐Ÿ‘๏ธ๐ŸŒŸ New Research Alert - ICCV 2025 (Oral)! ๐ŸŒŸ๐Ÿ‘๏ธ๐Ÿš€
๐Ÿ“„ Title: Token Activation Map to Visually Explain Multimodal LLMs ๐Ÿ”

๐Ÿ“ Description: The Token Activation Map (TAM) is an advanced explainability method for multimodal LLMs. Using causal inference and a Rank Gaussian Filter, TAM reveals token-level interactions and eliminates redundant activations. The result is clearer, high-quality visualizations that enhance understanding of object localization, reasoning and multimodal alignment across models.

๐Ÿ‘ฅ Authors: Yi Li, Hualiang Wang, Xinpeng Ding, Haonan Wang, and Xiaomeng Li

๐Ÿ“… Conference: ICCV, 19 โ€“ 23 Oct, 2025 | Honolulu, Hawai'i, USA ๐Ÿ‡บ๐Ÿ‡ธ

๐Ÿ“„ Paper: Token Activation Map to Visually Explain Multimodal LLMs (2506.23270)

๐Ÿ“ Repository: https://github.com/xmed-lab/TAM

๐Ÿš€ ICCV-2023-25-Papers: https://github.com/DmitryRyumin/ICCV-2023-25-Papers

๐Ÿš€ Added to the Multi-Modal Learning Section: https://github.com/DmitryRyumin/ICCV-2023-25-Papers/blob/main/sections/2025/main/multi-modal-learning.md

๐Ÿ“š More Papers: more cutting-edge research presented at other conferences in the DmitryRyumin/NewEraAI-Papers curated by @DmitryRyumin

๐Ÿ” Keywords: #TokenActivationMap #TAM #CausalInference #VisualReasoning #Multimodal #Explainability #VisionLanguage #LLM #XAI #AI #ICCV2025 #ResearchHighlight
  • 2 replies
ยท

Use Space Secrets

#2 opened over 1 year ago by
freddyaboulton