FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation Paper • 2609.16591 • Published 12 days ago • 16
Diagnosing Visual Reasoning: Challenges, Insights, and a Path Forward Paper • 2510.20696 • Published Oct 23, 2025
From Frames to Clips: Efficient Key Clip Selection for Long-Form Video Understanding Paper • 2510.02262 • Published Oct 2, 2025 • 3
FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation Paper • 2609.16591 • Published 12 days ago • 16
FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation Paper • 2609.16591 • Published 12 days ago • 16
Seedance 2.0: Advancing Video Generation for World Complexity Paper • 2604.14148 • Published Apr 15 • 171
VIDEOP2R: Video Understanding from Perception to Reasoning Paper • 2511.11113 • Published Nov 14, 2025 • 113
Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark Paper • 2511.13853 • Published Nov 17, 2025 • 37
VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity Paper • 2503.11557 • Published Mar 14, 2025 • 22
FedPerfix: Towards Partial Model Personalization of Vision Transformers in Federated Learning Paper • 2308.09160 • Published Aug 17, 2023
Exploring Parameter-Efficient Fine-Tuning to Enable Foundation Models in Federated Learning Paper • 2210.01708 • Published Oct 4, 2022
ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback Paper • 2404.07987 • Published Apr 11, 2024 • 49