FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation Paper • 2609.11486 • Published 12 days ago • 34
Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts Paper • 2608.20061 • Published Aug 20 • 47
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning Paper • 2608.09888 • Published Aug 10 • 787
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published Jul 16 • 120
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 165
TIPSv2 Collection TIPSv2 foundational vision-language models. Webpage: https://gdm-tipsv2.github.io/ • 9 items • Updated Jul 21 • 44
RDP LoRA: Geometry-Driven Identification for Parameter-Efficient Adaptation in Large Language Models Paper • 2604.19321 • Published Apr 21 • 8
SigLino: Vision Foundation Models (SigLIP2 + DINOv3) Collection Vision encoders distilled from DINOv3 and SigLIP2 (MoE & Dense). CVPR 2026. • 6 items • Updated Aug 13 • 17
Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking Paper • 2601.04720 • Published Jan 8 • 59
view article Article Fine-Tuning MetaCLIP-2 for Image Classification on Downstream Tasks prithivMLmods • Nov 15, 2025 • 7
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning Paper • 2511.02818 • Published Nov 4, 2025 • 15
SelectMix: Enhancing Label Noise Robustness through Targeted Sample Mixing Paper • 2509.11265 • Published Sep 14, 2025 • 1
Intra-Cluster Mixup: An Effective Data Augmentation Technique for Complementary-Label Learning Paper • 2509.17971 • Published Sep 22, 2025 • 1
Token Activation Map to Visually Explain Multimodal LLMs Paper • 2506.23270 • Published Jun 29, 2025 • 5