Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published 18 days ago • 37
Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers Paper • 2607.28611 • Published 15 days ago • 22
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Paper • 2607.21653 • Published 23 days ago • 32
Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering Paper • 2607.21848 • Published 22 days ago • 9
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Paper • 2607.20465 • Published May 19 • 56
Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models Paper • 2607.19604 • Published 24 days ago • 17
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD Paper • 2607.20145 • Published 23 days ago • 74
Environment-free Synthetic Data Generation for API-Calling Agents Paper • 2607.16900 • Published 27 days ago • 21
SciForma: Structure-Faithful Generation of Scientific Diagrams Paper • 2607.18091 • Published 25 days ago • 23
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published 29 days ago • 172
ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation Paper • 2607.13124 • Published about 1 month ago • 20
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published 29 days ago • 106
Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning Paper • 2607.08393 • Published Jul 9 • 18