Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification Paper • 2609.30467 • Published 12 days ago • 11
EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents Paper • 2609.17632 • Published 21 days ago • 46
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 29 days ago • 376
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets Paper • 2609.05663 • Published Sep 4 • 21
Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding Paper • 2609.04131 • Published Sep 3 • 31
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes Paper • 2609.03796 • Published Sep 3 • 186
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published Aug 18 • 161
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published Aug 14 • 287
SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries Paper • 2608.05604 • Published Aug 6 • 81
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 266
Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging Paper • 2608.03316 • Published Aug 4 • 26
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Paper • 2608.02023 • Published Aug 3 • 161
ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction Paper • 2607.29677 • Published Jul 31 • 26