DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 13 days ago • 190
view article Article Trained 210M text-to-image model from scratch on one GPU: what actually mattered ivanmikhnenkov • 20 days ago • 10
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Paper • 2609.11638 • Published 20 days ago • 702
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 23 days ago • 375
Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling Paper • 2608.30821 • Published 30 days ago • 69
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 29 days ago • 221
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published 28 days ago • 403
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 27 days ago • 248
Compile by Training: Turning Natural-Language Specifications into Local Neural Functions Paper • 2609.04199 • Published 27 days ago • 332
view article Article NeoMME: an efficient Multimodal-native and Multilingual Encoder Hcompany • 26 days ago • 110
view article Article Training a coding model to paint watercolours with TRL and OpenEnv sergiopaniego • 27 days ago • 75
view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • 27 days ago • 139
view article Article Open Yap 1K: 1,000 hours of full-duplex natural conversation, free for commercial use TheAgenticDataCompany • 26 days ago • 20
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published 30 days ago • 63
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 30 days ago • 97
DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution Paper • 2608.31106 • Published 30 days ago • 96