WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory Paper • 2609.24984 • Published 4 days ago • 143
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay Paper • 2609.25001 • Published 4 days ago • 126
Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation Paper • 2609.20758 • Published 8 days ago • 4
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 7 days ago • 131
xTayyub/High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning Updated Dec 9, 2025 • 33 • 6
Yuhan123/vicuna-13b-self_consistency_neg_exp_var_5 Text Generation • 13B • Updated Mar 14, 2025 • 25 • 5
Yuhan123/vicuna-13b-self_consistency_neg_exp_var_4 Text Generation • 13B • Updated Mar 14, 2025 • 28 • 3
Learning Foresight without Explicit Trajectories for 3D Diffusion Policies Paper • 2609.20669 • Published 8 days ago • 8
ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks Paper • 2609.18805 • Published 9 days ago • 62
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents Paper • 2609.17708 • Published 10 days ago • 71
SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization Paper • 2609.14320 • Published 12 days ago • 33
Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration Paper • 2609.16204 • Published 11 days ago • 4
open-llm-leaderboard-old/details_alexredna__Tukan-1.1B-Chat-reasoning-sft-COLA Updated Jan 22, 2024 • 81 • 3