Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference Paper • 2609.05275 • Published 4 days ago • 12
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 5 days ago • 276
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers Paper • 2609.01343 • Published 7 days ago • 98
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published 8 days ago • 53
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation Paper • 2607.05147 • Published Jul 6 • 49
Don't Retrain, Align: Adapting Autoregressive LMs to Diffusion LMs via Representation Alignment Paper • 2605.06885 • Published May 7 • 1
ParaTempo: Efficient Parallel Reasoning via Temporal Confidence Paper • 2608.16425 • Published 21 days ago • 40
TinyStories: How Small Can Language Models Be and Still Speak Coherent English? Paper • 2305.07759 • Published May 12, 2023 • 48
Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts Paper • 2608.20061 • Published 18 days ago • 46
FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution Paper • 2608.16157 • Published 22 days ago • 107
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning Paper • 2608.09888 • Published 29 days ago • 775
AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling Paper • 2608.02602 • Published Aug 3 • 82