Recognition-Refusal Misalignment in LLMs: Why Models Answer Structurally Unanswerable Questions
Paper • 2608.29109 • Published • 17
None defined yet.
Adaptive Generate-Rank-Verify: Inference-Time Search with Costly Verification
Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning