ThinkTwice: Jointly Optimizing Large Language Models for Reasoning and Self-Refinement Paper โข 2604.01591 โข Published Apr 2 โข 42