InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language Models
Paper • 2503.06692 • Published • 2
InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning
Note SFT model trained with InftyThink-style OpenThoughts-114k from DeepSeek-R1-Distill-Qwen-1.5B.
Note RL model trained with DeepScaleR from yanyc/InftyThink-1.5B using Task reward.
Note RL model trained with DeepScaleR from yanyc/InftyThink-1.5B using Task reward and Efficiency reward.
Note SFT model trained with InftyThink-style OpenThoughts-114k from Qwen3-4B-Base.
Note RL model trained with DeepScaleR from yanyc/InftyThink-4B using Task reward.
Note RL model trained with DeepScaleR from yanyc/InftyThink-4B using Task reward and Efficiency reward.
Note InftyThink-style cold start data.