Model Description
AgenticQwen-30B-A3B is a small agentic language model trained on Qwen3-30B-A3B-Instruct, designed for multi-step reasoning and tool use. It is trained with a multi-round reinforcement learning (GRPO-style) pipeline and a dual "data flywheel" mechanism that continually increases task difficulty for both reasoning and agentic workflows.