LessThink-Qwen3-4B-v1

Qwen3-4B (thinking mode) post-trained with GRPO to think less: a reward that keeps correctness and penalizes thinking tokens only (everything before </think>), measured relative to the base model's own thinking length on each prompt.

Results (32K token cap, full benchmark sets, multi-seed)

Benchmark Qwen3-4B LessThink Δ acc (pp) Δ thinking tokens (matched)
GSM8K 95.0 94.8 -0.2 -63%
MMLU-Pro 72.0 70.5 -1.6 -57%
MATH-500 96.8 96.2 -0.7 -52%
GPQA-Diamond 54.9 53.0 -1.9 -50%
AIME 2024 72.7 70.2 -2.5 -34%
HMMT Feb 2025 46.2 39.6 -6.7 -34%
AIME 2025 63.5 55.2 -8.3 -34%

Matched Δ = geometric mean over questions of the per-question thinking-length ratio (seeds averaged first). Truncation on competition math drops from 8-9% to 1.5-2.5%; loop rate falls to ~0.

Accuracy under a token budget

A response counts as correct at budget B only if it finished within B total tokens.

Budget AIME24 base / ours AIME25 base / ours MATH-500 base / ours
4K 1.9 / 15.8 0.6 / 15.6 56.4 / 79.8
8K 23.5 / 43.5 17.1 / 32.5 83.4 / 91.6
16K 57.9 / 65.4 48.8 / 50.8 94.7 / 95.1
32K 72.7 / 70.2 63.5 / 55.2 96.8 / 96.2

Full tables and budget curves for all benchmarks are in eval/.

Limitations

  • On competition math with an unlimited budget, accuracy drops 2.5-8 pp.
  • Final answers are 5-20% shorter than the base model's, although only thinking tokens were penalized.
  • No coding, agentic or instruction-following evaluation yet.

Training

  • GRPO (TRL 0.21), LoRA r=64 / alpha=128 on all linear layers, merged into the base weights. Released checkpoint: step 50 of 400.
  • Reward: correct -> 1 + aclip(1 - L/L_ref, -1, 1); wrong -> -aclip(L/L_ref - 1, 0, 1), where L = thinking tokens, L_ref = the base model's mean thinking length on correct samples for that prompt, a = 0.3 (linear warmup over 50 steps).
  • Data: 5,376 prompts (GSM8K train, DeepScaleR subset, ARC-Challenge, SciQ, CommonsenseQA), 13-gram decontaminated against every evaluation set.
  • lr 1e-5, KL beta 0.02, 8 generations per prompt, 64 completions per step, max completion 8,192 tokens, 1x H100.

Evaluation setup

vLLM 0.10.0, temperature 1.0, top_p 0.95, top_k 20, min_p 0, 32,768-token cap, truncated responses scored wrong. Seeds: AIME/HMMT 16, GPQA 8, MATH-500 4, GSM8K 2, MMLU-Pro 1.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
m = "5ivatej/LessThink-Qwen3-4B-v1"
tok = AutoTokenizer.from_pretrained(m)
model = AutoModelForCausalLM.from_pretrained(m, torch_dtype="auto", device_map="auto")
msgs = [{"role": "user", "content": "What is 17*23?"}]
x = tok.apply_chat_template(msgs, add_generation_prompt=True, enable_thinking=True, return_tensors="pt").to(model.device)
out = model.generate(x, max_new_tokens=4096, temperature=1.0, top_p=0.95, top_k=20, do_sample=True)
print(tok.decode(out[0][x.shape[1]:], skip_special_tokens=True))

Recommended sampling: temperature 1.0, top_p 0.95, top_k 20, min_p 0.

Downloads last month
221
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 5ivatej/LessThink-Qwen3-4B-v1

Finetuned
Qwen/Qwen3-4B
Finetuned
(1135)
this model
Quantizations
1 model