Math12k-GRPO-Control-7B

Size- and step-budget-matched general-math GRPO control from WirelessMathBench-XL: An Auditable Benchmark for Wireless Mathematical Reasoning (NeurIPS 2026, Evaluations and Datasets Track). It is the comparison point for WirelessMathLM-7B, not a wireless-math model.

Project page · arXiv · OpenReview · Code · Dataset

Model

  • Base: Qwen/Qwen2.5-7B base checkpoint (no SFT warm-start).
  • Training data: 3,227 problems sampled from hiyouga/math12k (train split, seed 20260501), matching the size of the WirelessMathBench-XL train split.
  • Training: EasyR1 GRPO, 40 epochs / 240 steps, format weight 0.1, same budget as WirelessMathLM-7B. Lower-level batch settings, prompt template, and reward implementation are not identical to the WirelessMathLM runs, so this is not a one-variable causal ablation.
  • Precision: bfloat16.

Results on WirelessMathBench-XL (800-item test split)

Model Accuracy
Math12k-GRPO-Control-7B (this model) 22.75% (182/800)
WirelessMathLM-7B 47.88% (383/800)

Paired problem-level bootstrap (10,000 resamples, seed 0): control − WirelessMathLM-7B = −25.13 pp, 95% CI [−29.00, −21.13]. One training run per condition, so the interval reflects test-item resampling, not training-seed variation. Protocol: raw completion endpoint, 2,048-token answer budget, T = 0.6, hierarchical verifier with GPT-4.1-mini fallback.

This supports domain-aligned learnability on the release split only. It does not establish intrinsic difficulty of wireless mathematics, a causal isolation of wireless constraints, or paper-disjoint generalisation (the release split shares source papers between train and test).

Citation

@inproceedings{
li2026wirelessmathbenchxl,
title={WirelessMathBench-XL: An Auditable Benchmark for Wireless Mathematical Reasoning},
author={Xin Li and Mengbing Liu and Yiyang Zhu and Wenhe Zhang and Li Wei and Jiancheng An and Chau Yuen},
booktitle={The Fortieth Annual Conference on Neural Information Processing Systems Evaluations and Datasets Track},
year={2026}
}
Downloads last month
-
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for XINLI1997/Math12k-GRPO-Control-7B

Base model

Qwen/Qwen2.5-7B
Finetuned
(969)
this model

Dataset used to train XINLI1997/Math12k-GRPO-Control-7B

Collection including XINLI1997/Math12k-GRPO-Control-7B

Paper for XINLI1997/Math12k-GRPO-Control-7B