WirelessMathLM-0.5B

GRPO-trained reference checkpoint from WirelessMathBench-XL: An Auditable Benchmark for Wireless Mathematical Reasoning (NeurIPS 2026, Evaluations and Datasets Track).

Project page · arXiv · OpenReview · Code · Dataset

Model

  • Base: Qwen/Qwen2.5-0.5B. Initialised from the Qwen2.5-0.5B base checkpoint (no SFT warm-start).
  • Training: GRPO with EasyR1 on the WirelessMathBench-XL train split (3,227 problems), 40 epochs / 240 steps, composite reward 0.1 × format + 0.9 × verifier accuracy, AdamW (lr 1e-6, cosine), KL coefficient 0.01, 4 × NVIDIA A6000. See Appendix B of the paper.
  • Precision: bfloat16.

Results

Evaluation (WirelessMathBench-XL test) Accuracy
Full 800-item test split (paper Tab. 3, row D14) 2.12%
310-item public test subset (paper Appendix L) 2.26%

Locked protocol: 2k-token answer budget, T = 0.6, hierarchical verifier (deterministic match, canonicalisation, GPT-4.1-mini fallback judge).

Training-time check under greedy model selection (paper Tab. 7): base 13.38% → GRPO 14.87%. These values use a different protocol from the locked scores.

Usage

The paper queries these checkpoints through a raw completion endpoint (no chat template), using the dataset's prompt field.

from datasets import load_dataset
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "XINLI1997/WirelessMathLM-0.5B"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto", device_map="auto")

ex = load_dataset("XINLI1997/WirelessMATHBench-XL", "cc_by", split="test")[0]
prompt = ex["prompt"]
inputs = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=2048, do_sample=True, temperature=0.6)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Limitations

  • These are release-split learnability checks, not evidence of general or verifier-independent reasoning. The train/test split is problem-level, so train and test share source papers.
  • The reward verifier and the evaluation fallback share the same checkpoint family and prompt style.
  • Intended for research on wireless mathematical reasoning; not a general-purpose assistant.
  • The 0.5B checkpoint stays near the floor on the locked protocol and is included for completeness.

License

Apache 2.0, following the base model.

Citation

@inproceedings{
li2026wirelessmathbenchxl,
title={WirelessMathBench-XL: An Auditable Benchmark for Wireless Mathematical Reasoning},
author={Xin Li and Mengbing Liu and Yiyang Zhu and Wenhe Zhang and Li Wei and Jiancheng An and Chau Yuen},
booktitle={The Fortieth Annual Conference on Neural Information Processing Systems Evaluations and Datasets Track},
year={2026}
}
Downloads last month
-
Safetensors
Model size
0.6B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for XINLI1997/WirelessMathLM-0.5B

Finetuned
(728)
this model

Dataset used to train XINLI1997/WirelessMathLM-0.5B

Collection including XINLI1997/WirelessMathLM-0.5B

Paper for XINLI1997/WirelessMathLM-0.5B