---
base_model: Qwen/Qwen2.5-7B-Instruct
library_name: peft
license: apache-2.0
datasets:
- openai/gsm8k
language:
- en
pipeline_tag: text-generation
tags:
- mathematical-reasoning
- chain-of-thought
- qlora
- peft
- gsm8k
---
# Qwen2.5-7B Positional Reasoning
A [QLoRA](https://arxiv.org/abs/2305.14314) fine-tune of
[Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct) that solves grade-school math
word problems by **committing to a prediction before each step is computed**, rather than narrating a
solution after the fact.
## 1. Introduction
Large language models solve multi-step arithmetic word problems by generating a chain of reasoning,
but that chain is produced *after* the model has effectively settled on an answer — the reasoning
narrates a conclusion rather than deriving one. This matters because a narrated chain gives little
signal about where it went wrong: an error sits buried somewhere in a paragraph of fluent prose,
undetectable until the final number is checked.
This model was fine-tuned on GSM8K reformatted into a structure called **Positional Reasoning
Format**: a chain of typed nodes where each node must commit to a `PREDICT` field describing what the
*next* step will accomplish, before that step runs. The design borrows from
[Hawkins & Ahmad's](https://www.frontiersin.org/articles/10.3389/fncir.2016.00023/full) account of the
cortex as a prediction machine, in which intelligence arises from continuous prediction and error
correction rather than passive pattern completion. Every reasoning step becomes a checkpoint, so a
wrong prediction is a localized, detectable failure instead of a buried one.
**A second, independent question was investigated in parallel.** The format includes a `filter` node
type intended to handle irrelevant information —
[a known weakness of LLMs](https://arxiv.org/abs/2302.00093), and one this project documented directly
when Llama-3-8B folded an unrelated quantity into its arithmetic simply because the number appeared in
the problem. Roughly 17% of training examples carry an injected distractor plus an explicit filter
step. This was evaluated as a **separate experiment on its own test set**, and unlike the main result
it came out inconclusive — see [Limitations](#8-limitations). The two questions are reported
separately throughout because they were tested separately; the GSM8K result below does not depend on
the distractor result.
**Main result:** on the full 1,319-problem GSM8K test split, accuracy rose from **75.8% to 81.7%
(+5.9 points)** over the base model — 78 additional problems solved, a paired McNemar
*p* = 2.4 × 10⁻⁶. For scale, that gain is larger than the 4.2-point gap separating
Llama-3-8B-Instruct from Qwen2.5-7B-Instruct, and it was obtained from 556 training examples in
44 minutes on a single NVIDIA T4. General capability declined slightly but measurably (LogiQA2 −2.5,
ARC-Challenge −1.0, MMLU −0.5). That combination is the interesting part: a math-specialized model of
the same size and family reaches higher raw GSM8K accuracy but sacrifices roughly eleven times more
general capability to get there (see [Evaluation](#4-evaluation)).
**Stated up front:** this study establishes that fine-tuning on 556 examples reformatted this way
improves GSM8K accuracy. It does *not* establish that the prediction mechanism is what caused the
improvement — no ablation separates it from structured formatting, longer outputs, or ordinary
in-domain fine-tuning. That limitation is discussed in [Limitations](#8-limitations).
## 2. Data
All data derives from [GSM8K](https://huggingface.co/datasets/openai/gsm8k). Because this project
investigated [two separate questions](#1-introduction), the materials are described separately below.
**Everything used for training came from GSM8K's `train` split; everything used for evaluation came
from its `test` split.** The two splits are disjoint by construction, so no test problem was ever
trained on.
### 2a. Training data — the primary (prediction) track
Built from the GSM8K **train** split (7,473 problems available):
| Stage | Result |
|---|---|
| Selection: every 8th example (deterministic, no seed needed) | 935 candidates |
| Reformatted into Positional Reasoning Format by Claude Opus 4.8, few-shot prompted with 3 hand-authored gold examples | 573 processed |
| Six mechanical validation checks | **556 accepted**, 17 rejected (97%) |
| Train/validation split, 90/10, `seed=42` | **500 train / 56 validation** |
The remaining 362 candidates were never processed. The 500-600 example target had been fixed in
advance during project planning, the run met it, and the reformatting API budget was exhausted at
roughly the same point — so the cost constraint did not, in the end, bind on the plan. Scaling the
corpus is nonetheless the most obvious lever for improving on these results, particularly for the
filter behavior described in [2b](#2b-distractor-materials--the-secondary-filter-track), which had
only ~95 supporting examples.
The six checks are implemented in [`scripts/validator.py`](https://huggingface.co/Niarfe/qwen2.5-7b-positional-reasoning/blob/main/scripts/validator.py), included in this repository. The most important is the **look-ahead detector**: a `PREDICT` field may not contain
a number that has not already been established by the question or an earlier `OBSERVE`/`STATE`. Without
it, the reformatter could silently solve each problem first and back-fill plausible-sounding
predictions — which would defeat the entire premise of the format.
**The 90/10 split above applies to all 556 training examples**, including the distractor-carrying ones
described next.
### 2b. Distractor materials — the secondary (filter) track
Two separate artifacts, both built by injecting exactly one irrelevant quantity into a problem while
leaving the correct answer unchanged:
| Artifact | Source split | Size | Purpose |
|---|---|---|---|
| Distractor **training** examples | `train` (subset of the 556 above) | ~95 (every 6th selected, ~17%) | Teach the `filter` node behavior |
| Distractor **test** set | `test` | **40** | Measure whether that behavior transferred |
The training subset is deliberately partial: clean problems keep clean chains, so filtering is learned
as a *conditional* skill rather than a ritual applied to every problem.
The 40-problem test set is **paired** — each problem is stored in both its original and
distractor-injected form, so each item serves as its own control. That pairing is what permits the
McNemar analysis reported in [Limitations](#8-limitations). Note that these injections were
LLM-generated and manually reviewed rather than drawn from a published benchmark; a standard dataset
for this purpose (GSM-IC, from the same work cited in the introduction) would have offered a larger
sample and results comparable to published numbers.
## 3. Methodology
QLoRA was not really optional here: full fine-tuning of a 7B model does not fit the 16GB VRAM of the
NVIDIA T4 this was trained on, so 4-bit quantization plus low-rank adapters was the only feasible
route. The hyperparameters, however, were chosen deliberately in response to a failed earlier attempt.
An earlier coursework experiment applied LoRA to
[Llama-3.2-1B](https://huggingface.co/meta-llama/Llama-3.2-1B) at `lr=1e-5` targeting only
`q_proj`/`v_proj`, and produced a nominal accuracy gain that masked *worse* underlying reasoning on
inspection. That result is suggestive rather than conclusive — at 1B parameters, weak multi-step
arithmetic is partly a capacity limitation and cannot be attributed to hyperparameters alone — but it
motivated a hypothesis worth testing directly. Following
[Hu et al. (2021)](https://arxiv.org/abs/2106.09685) and
[Dettmers et al. (2023)](https://arxiv.org/abs/2305.14314), the hypothesis was that `1e-5` sits well
below the `~1e-4`–`3e-4` range recommended for adapters, and that attention-only adaptation
under-powers multi-step reasoning. Three combinations tested this directly:
| Combo | Learning rate | Target modules | Final eval loss | Eval token accuracy |
|---|---|---|---|---|
| A | 1e-5 | `q_proj`, `v_proj` | 0.718 | 83.2% |
| B | 1e-4 | `q_proj`, `v_proj` | 0.191 | 94.2% |
| **C (released)** | **2e-4** | **all-linear** | **0.138** | **95.6%** |
The diagnosis held monotonically: fixing the learning rate alone closed most of the gap, and widening
to all linear projections closed the rest. **Combo C is the model published here.**
Full configuration for reproducibility:
| | |
|---|---|
| Base model | `Qwen/Qwen2.5-7B-Instruct` |
| Quantization | 4-bit NF4, double quantization, fp16 compute |
| LoRA | `r=64`, `lora_alpha=64`, `lora_dropout=0.05`, `bias="none"` |
| Target modules | all-linear (`q,k,v,o,gate,up,down_proj`) — 161M trainable params (2.08%) |
| Learning rate | 2e-4 |
| Epochs | 2 (126 optimizer steps) |
| Batch | 1 × grad-accum 8 |
| Optimizer | `paged_adamw_8bit`, gradient checkpointing enabled |
| Max sequence length | 1600 |
| Seed | 42 |
| Trainer | [TRL](https://github.com/huggingface/trl) `SFTTrainer`, prompt/completion format (completion-only loss) |
| Training script | [`scripts/train_qlora.py`](https://huggingface.co/Niarfe/qwen2.5-7b-positional-reasoning/blob/main/scripts/train_qlora.py) (included in this repo) |
| Hardware | 1 × NVIDIA T4 (16GB), ~44 minutes |
## 4. Evaluation
**Table 1 — Headline comparison.** Five models on four benchmarks. Rows 1-4 are off-the-shelf
models benchmarked here rather than quoted from published results; row 4 is also the base this
model was fine-tuned from. Best score per column in bold.
| Model | GSM8K (test split) | LogiQA2 | ARC-Challenge | MMLU |
|---|---|---|---|---|
| [Mistral-7B-Instruct-v0.3](https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.3) | 47.1% | 40.6% | **58.6%** | 59.3% |
| [Meta-Llama-3-8B-Instruct](https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct) | 71.6% | 30.9% | 54.6% | 62.9% |
| [Qwen2.5-Math-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Math-7B-Instruct) | **89.6%** | 30.3% | 41.2% | 53.0% |
| [Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct) (base) | 75.8% | **43.8%** | 54.0% | **71.0%** |
| **This model** | 81.7% | 41.3% | 53.0% | 70.5% |
Every row was produced by the same script —
[`scripts/benchmark_suite.py`](https://huggingface.co/Niarfe/qwen2.5-7b-positional-reasoning/blob/main/scripts/benchmark_suite.py),
included in this repository — wrapping
[lm-evaluation-harness](https://github.com/EleutherAI/lm-evaluation-harness). Identical tasks,
identical 4-bit NF4 quantization, **no sampling cap** (GSM8K n=1,319; LogiQA2 n=1,572; ARC-Challenge
n=1,172; MMLU n=14,042), and no chat template applied to any model. Only the `--model` argument
differed between rows.
**Why these benchmarks.** GSM8K is the testing split of the training data and measures the target
skill directly. The other three probe whether the fine-tune damaged general ability:
[LogiQA2](https://huggingface.co/datasets/baber/logiqa2) for logical reasoning,
[ARC-Challenge](https://huggingface.co/datasets/allenai/ai2_arc) for science reasoning, and
[MMLU](https://huggingface.co/datasets/cais/mmlu) as a broad catastrophic-forgetting canary at its
full 14,042-example size.
**Why these comparison models.** Mistral-7B-Instruct-v0.3 is the standard general-purpose 7B instruct
peer. Llama-3-8B-Instruct is the model whose documented distractor failure originally motivated this
project's filter-node design. Qwen2.5-Math-7B-Instruct is the toughest available comparison — the
math-specialized sibling of this model's own base — which asks directly: *why fine-tune a general model
when a purpose-built math model already exists?*
**Reading the results honestly.** This model leads GSM8K among the general-purpose models, but most of
its margin over Mistral and Llama-3 comes from the **base model choice**, not from this fine-tune:
Qwen2.5-7B-Instruct already scored 75.8%. The fine-tune's own contribution is **+5.9 points**. This
model also does not win every column and does not claim to — Mistral leads ARC-Challenge, and the
untuned base leads LogiQA2 and MMLU. Leading only on the task actually fine-tuned for is the expected
profile for a task-specific adapter.
**Table 2 — Effect of the fine-tune: base vs. this model only.** This table isolates rows 4 and 5
of Table 1 — the only pair where everything except the fine-tuning is held constant, and therefore
the only comparison that measures what the training actually did. Because both were evaluated on
identical problem sets, the appropriate test is paired (McNemar) rather than a comparison of
independent proportions. "Net" is problems gained minus problems lost.
| Benchmark | n | Base correct | This model | Net | McNemar *p* |
|---|---|---|---|---|---|
| GSM8K | 1,319 | 1,000 | 1,078 | **+78** | **2.4 × 10⁻⁶** |
| LogiQA2 | 1,572 | 689 | 650 | −39 | 0.011 |
| MMLU | 14,042 | 9,971 | 9,896 | −75 | 0.006 |
| ARC-Challenge | 1,172 | 601 | 588 | −13 | 0.298 (n.s.) |
The GSM8K improvement is the strongest effect measured in this project by roughly three orders of
magnitude. The LogiQA2 and MMLU declines are also genuine rather than noise, and are reported as such;
the ARC-Challenge change is not statistically distinguishable from zero.
One caveat worth stating plainly: counted as raw problems rather than rates, this model gains 78 and
loses 127 across the other three benchmarks. That framing is misleading — MMLU is more than ten times
the size of GSM8K, so summing raw counts lets benchmark size, not capability, drive the conclusion —
but the numbers are given here so readers can judge for themselves. Measured as rates, the target-task
gain (+7.8% relative) exceeds every individual decline (−5.7%, −1.9%, −0.8%).
**Table 3 — Specialization tradeoff: this model vs. the math specialist.** Both are derived from
the same Qwen2.5-7B family and scale, so each is measured as its change *from the shared base*
(row 4 of Table 1). "Cost" sums the declines across LogiQA2, ARC-Challenge and MMLU in percentage
points.
| | GSM8K gain | Cost across the other three | Gain per point sacrificed |
|---|---|---|---|
| This model | +5.9 | 4.0 | **1.46** |
| Qwen2.5-Math-7B-Instruct | +13.8 | 44.3 | 0.31 |
Qwen2.5-Math wins on raw GSM8K accuracy and should be preferred if grade-school arithmetic is the only
thing that matters. But it pays with an 18-point MMLU collapse. This model buys a smaller gain at
roughly one-eleventh the cost to general capability — which is the actual claim being made here:
predict-before-observe fine-tuning delivers a real target-task improvement while leaving the model
broadly intact.
## 5. Usage and Intended Uses
This is a **LoRA adapter**, not a standalone model. Loading it requires the base model plus
[PEFT](https://github.com/huggingface/peft):
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained(
"Qwen/Qwen2.5-7B-Instruct", device_map="auto", torch_dtype="float16"
)
model = PeftModel.from_pretrained(base, "Niarfe/qwen2.5-7b-positional-reasoning")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-7B-Instruct")
SYSTEM = ("You solve math word problems using positional reasoning: a chain of typed nodes, "
"each committing to a prediction before the next step is computed. "
"End with the final answer in tags.")
messages = [
{"role": "system", "content": SYSTEM},
{"role": "user", "content": "Zhang is twice as old as Li. Li is 12 years old. "
"Zhang's brother Jung is 2 years older than Zhang. How old is Jung?"},
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=900, do_sample=False)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
```
**Intended uses.** Grade-school arithmetic word problems where an *auditable* reasoning trace matters
more than raw speed — tutoring tools that need to show their work, systems that verify intermediate
steps, or research into structured reasoning formats. The typed-node output is designed to be parsed
programmatically, so each step can be checked in isolation.
**Not intended for.** Competition-level or higher mathematics, non-mathematical reasoning, or
production use without human review. See [Limitations](#8-limitations).
## 6. Prompt Format
Use the standard Qwen2.5 chat template with a system message requesting positional reasoning. The user
turn is the raw word problem — no worked examples required.
```
<|im_start|>system
You solve math word problems using positional reasoning: a chain of typed nodes, each committing to a prediction before the next step is computed. End with the final answer in tags.<|im_end|>
<|im_start|>user
Zhang is twice as old as Li. Li is 12 years old. Zhang's brother Jung is 2 years older than Zhang. How old is Jung?<|im_end|>
<|im_start|>assistant
```
## 7. Expected Output Format
The model emits a numbered chain of typed nodes — `entry`, `filter`, `resolve`, `combine`, `check`,
`branch` — followed by the final answer in `` tags. The entry node declares `STATE` and `GOAL`;
every later node carries `MOVE`, `OBSERVE`, `STATE`, and `PREDICT`. Each `PREDICT` commits to what the
*next* node will do, and the final node's `PREDICT` is always `goal resolved`.
```
STATE: Li's age known (12). Zhang's age unknown (twice Li). Jung's age unknown (2 more than Zhang).
GOAL: Find Jung's age.
PREDICT: Next node will resolve Zhang's age using the twice-Li relationship.
MOVE: Apply Zhang = Li x 2.
OBSERVE: 2 x 12 = 24
STATE: Li's age known (12). Zhang's age known (24). Jung's age unknown.
PREDICT: Next node will resolve Jung's age by adding 2 to Zhang's age.
MOVE: Apply Jung = Zhang + 2.
OBSERVE: 24 + 2 = 26
STATE: Jung's age known (26). Goal resolved.
PREDICT: goal resolved
26
```
When a problem contains irrelevant information, a `filter` node appears at position 1 and explicitly
discards it before any arithmetic begins. Given *"Jana is 5 inches taller than Kelly, and Kelly is 3
inches shorter than Jess. **Their coach Diego is 68 inches tall.** If Jess is 72 inches tall, how tall
is Jana?"* — the first two nodes are:
```
STATE: Jana = Kelly + 5. Kelly = Jess - 3. Diego's height mentioned (68 inches). Jess known (72 inches). Kelly unknown. Jana unknown.
GOAL: Find Jana's height.
PREDICT: Next node will filter the given information against the goal.
MOVE: Evaluate all given information against the stated goal.
OBSERVE: Jess's height (72): relevant. Kelly = Jess - 3: relevant. Jana = Kelly + 5: relevant. Diego's height (68): NOT relevant to Jana's height — discard.
STATE: Active scope: Jess (72), Kelly = Jess - 3, Jana = Kelly + 5. Discarded: Diego's 68 — not referenced by goal.
PREDICT: Next node will resolve Kelly's height using the shorter-than-Jess relationship.
```
The remaining nodes then proceed as normal. Note the design weakness discussed in
[Limitations](#8-limitations): the filter node *names* the discarded quantity, so "Diego" and "68"
remain in context for every subsequent step rather than being removed from it.
## 8. Limitations
**The central mechanism is unverified — the accuracy gain cannot be attributed to prediction.**
This is the most important caveat on this page. The format constrains the *order tokens are emitted
in*; it does not constrain the computation that produces them. A model that has already internally
settled on an answer can emit a fluent `PREDICT` field describing the upcoming step as pure narration,
and nothing measured here would distinguish that from a genuine prior commitment. The look-ahead
validator described in [Data](#2-data) policed the *training corpus* — it guaranteed the reformatter
never leaked results into `PREDICT` fields — but it says nothing about the trained model's behavior at
inference time.
The problem compounds at the level of attribution. At least four explanations are consistent with the
+5.9 point gain, and this study cannot separate them:
1. The predict-before-observe mechanism does something real;
2. Any consistent structured format aids problem decomposition;
3. The node format emits substantially more tokens than plain chain-of-thought, and additional
generated tokens improve accuracy independently of their content;
4. Fine-tuning on 556 in-domain GSM8K training problems would improve GSM8K test accuracy in
*any* format.
**The controlled experiment that would settle this was not run:** fine-tuning the same base model, on
the same 556 problems, with the same hyperparameters, but using GSM8K's *original* solution text
instead of the reformatted chains. If that ablation also reached ~81%, the format contributed nothing
and the gain is simply in-domain fine-tuning. Absent that comparison, the honest claim is narrow —
*fine-tuning on 556 examples reformatted this way improves GSM8K by 5.9 points* — not that
prediction-before-computation is the cause.
There is also a suggestive counter-signal in the two-call result below. If `PREDICT` were a genuine,
separable commitment, handing it to a fresh context should be unproblematic — any executor could carry
it out. Accuracy instead fell 25 points, which is at least consistent with the prediction being
entangled with the model's own rolling state rather than functioning as an independent commitment.
(An out-of-distribution formatting effect explains the same observation, so this is not decisive
either way.) A cheap discriminating test, not yet run: supply a deliberately *incorrect* `PREDICT` and
observe whether the model follows it into a wrong answer (the field steers computation) or silently
ignores it and answers correctly (the field is decorative).
**Mild but genuine catastrophic forgetting.** General capability declined: LogiQA2 −2.5, ARC-Challenge
−1.0, MMLU −0.5 points. Two of the three are statistically real, not sampling noise — paired McNemar
gives *p* = 0.011 for LogiQA2 (39 net problems lost) and *p* = 0.006 for MMLU (75 net lost), both on
full-size benchmarks. The ARC-Challenge change (*p* = 0.298) is not distinguishable from zero. The
effect is small in absolute terms but should not be dismissed: this adapter does make the model
slightly worse at everything other than the task it was trained on.
**The distractor-filtering skill was not demonstrated.** The training data deliberately includes ~95
examples with a `filter` node, and a held-out 40-problem distractor set was built to test whether that
skill transferred. It produced no detectable improvement — but the test was **underpowered, and the
result is inconclusive rather than negative**. A paired McNemar test gives *p* = 0.607 for the base
model and *p* = 0.180 for this one; the apparent difference between them amounts to two problems out
of forty, and several problems flip to *correct* when an irrelevant sentence is added, which cannot be
a real effect. A conclusive test would need several hundred problems per condition. There is also a
design weakness: the `filter` node *names* the distractor, but those tokens remain in context for every
subsequent step, so attention can still reach them. A stronger approach would restate the problem with
the distractor removed and solve the restatement.
**Poor robustness to an externally supplied reasoning prefix.** When inference is split into two calls
— one to produce the entry node, a second in a fresh context to continue from it — accuracy drops
**25 points on a 100-problem GSM8K subset** (65.0% → 40.0%), a statistically solid effect at roughly
3.7 standard errors. The model was trained exclusively on complete, self-generated chains, so it
appears to depend on continuity with its own prior context and does not gracefully continue someone
else's reasoning.
**Small training set.** 556 examples, roughly 7% of the GSM8K train split, reformatted by an LLM rather
than written by humans. The reformatting was mechanically validated but not exhaustively hand-reviewed.
**Quantized throughout.** Both training and all reported evaluations used 4-bit NF4 quantization.
Behavior at full precision was not measured and may differ from the numbers above.
**Domain-narrow.** Trained and evaluated only on grade-school arithmetic word problems. Nothing here
demonstrates transfer to algebra, geometry, competition mathematics, or non-mathematical reasoning —
and the OOD benchmark results suggest the fine-tune slightly *reduces* general reasoning rather than
improving it.
## Citation and provenance
Built for **UVA DS 5002: How to Train Your LLM** (Summer 2026). Training data derived from
[GSM8K](https://huggingface.co/datasets/openai/gsm8k) (Cobbe et al., 2021,
[arXiv:2110.14168](https://arxiv.org/abs/2110.14168)); base model
[Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct); method after
[QLoRA](https://arxiv.org/abs/2305.14314) (Dettmers et al., 2023) and
[LoRA](https://arxiv.org/abs/2106.09685) (Hu et al., 2021).