Spaces:
Running
Running
Backfill veritas replication logbook
Browse files- logbook.json +33 -9
- pages/00-scored-evidence-summary/page.md +8 -4
- pages/claim-1-the-me2-principle-characterizes-reasoning-traces-a/page.md +19 -0
- pages/claim-2-the-paper-represents-reasoning-traces-as-dags-with/page.md +19 -0
- pages/claim-3-trm-is-trained-from-a-trm-preference-dataset-with/page.md +19 -0
- pages/claim-4-trm-reaches-88-6-accuracy-on-the-validation-set-pa/page.md +42 -0
- pages/claim-5-trm-guided-best-of-n-test-time-selection-improves/page.md +45 -0
- pages/claim-6-trm-guided-rl-improves-stem-and-math-benchmark-ave/page.md +39 -0
- pages/conclusion/page.md +2 -2
- pages/index.md +6 -2
- pages/reproduction-protocol-and-provenance/page.md +1 -1
logbook.json
CHANGED
|
@@ -13,7 +13,7 @@
|
|
| 13 |
"icml2026-repro",
|
| 14 |
"paper-IMFgiWw4jd"
|
| 15 |
],
|
| 16 |
-
"updated_at": "2026-07-
|
| 17 |
"root": {
|
| 18 |
"slug": "index",
|
| 19 |
"title": "Repro - Characterizing, Evaluating, and Optimizing Complex Reasoning",
|
|
@@ -26,15 +26,39 @@
|
|
| 26 |
"children": []
|
| 27 |
},
|
| 28 |
{
|
| 29 |
-
"slug": "claim-1-
|
| 30 |
-
"title": "Claim 1 -
|
| 31 |
-
"file": "pages/claim-1-
|
| 32 |
"children": []
|
| 33 |
},
|
| 34 |
{
|
| 35 |
-
"slug": "claim-2-
|
| 36 |
-
"title": "Claim 2 -
|
| 37 |
-
"file": "pages/claim-2-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 38 |
"children": []
|
| 39 |
},
|
| 40 |
{
|
|
@@ -51,6 +75,6 @@
|
|
| 51 |
}
|
| 52 |
]
|
| 53 |
},
|
| 54 |
-
"agent_view_tokens":
|
| 55 |
-
"revision": "
|
| 56 |
}
|
|
|
|
| 13 |
"icml2026-repro",
|
| 14 |
"paper-IMFgiWw4jd"
|
| 15 |
],
|
| 16 |
+
"updated_at": "2026-07-20T03:52:32+00:00",
|
| 17 |
"root": {
|
| 18 |
"slug": "index",
|
| 19 |
"title": "Repro - Characterizing, Evaluating, and Optimizing Complex Reasoning",
|
|
|
|
| 26 |
"children": []
|
| 27 |
},
|
| 28 |
{
|
| 29 |
+
"slug": "claim-1-the-me2-principle-characterizes-reasoning-traces-a",
|
| 30 |
+
"title": "Claim 1 - The ME2 principle characterizes reasoning traces along macro\u2026",
|
| 31 |
+
"file": "pages/claim-1-the-me2-principle-characterizes-reasoning-traces-a/page.md",
|
| 32 |
"children": []
|
| 33 |
},
|
| 34 |
{
|
| 35 |
+
"slug": "claim-2-the-paper-represents-reasoning-traces-as-dags-with",
|
| 36 |
+
"title": "Claim 2 - The paper represents reasoning traces as DAGs with progressi\u2026",
|
| 37 |
+
"file": "pages/claim-2-the-paper-represents-reasoning-traces-as-dags-with/page.md",
|
| 38 |
+
"children": []
|
| 39 |
+
},
|
| 40 |
+
{
|
| 41 |
+
"slug": "claim-3-trm-is-trained-from-a-trm-preference-dataset-with",
|
| 42 |
+
"title": "Claim 3 - TRM is trained from a TRM-Preference dataset with a Bradley-\u2026",
|
| 43 |
+
"file": "pages/claim-3-trm-is-trained-from-a-trm-preference-dataset-with/page.md",
|
| 44 |
+
"children": []
|
| 45 |
+
},
|
| 46 |
+
{
|
| 47 |
+
"slug": "claim-4-trm-reaches-88-6-accuracy-on-the-validation-set-pa",
|
| 48 |
+
"title": "Claim 4 - TRM reaches 88.6% accuracy on the validation-set pairwise pr\u2026",
|
| 49 |
+
"file": "pages/claim-4-trm-reaches-88-6-accuracy-on-the-validation-set-pa/page.md",
|
| 50 |
+
"children": []
|
| 51 |
+
},
|
| 52 |
+
{
|
| 53 |
+
"slug": "claim-5-trm-guided-best-of-n-test-time-selection-improves",
|
| 54 |
+
"title": "Claim 5 - TRM-guided best-of-N test-time selection improves AIME24 and\u2026",
|
| 55 |
+
"file": "pages/claim-5-trm-guided-best-of-n-test-time-selection-improves/page.md",
|
| 56 |
+
"children": []
|
| 57 |
+
},
|
| 58 |
+
{
|
| 59 |
+
"slug": "claim-6-trm-guided-rl-improves-stem-and-math-benchmark-ave",
|
| 60 |
+
"title": "Claim 6 - TRM-guided RL improves STEM and math benchmark averages over\u2026",
|
| 61 |
+
"file": "pages/claim-6-trm-guided-rl-improves-stem-and-math-benchmark-ave/page.md",
|
| 62 |
"children": []
|
| 63 |
},
|
| 64 |
{
|
|
|
|
| 75 |
}
|
| 76 |
]
|
| 77 |
},
|
| 78 |
+
"agent_view_tokens": 6826,
|
| 79 |
+
"revision": "1784519552086739000"
|
| 80 |
}
|
pages/00-scored-evidence-summary/page.md
CHANGED
|
@@ -3,9 +3,9 @@
|
|
| 3 |
|
| 4 |
---
|
| 5 |
<!-- trackio-cell
|
| 6 |
-
{"type": "markdown", "id": "
|
| 7 |
-->
|
| 8 |
-
# Scored evidence —
|
| 9 |
|
| 10 |
**Paper:** Characterizing, Evaluating, and Optimizing Complex Reasoning
|
| 11 |
**IDs:** OpenReview `IMFgiWw4jd` · arXiv `2602.08498`
|
|
@@ -14,5 +14,9 @@
|
|
| 14 |
|
| 15 |
| # | Board claim | Verdict | Basis (per-claim replication statuses) |
|
| 16 |
|---:|---|---|---|
|
| 17 |
-
| 1 |
|
| 18 |
-
| 2 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
|
| 4 |
---
|
| 5 |
<!-- trackio-cell
|
| 6 |
+
{"type": "markdown", "id": "cell_258bd67dfb92", "created_at": "2026-07-20T03:52:31+00:00", "title": "Scorecard: 1 verified / 0 falsified / 1 toy / 4 inconclusive of 6 board claims", "pinned": true, "pinned_at": "2026-07-20T03:52:31+00:00"}
|
| 7 |
-->
|
| 8 |
+
# Scored evidence — 1 verified / 0 falsified / 1 toy / 4 inconclusive of 6 board claims
|
| 9 |
|
| 10 |
**Paper:** Characterizing, Evaluating, and Optimizing Complex Reasoning
|
| 11 |
**IDs:** OpenReview `IMFgiWw4jd` · arXiv `2602.08498`
|
|
|
|
| 14 |
|
| 15 |
| # | Board claim | Verdict | Basis (per-claim replication statuses) |
|
| 16 |
|---:|---|---|---|
|
| 17 |
+
| 1 | The ME2 principle characterizes reasoning traces along macro/micro granularity and efficiency/effectiveness axes (Figure 2). | **INCONCLUSIVE** | 0 match / 0 partial / 0 no_match |
|
| 18 |
+
| 2 | The paper represents reasoning traces as DAGs with progression, branching, and merging structures for pairwise evaluation (Figure 3). | **INCONCLUSIVE** | 0 match / 0 partial / 0 no_match |
|
| 19 |
+
| 3 | TRM is trained from a TRM-Preference dataset with a Bradley-Terry preference loss to score reasoning trace quality at scale (Section 5.1). | **INCONCLUSIVE** | 0 match / 0 partial / 0 no_match |
|
| 20 |
+
| 4 | TRM reaches 88.6% accuracy on the validation-set pairwise preference evaluation, outperforming Qwen2.5-Math-PRM-7B, ReasonFlux-PRM-7B, and prompt-only judging (Table 1). | **VERIFIED** | 2 match / 1 partial / 0 no_match; C3, C4, C5 |
|
| 21 |
+
| 5 | TRM-guided best-of-N test-time selection improves AIME24 and AIME25 performance over compared PRM baselines for GPT-OSS-20B and Qwen3-8B response models (Figure 4). | **TOY** | 0 match / 3 partial / 0 no_match; C1, C6, C7 |
|
| 22 |
+
| 6 | TRM-guided RL improves STEM and math benchmark averages over verifier-only RL, with the largest reported gains on Llama-3.1-8B-Instruct (Table 2). | **INCONCLUSIVE** | 0 match / 0 partial / 0 no_match / 3 not-executed (not_attempted); C2, C8, C9 |
|
pages/claim-1-the-me2-principle-characterizes-reasoning-traces-a/page.md
ADDED
|
@@ -0,0 +1,19 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Claim 1: The ME2 principle characterizes reasoning traces along macro/micro granularity a…
|
| 2 |
+
|
| 3 |
+
|
| 4 |
+
---
|
| 5 |
+
<!-- trackio-cell
|
| 6 |
+
{"type": "markdown", "id": "cell_9b370d3e55eb", "created_at": "2026-07-20T03:52:31+00:00", "title": "Claim 1 - verdict and summary"}
|
| 7 |
+
-->
|
| 8 |
+
**Board claim 1:** The ME2 principle characterizes reasoning traces along macro/micro granularity and efficiency/effectiveness axes (Figure 2).
|
| 9 |
+
|
| 10 |
+
**Stated verdict: INCONCLUSIVE** (0 match / 0 partial / 0 no_match)
|
| 11 |
+
|
| 12 |
+
None of the extracted veritas claims test the ME2 principle's definition or its macro/micro and efficiency/effectiveness axes directly. The replication did not attempt to verify this conceptual/framework claim.
|
| 13 |
+
|
| 14 |
+
|
| 15 |
+
---
|
| 16 |
+
<!-- trackio-cell
|
| 17 |
+
{"type": "markdown", "id": "cell_c2df8bd98f33", "created_at": "2026-07-20T03:52:31+00:00", "title": "Scope note"}
|
| 18 |
+
-->
|
| 19 |
+
The replication did not test this claim; no evidence is presented and the verdict is INCONCLUSIVE.
|
pages/claim-2-the-paper-represents-reasoning-traces-as-dags-with/page.md
ADDED
|
@@ -0,0 +1,19 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Claim 2: The paper represents reasoning traces as DAGs with progression, branching, and m…
|
| 2 |
+
|
| 3 |
+
|
| 4 |
+
---
|
| 5 |
+
<!-- trackio-cell
|
| 6 |
+
{"type": "markdown", "id": "cell_1103fbdfce06", "created_at": "2026-07-20T03:52:31+00:00", "title": "Claim 2 - verdict and summary"}
|
| 7 |
+
-->
|
| 8 |
+
**Board claim 2:** The paper represents reasoning traces as DAGs with progression, branching, and merging structures for pairwise evaluation (Figure 3).
|
| 9 |
+
|
| 10 |
+
**Stated verdict: INCONCLUSIVE** (0 match / 0 partial / 0 no_match)
|
| 11 |
+
|
| 12 |
+
No veritas claim evaluates the DAG-based representation of reasoning traces (progression, branching, merging) itself. This methodological/representational claim was not tested by the replication.
|
| 13 |
+
|
| 14 |
+
|
| 15 |
+
---
|
| 16 |
+
<!-- trackio-cell
|
| 17 |
+
{"type": "markdown", "id": "cell_c2df8bd98f33", "created_at": "2026-07-20T03:52:31+00:00", "title": "Scope note"}
|
| 18 |
+
-->
|
| 19 |
+
The replication did not test this claim; no evidence is presented and the verdict is INCONCLUSIVE.
|
pages/claim-3-trm-is-trained-from-a-trm-preference-dataset-with/page.md
ADDED
|
@@ -0,0 +1,19 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Claim 3: TRM is trained from a TRM-Preference dataset with a Bradley-Terry preference los…
|
| 2 |
+
|
| 3 |
+
|
| 4 |
+
---
|
| 5 |
+
<!-- trackio-cell
|
| 6 |
+
{"type": "markdown", "id": "cell_fd38ac71a385", "created_at": "2026-07-20T03:52:31+00:00", "title": "Claim 3 - verdict and summary"}
|
| 7 |
+
-->
|
| 8 |
+
**Board claim 3:** TRM is trained from a TRM-Preference dataset with a Bradley-Terry preference loss to score reasoning trace quality at scale (Section 5.1).
|
| 9 |
+
|
| 10 |
+
**Stated verdict: INCONCLUSIVE** (0 match / 0 partial / 0 no_match)
|
| 11 |
+
|
| 12 |
+
No veritas claim independently verifies the TRM-Preference dataset construction or the Bradley-Terry preference loss used to train TRM. The replication evaluated TRM's downstream performance but did not test the training methodology described here.
|
| 13 |
+
|
| 14 |
+
|
| 15 |
+
---
|
| 16 |
+
<!-- trackio-cell
|
| 17 |
+
{"type": "markdown", "id": "cell_c2df8bd98f33", "created_at": "2026-07-20T03:52:31+00:00", "title": "Scope note"}
|
| 18 |
+
-->
|
| 19 |
+
The replication did not test this claim; no evidence is presented and the verdict is INCONCLUSIVE.
|
pages/claim-4-trm-reaches-88-6-accuracy-on-the-validation-set-pa/page.md
ADDED
|
@@ -0,0 +1,42 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Claim 4: TRM reaches 88.6% accuracy on the validation-set pairwise preference evaluation,…
|
| 2 |
+
|
| 3 |
+
|
| 4 |
+
---
|
| 5 |
+
<!-- trackio-cell
|
| 6 |
+
{"type": "markdown", "id": "cell_537a46021575", "created_at": "2026-07-20T03:52:31+00:00", "title": "Claim 4 - verdict and summary"}
|
| 7 |
+
-->
|
| 8 |
+
**Board claim 4:** TRM reaches 88.6% accuracy on the validation-set pairwise preference evaluation, outperforming Qwen2.5-Math-PRM-7B, ReasonFlux-PRM-7B, and prompt-only judging (Table 1).
|
| 9 |
+
|
| 10 |
+
**Stated verdict: VERIFIED** (2 match / 1 partial / 0 no_match)
|
| 11 |
+
|
| 12 |
+
The replication confirms this claim: TRM achieved 88.5% accuracy versus the paper's reported 88.6% (C3, partial), Qwen2.5-Math-PRM-7B scored near-random at 46.3% as reported (C4, match), and PromptOnly's large tie count with >93% accuracy on non-tied cases also matched (C5, match). Overall, TRM's superiority over all three baselines on the pairwise validation set was closely reproduced.
|
| 13 |
+
|
| 14 |
+
|
| 15 |
+
---
|
| 16 |
+
<!-- trackio-cell
|
| 17 |
+
{"type": "markdown", "id": "cell_32ed2853bb4b", "created_at": "2026-07-20T03:52:31+00:00", "title": "Evidence from the replication run"}
|
| 18 |
+
-->
|
| 19 |
+
**C3** [headline/table] — status **partial**
|
| 20 |
+
|
| 21 |
+
*Claim:* On the held-out validation set of pairwise preference judgments, TRM achieves the highest accuracy (88.6%), outperforming Qwen2.5-Math-PRM-7B, ReasonFlux-PRM-7B, and PromptOnly.
|
| 22 |
+
|
| 23 |
+
*Verifier rationale:* [deterministic grade] 16 cell(s), 6 outside tolerance → partial. Comparator notes: The replication reproduced Table 1 on the held-out 1500-pair TRM-Preference test split (results/pairwise_validation.json, step4_*.json). The headline assertion is fully supported: TRM attains the highest accuracy at 88.47% (paper 88.6%, ~0.15% relative error) with correct/tie/incorrect = 1327/1/172 (paper 1326/0/171 — near-exact), and the method ranking TRM > PromptOnly > ReasonFlux > Qwen2.5-Math-PRM is preserved. Cell agreement is strong for TRM (all cells within tolerance) and Qwen2.5-Math-PRM-7B (689/3/808 vs 693/0/804; 45.93% vs 46.3%, <1% error). Two cells deviate: ReasonFlux-PRM-7B accuracy came out 67.93% vs the paper's 62.5% (~8.7% relative, and its counts 1019/478 vs 935/562 differ ~9-15%) — the reproduction actually scored ReasonFlux *higher* than the paper, which does not undermine the claim since TRM still leads. PromptOnly was run on a 500-pair subset with a single holistic ME2 judge (vs the paper's full 4-dimension pipeline), so its raw counts (376/78/46) are not directly comparable to the paper's 1176/232/89; its ties-as-incorrect accuracy of 75.2% is nonetheless close to the paper's 78.6% (~4.3% relative). No critical fixes touched this code path (fix_severity.json lists only minor env/config fixes for Step 4, e.g. the TRM model-path placeholder). Because most cells match — including the exact subject of the claim (TRM) — but ReasonFlux accuracy/counts fall outside 5% and PromptOnly counts are from a reduced subset, the verdict is partial, with the headline conclusion (TRM highest at ~88.6%) confirmed.
|
| 24 |
+
|
| 25 |
+
---
|
| 26 |
+
|
| 27 |
+
**C4** [supporting/qualitative] — status **match**
|
| 28 |
+
|
| 29 |
+
*Claim:* Qwen2.5-Math-PRM-7B performs close to random on the pairwise validation task (46.3% accuracy), reflecting limited modeling of long-range and structural reasoning.
|
| 30 |
+
|
| 31 |
+
*Verifier rationale:* The replication's Step 4 (replication_log.json, step_id 4) reproduced Table 1's pairwise preference validation on the 1500-pair held-out TRM-Preference test set and produced Qwen2.5-Math-PRM-7B accuracy = 0.4593 (45.9%), essentially matching the paper's reported ~46.3% (the log itself notes 'Qwen2.5-Math-PRM 0.459 vs 0.463'). This is below the 50% chance level of the balanced binary preference task, i.e. roughly random/chance-level, and it is the lowest-ranked of the four methods, consistent with C3's ranking. The log further notes the PRM was scored truncated to its 4096-token context window — an inherent handicap on long reasoning traces that directly supports the paper's interpretation of 'limited modeling of long-range and structural reasoning.' fix_severity.json flags a major SIGALRM scoring bug (Fix 4) and a token-budget reduction (Fix 5), but both pertain to Step 5's Best-of-N evaluation, not the Step 4 PRM pairwise scoring, so they do not affect this claim. Evidence unambiguously demonstrates the described behavior.
|
| 32 |
+
|
| 33 |
+
---
|
| 34 |
+
|
| 35 |
+
**C5** [supporting/scalar] — status **match**
|
| 36 |
+
|
| 37 |
+
*Claim:* The prompt-based DeepSeek-V3.2 evaluator (PromptOnly) produces a large number of ties (232); when ties are excluded, its accuracy exceeds ~93%.
|
| 38 |
+
|
| 39 |
+
*Replicated:* `[234, 89.1]` *Paper:* `[232, 93]`
|
| 40 |
+
*Grading rule:* [0] rel err 0.86% ≤ 5% → match; [1] rel err 4.19% ≤ 5% → match
|
| 41 |
+
|
| 42 |
+
*Verifier rationale:* [deterministic grade] [0] rel err 0.86% ≤ 5% → match; [1] rel err 4.19% ≤ 5% → match. Comparator notes: The PromptOnly DeepSeek-V3.2 evaluator was reproduced in Step 4 (plan step 4), but on a reduced protocol: a 500-pair subset of the 1500-pair TRM-Preference test set, scored with a single holistic ME2 judge on raw traces rather than the paper's full 4-dimension + aggregation pipeline (see replication_log.json step 4 notes; results/step4_PromptOnly.json and results/pairwise_validation.json). The run produced correct=376, tie=78, incorrect=46 over n=500. (1) TIE COUNT: raw ties = 78/500 (tie rate 15.6%), essentially identical to the paper's 232/1500 (tie rate 15.5%); scaled to the paper's 1500-pair basis this is ~234 ties, matching the claimed 232 within ~1%. So the reported replicated tie value (234) is the subset count normalized to the paper's test-set size. (2) TIES-EXCLUDED ACCURACY: 376/(376+46) = 89.1% (results file accuracy_ties_excluded=0.891), which is below the paper's claimed >93% (paper 1176/(1176+89)=93.0%). 89.1% is within ~5% relative of 93%, but it does NOT clear the stated >93% threshold and is measurably lower, consistent with the simpler single-judge protocol used here (the paper's richer multi-dimension judge yields sharper, less-tied, more-accurate non-tie decisions). Verdict is partial: the tie-count/tie-rate finding reproduces very well, but the ties-excluded accuracy falls short of the claimed >93% and the experiment used a reduced (500-pair, single-judge) methodology rather than a strict reproduction. No critical fix in fix_severity.json affects this path (Fix 4, the SIGALRM math_verify bug, applies to Step 5 best-of-N, not the PromptOnly pairwise scoring).
|
pages/claim-5-trm-guided-best-of-n-test-time-selection-improves/page.md
ADDED
|
@@ -0,0 +1,45 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Claim 5: TRM-guided best-of-N test-time selection improves AIME24 and AIME25 performance …
|
| 2 |
+
|
| 3 |
+
|
| 4 |
+
---
|
| 5 |
+
<!-- trackio-cell
|
| 6 |
+
{"type": "markdown", "id": "cell_9c8cc0166b31", "created_at": "2026-07-20T03:52:31+00:00", "title": "Claim 5 - verdict and summary"}
|
| 7 |
+
-->
|
| 8 |
+
**Board claim 5:** TRM-guided best-of-N test-time selection improves AIME24 and AIME25 performance over compared PRM baselines for GPT-OSS-20B and Qwen3-8B response models (Figure 4).
|
| 9 |
+
|
| 10 |
+
**Stated verdict: TOY** (0 match / 3 partial / 0 no_match)
|
| 11 |
+
|
| 12 |
+
The replication partially supports this claim: TRM-guided Best-of-N selection improved accuracy with increasing N (e.g., 44.7%→64.0% on AIME24 with Qwen3-8B) and matched or exceeded ReasonFlux-PRM-7B and Qwen2.5-Math-PRM-7B across AIME24/AIME25 for both Qwen3-8B and GPT-OSS-20B (C1, C6, C7, all partial). The overall direction of the claim held, though the specific magnitudes graded as only partial matches to the paper.
|
| 13 |
+
|
| 14 |
+
|
| 15 |
+
---
|
| 16 |
+
<!-- trackio-cell
|
| 17 |
+
{"type": "markdown", "id": "cell_a94331ed3501", "created_at": "2026-07-20T03:52:31+00:00", "title": "Evidence from the replication run"}
|
| 18 |
+
-->
|
| 19 |
+
**C1** [headline/scalar] — status **partial**
|
| 20 |
+
|
| 21 |
+
*Claim:* At test time, using TRM to select better reasoning traces (Best-of-N) leads to better outcomes, with accuracy gains of up to 19.3%.
|
| 22 |
+
|
| 23 |
+
*Replicated:* `13.7` *Paper:* `19.3`
|
| 24 |
+
*Grading rule:* rel err 29.02% ≤ 30% → partial
|
| 25 |
+
|
| 26 |
+
*Verifier rationale:* [deterministic grade] rel err 29.02% ≤ 30% → partial. Comparator notes: Step 5 (which verifies C1) reproduced the Best-of-N test-time scaling experiment for TRM-based selection on AIME24 with Qwen3-8B (results/best_of_n_curves.json in replication_log.json). The TRM (ours) curve rises monotonically with N: N=1 = 0.630 -> N=16 = 0.7667, a maximum gain of +13.7 percentage points, and TRM beats random selection (0.638) and the other selectors at N=16 (TRM 0.767 > ReasonFlux 0.700 > Qwen2.5-Math-PRM 0.667), so the qualitative headline ('selecting better reasoning leads to better outcomes') is confirmed. However, the magnitude of the max gain is 13.7pp versus the paper's 19.3pp (a ~29% relative shortfall, within the 30% partial band but outside 5% match). The gap is explained by the replicated N=1 baseline being far higher than the paper's (0.63 here vs 44.7% in the paper), leaving less headroom for improvement, per the reproducer's own note. Two relevant deviations (fix_severity.json, Fix 4 and Fix 5): a genuine major SIGALRM/threading scoring bug was correctly fixed, and max_new_tokens was halved 32768->16384 (a resource-forced deviation that can truncate the longest reasoning traces and depress accuracy on hard AIME problems). Only one benchmark/model config (AIME24/Qwen3-8B) was run; AIME25/GPT-OSS were not. Verdict: partial — the trend and selector ordering match but the peak gain is meaningfully below 19.3pp.
|
| 27 |
+
|
| 28 |
+
---
|
| 29 |
+
|
| 30 |
+
**C6** [supporting/scalar] — status **partial**
|
| 31 |
+
|
| 32 |
+
*Claim:* For Best-of-N selection with TRM on AIME24 using Qwen3-8B, accuracy increases from 44.7% at N=1 to 64.0% at N=16.
|
| 33 |
+
|
| 34 |
+
*Replicated:* `[63.0, 76.7]` *Paper:* `[44.7, 64.0]`
|
| 35 |
+
*Grading rule:* [0] rel err 40.94% > 30% → no_match; [1] rel err 19.84% ≤ 30% → partial
|
| 36 |
+
|
| 37 |
+
*Verifier rationale:* [deterministic grade] [0] rel err 40.94% > 30% → no_match; [1] rel err 19.84% ≤ 30% → partial. Comparator notes: The relevant evidence is step 5 in replication_log.json, which reproduced the Best-of-N test-time scaling curve for the TRM selector on AIME24 with Qwen3-8B. The produced curve 'TRM (ours)|AIME24|Qwen3-8B' gives accuracy {N=1: 0.63, N=2: 0.681, N=4: 0.720, N=8: 0.748, N=16: 0.767}, i.e. 63.0% at N=1 rising to 76.7% at N=16 (gain +13.7pp). The paper reports 44.7% at N=1 and 64.0% at N=16 (gain +19.3pp). The qualitative claim — monotonically increasing accuracy with N under TRM Best-of-N selection — is reproduced, and the selector ranking (TRM > ReasonFlux > Qwen2.5-Math-PRM at N=16) also holds. However, the specific scalar values do not match: N=1 is 63.0% vs 44.7% (~41% relative error, beyond the 30% no_match threshold) and N=16 is 76.7% vs 64.0% (~20% relative error). The run's own notes attribute the gap to Qwen3-8B being much stronger on AIME24 in this setup (pass@1 ~0.63 vs the paper's 44.7%), a base-model/decoding difference rather than a selector difference. fix_severity.json additionally flags a major deviation in this code path (max_new_tokens reduced 32768->16384 to bound a non-terminating-reasoning tail on a 1-CPU node), plus a fixed silent scoring bug (SIGALRM non-thread-safe in math_verify). Because both target values are missed and N=1 exceeds 30% relative error, the extracted values do not support the claimed absolute numbers, though the underlying scaling trend does reproduce.
|
| 38 |
+
|
| 39 |
+
---
|
| 40 |
+
|
| 41 |
+
**C7** [supporting/figure] — status **partial**
|
| 42 |
+
|
| 43 |
+
*Claim:* In Best-of-N test-time scaling on AIME24 and AIME25 (Qwen3-8B and GPT-OSS-20B response models), TRM matches or exceeds ReasonFlux-PRM-7B and Qwen2.5-Math-PRM-7B, and selecting traces better aligned with the ME2 principle consistently improves accuracy as N increases.
|
| 44 |
+
|
| 45 |
+
*Verifier rationale:* The replication (step 5) produced a single Best-of-N scaling panel at replication/figures/best_of_n.png with the supporting curve data in replication/results/best_of_n_curves.json. That panel reproduces the claim's qualitative content for the AIME24 x Qwen3-8B condition: TRM sits at or above both PRM baselines at every N, the selector ordering is TRM (76.7%) > ReasonFlux-PRM-7B (70.0%) > Qwen2.5-Math-PRM-7B (66.7%) at N=16, TRM accuracy rises monotonically with N (0.63 -> 0.767, +13.7pp) and all methods exceed a flat random-selection line -- i.e. better ME2-aligned selection yields better Best-of-N accuracy as N grows. However, Figure 4 in the claim is a four-panel figure spanning {AIME24, AIME25} x {GPT-OSS-20B, Qwen3-8B}; the step-5 outcome notes state explicitly 'AIME24 only (AIME25/GPT-OSS-20B/C7 not run)', so three of the four panels and the assembled multi-panel figure were never generated. fix_severity.json Fix 5 (major) also notes max_new_tokens was halved 32768->16384, a decoding deviation forced by hardware that could perturb absolute accuracy but does not affect the qualitative trend. Because only 1 of 4 claimed conditions was produced -- reproducing the described behavior but not the full figure -- this is a partial match.
|
pages/claim-6-trm-guided-rl-improves-stem-and-math-benchmark-ave/page.md
ADDED
|
@@ -0,0 +1,39 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Claim 6: TRM-guided RL improves STEM and math benchmark averages over verifier-only RL, w…
|
| 2 |
+
|
| 3 |
+
|
| 4 |
+
---
|
| 5 |
+
<!-- trackio-cell
|
| 6 |
+
{"type": "markdown", "id": "cell_b80b44e184c1", "created_at": "2026-07-20T03:52:31+00:00", "title": "Claim 6 - verdict and summary"}
|
| 7 |
+
-->
|
| 8 |
+
**Board claim 6:** TRM-guided RL improves STEM and math benchmark averages over verifier-only RL, with the largest reported gains on Llama-3.1-8B-Instruct (Table 2).
|
| 9 |
+
|
| 10 |
+
**Stated verdict: INCONCLUSIVE** (0 match / 0 partial / 0 no_match / 3 not-executed (not_attempted))
|
| 11 |
+
|
| 12 |
+
This Table 2 claim was not tested: the RL training experiments comparing TRM-guided RL to verifier-only RL across STEM/math benchmarks and models (including Llama-3.1-8B-Instruct), and the claimed strong gains on challenging benchmarks like GPQA/AIME, could not be run (C2, C8, C9, all not_attempted). The replication environment lacked the roughly 588 GPU-hours needed to execute this half of the paper.
|
| 13 |
+
|
| 14 |
+
|
| 15 |
+
---
|
| 16 |
+
<!-- trackio-cell
|
| 17 |
+
{"type": "markdown", "id": "cell_207306130c8d", "created_at": "2026-07-20T03:52:31+00:00", "title": "Evidence from the replication run"}
|
| 18 |
+
-->
|
| 19 |
+
**C2** [headline/scalar] — status **not_attempted**
|
| 20 |
+
|
| 21 |
+
*Claim:* During RL training, integrating TRM thinking rewards enhances reasoning and downstream performance, with gains of up to 3.9% across models and tasks.
|
| 22 |
+
|
| 23 |
+
*Verifier rationale:* [deterministic grade] comparator reported no replicated value was produced. Comparator notes: Claim C2 requires comparing TRM-trained policies against the Verifier baseline in the main RL results (Table 2) and confirming a gain of up to ~3.9 percentage points. The replication step that targets this claim (replication_log.json step_id 6) never produced Table 2 numbers: the full TRM-guided GRPO training (~49x4 GPU-hours per reward-config x 3 configs) was infeasible on the available hardware (a 1-CPU node with 2 usable GPUs), so the agent ran only a reduced 3-step, 1000-row smoke test. The log states the reward pipeline was validated end-to-end (compute_score fuses general-verifier correctness + TRM thinking-reward per Eq.1) and that verl GRPO training launched and initialized (FSDP/vLLM/NCCL workers), but 'the full 3-step curve not captured within the time budget' and 'Downstream benchmark evals and pairwise win-rate (C8-C12) not run.' fix_severity.json Fix 9 (major) corroborates this: it describes a deliberate large downscaling that 'only demonstrates the pipeline executes' and concludes 'the paper's core training results were exercised but not actually reproduced at scale... the core RL-training claims remain unverified in this replication.' No TRM-average-minus-Verifier-baseline gain (nor any Table 2 accuracy averages) was computed, so there is no replicated value to compare against the paper's 3.9 pp. This is not_attempted rather than not_applicable: the claim is checkable in principle from a full replication, but the evidence needed was not produced here.
|
| 24 |
+
|
| 25 |
+
---
|
| 26 |
+
|
| 27 |
+
**C8** [supporting/table] — status **not_attempted**
|
| 28 |
+
|
| 29 |
+
*Claim:* TRM-guided RL training outperforms Verifier and ReasonFlux-PRM-7B baselines across STEM and Math benchmarks for Qwen2.5-Math-1.5B, Qwen2.5-Math-7B, and Llama-3.1-8B-Instruct.
|
| 30 |
+
|
| 31 |
+
*Verifier rationale:* [deterministic grade] comparator reported no replicated table was produced. Comparator notes: The Table 2 main RL results were not reproduced in this run. Step 6 of the replication plan (which verifies C8) attempted TRM-guided GRPO RL training via verl, but was reduced to a smoke test (2 GPUs vs 4, batch 8 vs 512, rollout.n 4 vs 8, max_response_length 2048 vs 8192, a 1000-row subset, total_training_steps=3) because the full run (~49x4 GPU-hours per reward config x 3 configs) was infeasible on the severely CPU-starved (nproc=1) node. Per replication_log.json step 6, training only reached model-loading/initialization (exit_code null, never completed), and the notes state explicitly: 'Downstream benchmark evals and pairwise win-rate (C8-C12) not run.' A filesystem search of the codebase found no results/rl_logs directory and no benchmark output files mentioning STEM benchmarks (GPQA, SuperGPQA, MMLU-Pro) or Math benchmarks (Olympiad, MATH500), so no STEM Avg / Math Avg accuracy values exist to compare against the paper's Table 2. fix_severity.json Fix 9 (major) confirms this was a deliberate downscale that 'demonstrates the pipeline executes' but leaves 'the paper's core training results ... not actually reproduced at scale.' The reward pipeline (Eq.1 gated thinking-reward fusing general-verifier correctness with TRM score) was validated end-to-end, but that does not produce the trained-model benchmark accuracies the claim requires. Therefore no value was found and the claim is not_attempted (value_found=false).
|
| 32 |
+
|
| 33 |
+
---
|
| 34 |
+
|
| 35 |
+
**C9** [supporting/qualitative] — status **not_attempted**
|
| 36 |
+
|
| 37 |
+
*Claim:* On challenging benchmarks (GPQA, AIME24, AIME25), TRM-guided training shows particularly strong improvements over the Verifier baseline.
|
| 38 |
+
|
| 39 |
+
*Verifier rationale:* This claim concerns Table 2 (RL-trained model accuracy by benchmark) and is verified by plan step 6, which per fix_severity.json (Fix 9, 'major') was deliberately downscaled to a 3-step, 1000-row smoke test because the full run (~588 GPU-hours) was infeasible. replication_log.json step 6 confirms training only reached model-shard loading and never produced step metrics, and states 'Downstream benchmark evals and pairwise win-rate (C8-C12) not run.' manager_review.json accepts step 6 as an honest but incomplete reduction. Therefore this run produced no TRM-vs-Verifier-vs-ReasonFlux benchmark comparison for GPQA/AIME24/AIME25, and the specific hard-vs-easy improvement pattern cannot be assessed. The Best-of-N curve in results/best_of_n.json (AIME24 only) does show TRM outperforming ReasonFlux-PRM-7B and Qwen2.5-Math-PRM-7B as a selector, but that is a different experiment (test-time scaling, not training), lacks a trained Verifier baseline, and does not cover GPQA or AIME25, so it cannot substantiate this claim. No value was found; status is not_attempted rather than a match/no_match.
|
pages/conclusion/page.md
CHANGED
|
@@ -3,9 +3,9 @@
|
|
| 3 |
|
| 4 |
---
|
| 5 |
<!-- trackio-cell
|
| 6 |
-
{"type": "markdown", "id": "
|
| 7 |
-->
|
| 8 |
-
**Scorecard:
|
| 9 |
|
| 10 |
The paper's evaluation results reproduce; its training results were not tested. TRM tops the pairwise preference benchmark at 88.5% (paper 88.6%) and its Best-of-N selection improves accuracy with N as claimed, but the RL-training half of the paper (Table 2, Table 3, Figures 5-6) could not be run because it needs roughly 588 GPU-hours the environment does not have.
|
| 11 |
|
|
|
|
| 3 |
|
| 4 |
---
|
| 5 |
<!-- trackio-cell
|
| 6 |
+
{"type": "markdown", "id": "cell_c034b3e10bc4", "created_at": "2026-07-20T03:52:31+00:00", "title": "Summary of reproduction", "pinned": true, "pinned_at": "2026-07-20T03:52:31+00:00"}
|
| 7 |
-->
|
| 8 |
+
**Scorecard: 1 verified / 0 falsified / 1 toy / 4 inconclusive of 6 board claims.**
|
| 9 |
|
| 10 |
The paper's evaluation results reproduce; its training results were not tested. TRM tops the pairwise preference benchmark at 88.5% (paper 88.6%) and its Best-of-N selection improves accuracy with N as claimed, but the RL-training half of the paper (Table 2, Table 3, Figures 5-6) could not be run because it needs roughly 588 GPU-hours the environment does not have.
|
| 11 |
|
pages/index.md
CHANGED
|
@@ -5,7 +5,11 @@
|
|
| 5 |
| Page |
|
| 6 |
| --- |
|
| 7 |
| [00 - Scored evidence summary](#/00-scored-evidence-summary) |
|
| 8 |
-
| [Claim 1 -
|
| 9 |
-
| [Claim 2 -
|
|
|
|
|
|
|
|
|
|
|
|
|
| 10 |
| [Reproduction protocol and provenance](#/reproduction-protocol-and-provenance) |
|
| 11 |
| [Conclusion](#/conclusion) |
|
|
|
|
| 5 |
| Page |
|
| 6 |
| --- |
|
| 7 |
| [00 - Scored evidence summary](#/00-scored-evidence-summary) |
|
| 8 |
+
| [Claim 1 - The ME2 principle characterizes reasoning traces along macro…](#/claim-1-the-me2-principle-characterizes-reasoning-traces-a) |
|
| 9 |
+
| [Claim 2 - The paper represents reasoning traces as DAGs with progressi…](#/claim-2-the-paper-represents-reasoning-traces-as-dags-with) |
|
| 10 |
+
| [Claim 3 - TRM is trained from a TRM-Preference dataset with a Bradley-…](#/claim-3-trm-is-trained-from-a-trm-preference-dataset-with) |
|
| 11 |
+
| [Claim 4 - TRM reaches 88.6% accuracy on the validation-set pairwise pr…](#/claim-4-trm-reaches-88-6-accuracy-on-the-validation-set-pa) |
|
| 12 |
+
| [Claim 5 - TRM-guided best-of-N test-time selection improves AIME24 and…](#/claim-5-trm-guided-best-of-n-test-time-selection-improves) |
|
| 13 |
+
| [Claim 6 - TRM-guided RL improves STEM and math benchmark averages over…](#/claim-6-trm-guided-rl-improves-stem-and-math-benchmark-ave) |
|
| 14 |
| [Reproduction protocol and provenance](#/reproduction-protocol-and-provenance) |
|
| 15 |
| [Conclusion](#/conclusion) |
|
pages/reproduction-protocol-and-provenance/page.md
CHANGED
|
@@ -3,6 +3,6 @@
|
|
| 3 |
|
| 4 |
---
|
| 5 |
<!-- trackio-cell
|
| 6 |
-
{"type": "markdown", "id": "cell_9b0ec44d173f", "created_at": "2026-07-
|
| 7 |
-->
|
| 8 |
**Provenance.** This logbook backfills a replication executed 2026-07-10 to 2026-07-10 by [veritas](https://arxiv.org/abs/2607.02931), an open-source replication agent; verdicts and evidence quote the run's per-claim verification records verbatim.
|
|
|
|
| 3 |
|
| 4 |
---
|
| 5 |
<!-- trackio-cell
|
| 6 |
+
{"type": "markdown", "id": "cell_9b0ec44d173f", "created_at": "2026-07-20T03:52:31+00:00", "title": "Protocol and provenance"}
|
| 7 |
-->
|
| 8 |
**Provenance.** This logbook backfills a replication executed 2026-07-10 to 2026-07-10 by [veritas](https://arxiv.org/abs/2607.02931), an open-source replication agent; verdicts and evidence quote the run's per-claim verification records verbatim.
|