ABForge-Qwen3-8B-RL

📄 ArXiv  ï½œ  💻 Code  ï½œ  🤗 Collection

About

This repository contains the RL-only ablation of ABForge, presented in ABForge: Post-Training for Paper-Grounded Ablation Design. Given a paper's methodology with its ablation content removed, ABForge proposes the ablation objectives the paper should investigate and designs a rigorous experiment plan for each — both from a single checkpoint.

This model applies rubric-guided GRPO directly to Qwen3-8B for 200 updates with no SFT warm start, on a 1:1 mixture of the two tasks with each rollout routed to its task's reward by data_source. It is the RL only row of the paper's post-training ablation; the released model ABForge-Qwen3-8B runs the same RL stage from the SFT checkpoint instead.

Links

Resource Link
Code SlowGuess/Abforge_1
Training & evaluation data SlowGuess/abforge-data
Released model (SFT → GRPO) SlowGuess/ABForge-Qwen3-8B
SFT checkpoint SlowGuess/ABForge-Qwen3-8B-SFT
Per-paper outputs & judge rationales outputs/task{1,2}/*/abforge-rl.jsonl in the data repo

Training data

SlowGuess/abforge-data ships a single table under train/, one row per paper, built by a semi-automated audit-in-the-loop pipeline over research papers from major ML, NLP and CV venues. This model trains on the rows flagged in_rl_task1 and in_rl_task2 — 30,000 papers each, disjoint from the SFT pool — mixed 1:1. The benchmark papers carry no training flag, so they cannot leak in.

Performance

AblationBench, automated rubric-based LLM-as-a-Judge evaluation (eval/ablationbench_200.jsonl, 200 papers, judge claude-sonnet-4-6). Task 1 is ablation objective identification (paper_score); Task 2 is ablation plan synthesis (design_score, ×100).

Model Task 1 Task 2
Qwen3-8B (base) 44.4 43.4
ABForge-Qwen3-8B-SFT (SFT only) 30.7 52.2
ABForge-Qwen3-8B-RL (this model, RL only) 52.2 54.9
ABForge-Qwen3-8B (SFT → GRPO) 55.9 62.4

GRPO alone already lifts both tasks over the base model, and unlike SFT it does not trade one for the other. Warm-starting the same RL stage from the SFT checkpoint is still worth +3.7 on Task 1 and +7.5 on Task 2 — SFT is an effective RL initialization even though it does not improve Task 1 on its own.

Evaluation

Reproduce the numbers above with the code release:

git clone https://github.com/SlowGuess/Abforge_1 && cd Abforge_1
huggingface-cli download SlowGuess/abforge-data --repo-type dataset \
  --include "eval/*" --local-dir data

python run_inference_local.py --task 1 \
  --input data/eval/ablationbench_200.jsonl \
  --output outputs/task1_infer.jsonl \
  --model-path SlowGuess/ABForge-Qwen3-8B-RL \
  --dtype bf16 --device-map auto \
  --max-new-tokens 5120 --temperature 0.0 --stop-on '</Result>'

export JUDGE_API_BASE=https://api.openai.com/v1
export JUDGE_API_KEY=...
export JUDGE_MODEL=...
scripts/evaluate_task1.sh outputs/task1_infer.jsonl

Swap --task 2, --stop-on '</Proposed_Plan>' and scripts/evaluate_task2.sh for Task 2. The model is trained on the prompt templates in the code release and the rubric evaluator expects the matching output structure, so use those templates and greedy decoding.

Citation

@misc{abforge2026,
  title={ABForge: Post-Training for Paper-Grounded Ablation Design},
  author={TODO},
  year={2026},
}
Downloads last month
351
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SlowGuess/ABForge-Qwen3-8B-RL

Finetuned
Qwen/Qwen3-8B
Finetuned
(1997)
this model

Dataset used to train SlowGuess/ABForge-Qwen3-8B-RL

Collection including SlowGuess/ABForge-Qwen3-8B-RL