Add model card, links to paper and official repository

#1
by nielsr HF Staff - opened
Files changed (1) hide show
  1. README.md +35 -0
README.md ADDED
@@ -0,0 +1,35 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: transformers
4
+ pipeline_tag: text-generation
5
+ base_model: meta-llama/Llama-3.1-8B-Instruct
6
+ ---
7
+
8
+ # SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering
9
+
10
+ This repository contains the fine-tuned Llama-3.1-8B model checkpoint developed using the SPADER reinforcement learning framework, as presented in the paper [SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering](https://huggingface.co/papers/2606.00593).
11
+
12
+ ## Model Description
13
+
14
+ SPADER is a reinforcement learning framework designed for long-horizon tool-use agents in Multi-Answer QA. It introduces:
15
+ - **Step-wise Peer Advantage (SPA)**: A critic-free step-level credit assignment mechanism that aligns parallel trajectories by decision step and estimates advantages from peer returns.
16
+ - **Diversity-Aware Exploration Reward**: Promotes long-tail entity discovery by upweighting rare findings and downweighting redundant ones.
17
+
18
+ This checkpoint represents the Llama-3.1-8B-Instruct base model trained with SPADER.
19
+
20
+ - **Repository:** [KhanCold/spader](https://github.com/KhanCold/spader)
21
+ - **Paper:** [arXiv:2606.00593](https://arxiv.org/abs/2606.00593)
22
+
23
+ ## Citation
24
+
25
+ ```bibtex
26
+ @misc{shi2026spaderstepwisepeeradvantage,
27
+ title={SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering},
28
+ author={Qiming Shi and Zhaolu Kang and Yunfan Zhou and Di Weng and Yingcai Wu},
29
+ year={2026},
30
+ eprint={2606.00593},
31
+ archivePrefix={arXiv},
32
+ primaryClass={cs.CL},
33
+ url={https://arxiv.org/abs/2606.00593},
34
+ }
35
+ ```