File size: 16,639 Bytes
d18fc11 48d49eb d18fc11 48d49eb d18fc11 48d49eb d18fc11 48d49eb 453ac47 48d49eb d18fc11 48d49eb f5c2e95 48d49eb d18fc11 48d49eb d18fc11 48d49eb d18fc11 48d49eb f5c2e95 48d49eb d18fc11 48d49eb f5c2e95 48d49eb d18fc11 48d49eb f5c2e95 48d49eb d18fc11 48d49eb d18fc11 48d49eb d18fc11 48d49eb d18fc11 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 | <!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<meta name="description" content="RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning.">
<meta name="keywords" content="RISE-RL, reinforcement learning, large language models, rubric, selective exploration">
<meta name="author" content="Jinkun Hou, Zhuo Liu, Huimin Ren, Hongsheng Xin, Pan Zhou, Kun Zhan">
<title>RISE-RL | Rubric-Informed Selective Exploration</title>
<style>
:root { --ink:#152033; --muted:#5b6879; --line:#dce3ec; --soft:#f5f8fc; --accent:#5146d8; --accent-dark:#352aa5; }
* { box-sizing:border-box; }
html { scroll-behavior:smooth; }
body { margin:0; color:var(--ink); background:#fff; font-family:Inter,ui-sans-serif,system-ui,-apple-system,BlinkMacSystemFont,"Segoe UI",sans-serif; line-height:1.65; }
a { color:var(--accent); }
img { display:block; max-width:100%; }
.wrap { width:min(1160px,calc(100% - 40px)); margin:0 auto; }
.topbar { position:sticky; z-index:10; top:0; border-bottom:1px solid rgba(220,227,236,.8); background:rgba(255,255,255,.9); backdrop-filter:blur(14px); }
.topbar .wrap { display:flex; align-items:center; justify-content:space-between; min-height:58px; }
.brand { color:var(--ink); font-weight:850; text-decoration:none; letter-spacing:-.02em; }
.nav { display:flex; gap:20px; }
.nav a { color:var(--muted); font-size:.92rem; font-weight:700; text-decoration:none; }
.hero { position:relative; overflow:hidden; padding:72px 0 68px; border-bottom:1px solid var(--line); background:radial-gradient(circle at 10% 10%,rgba(255,183,51,.25),transparent 28%),radial-gradient(circle at 90% 8%,rgba(81,70,216,.22),transparent 31%),linear-gradient(145deg,#fffdf8,#f3f4ff 56%,#eff8ff); }
.hero::after { content:""; position:absolute; inset:0; pointer-events:none; background-image:linear-gradient(rgba(21,32,51,.03) 1px,transparent 1px),linear-gradient(90deg,rgba(21,32,51,.03) 1px,transparent 1px); background-size:28px 28px; mask-image:linear-gradient(to bottom,#000,transparent 92%); }
.hero-grid { position:relative; z-index:1; display:grid; grid-template-columns:.92fr 1.08fr; gap:45px; align-items:center; }
.eyebrow { display:inline-flex; padding:7px 12px; border:1px solid rgba(81,70,216,.24); border-radius:999px; color:var(--accent-dark); background:rgba(255,255,255,.78); font-size:.8rem; font-weight:800; letter-spacing:.06em; text-transform:uppercase; }
h1 { margin:20px 0 10px; font-size:clamp(3.1rem,7vw,6.2rem); line-height:.98; letter-spacing:-.065em; }
.rise { color:var(--accent-dark); }
.subtitle { margin:0 0 22px; color:#354258; font-size:1.17rem; line-height:1.5; }
.authors { font-weight:690; }
.authors span { display:inline-block; margin:2px 7px 2px 0; }
.affiliations { margin-top:8px; color:var(--muted); font-size:.9rem; }
.links { display:flex; flex-wrap:wrap; gap:10px; margin-top:27px; }
.button { display:inline-flex; align-items:center; justify-content:center; padding:10px 17px; border:1px solid var(--ink); border-radius:999px; color:#fff; background:var(--ink); font-weight:780; text-decoration:none; transition:transform .16s ease,background .16s ease; }
.button:hover { transform:translateY(-2px); background:var(--accent-dark); }
.button.secondary { color:var(--ink); background:#fff; }
.hero-figure,.paper-figure { margin:0; overflow:hidden; border:1px solid var(--line); border-radius:20px; background:#fff; box-shadow:0 18px 55px rgba(31,47,75,.13); }
.hero-figure img,.wide-figure img { width:100%; }
.figure-placeholder {
display:grid;
place-items:center;
min-height:320px;
padding:34px;
color:var(--muted);
background:repeating-linear-gradient(135deg,#f8faff,#f8faff 14px,#f1f4fa 14px,#f1f4fa 28px);
text-align:center;
}
.figure-placeholder strong { display:block; margin-bottom:7px; color:var(--ink); font-size:1.05rem; }
.figure-placeholder code { display:inline-block; margin-top:10px; padding:3px 8px; border-radius:6px; color:var(--accent-dark); background:#fff; }
figcaption { padding:13px 17px 15px; color:var(--muted); font-size:.88rem; line-height:1.45; }
main { width:min(1060px,calc(100% - 40px)); margin:0 auto; padding:68px 0 80px; }
section { margin-bottom:74px; scroll-margin-top:85px; }
.section-kicker { margin-bottom:5px; color:var(--accent); font-size:.79rem; font-weight:850; letter-spacing:.08em; text-transform:uppercase; }
h2 { margin:0 0 20px; font-size:clamp(1.85rem,3.4vw,2.65rem); line-height:1.15; letter-spacing:-.035em; }
.lead { color:#2e3b4e; font-size:1.08rem; text-align:justify; }
.lead p { margin:0 0 16px; }
.takeaways { display:grid; grid-template-columns:repeat(3,1fr); gap:16px; margin-top:28px; }
.card { padding:23px; border:1px solid var(--line); border-radius:17px; background:var(--soft); }
.card .metric { color:var(--accent-dark); font-size:1.65rem; font-weight:880; letter-spacing:-.03em; }
.card h3 { margin:4px 0 7px; font-size:1rem; }
.card p { margin:0; color:var(--muted); font-size:.93rem; }
.wide-figure { margin-top:28px; }
.explain-grid { display:grid; grid-template-columns:1fr 1fr; gap:17px; margin-top:23px; }
.explain { padding:21px; border-left:4px solid var(--accent); border-radius:0 14px 14px 0; background:var(--soft); }
.explain strong { display:block; margin-bottom:6px; }
.explain p { margin:0; color:var(--muted); font-size:.95rem; }
.result-note { display:grid; grid-template-columns:repeat(3,1fr); gap:12px; margin-top:19px; }
.result-note div { padding:16px 18px; border-radius:13px; color:#fff; background:linear-gradient(135deg,#362ba6,#6257df); }
.result-note b { display:block; font-size:1.25rem; }
.dynamics { display:grid; grid-template-columns:1.06fr .94fr; gap:28px; align-items:center; }
.dynamics .paper-figure { box-shadow:0 12px 35px rgba(31,47,75,.08); }
.dynamics-copy ul { padding-left:1.2rem; color:#344157; }
.dynamics-copy li + li { margin-top:10px; }
pre { overflow-x:auto; padding:22px; border:1px solid var(--line); border-radius:15px; color:#dce7ff; background:#131b2b; font:.9rem/1.55 ui-monospace,SFMono-Regular,Menlo,Monaco,Consolas,monospace; }
footer { padding:35px 20px 44px; border-top:1px solid var(--line); color:var(--muted); background:var(--soft); text-align:center; }
footer p { margin:5px 0; }
@media (max-width:900px) { .hero-grid,.dynamics { grid-template-columns:1fr; } .hero-copy { text-align:center; } .links { justify-content:center; } .hero-figure { max-width:690px; margin:0 auto; } }
@media (max-width:700px) { .nav { display:none; } .hero { padding:48px 0 52px; } .takeaways,.explain-grid,.result-note { grid-template-columns:1fr; } main { padding-top:48px; } section { margin-bottom:58px; } }
</style>
</head>
<body>
<header class="topbar">
<div class="wrap">
<a class="brand" href="#top">RISE-RL</a>
<nav class="nav" aria-label="Page sections">
<a href="#abstract">Abstract</a><a href="#method">Method</a><a href="#results">Results</a><a href="#dynamics">Dynamics</a><a href="#citation">Citation</a>
</nav>
</div>
</header>
<header class="hero" id="top">
<div class="wrap hero-grid">
<div class="hero-copy">
<div class="eyebrow">arXiv:2608.09123 · Research Project</div>
<h1><span class="rise">RISE-RL</span></h1>
<p class="subtitle">Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning</p>
<div class="authors">
<span>Jinkun Hou<sup>1*</sup>,</span><span>Zhuo Liu<sup>2*</sup>,</span><span>Huimin Ren<sup>3†</sup>,</span><span>Hongsheng Xin<sup>3</sup>,</span><span>Pan Zhou<sup>3</sup>,</span><span>Kun Zhan<sup>3</sup></span>
</div>
<div class="affiliations">
<div>
<sup>1</sup>Peking University ·
<sup>2</sup>Beijing Institute of Technology ·
<sup>3</sup>Li Auto Inc.
</div>
<div style="margin-top: 10px;">
<sup>*</sup> This work was done during an internship at Li Auto Inc.
</div>
<div>
<sup>†</sup> Corresponding author
</div>
</div>
<nav class="links" aria-label="Project links">
<a class="button" href="https://arxiv.org/pdf/2608.09123" target="_blank" rel="noopener">Paper PDF</a>
<a class="button secondary" href="https://arxiv.org/abs/2608.09123" target="_blank" rel="noopener">arXiv</a>
<a class="button secondary" href="https://huggingface.co/papers/2608.09123" target="_blank" rel="noopener">HF Paper</a>
</nav>
</div>
<figure class="hero-figure">
<!-- Export Figure 1 from your local PDF and replace this div with:
<img src="./static/images/overview.jpg" alt="RISE-RL guidance and exploration overview">
-->
<img src="./static/images/overview.jpg"
alt="RISE-RL guidance and exploration overview">
<figcaption><strong>Guidance expands the reachable policy space.</strong> RISE-RL targets repeatedly missed rubric criteria early, then returns to autonomous on-policy exploration.</figcaption>
</figure>
</div>
</header>
<main>
<section id="abstract">
<div class="section-kicker">Motivation</div>
<h2>Selective guidance for open-ended alignment</h2>
<div class="lead">
<p>Aligning Large Language Models (LLMs) for open-ended tasks is challenging because responses must satisfy multidimensional criteria without following a single correct generation trajectory. Existing rubric-based reinforcement learning methods compress fine-grained criterion-level feedback into scalar rewards, making persistent capability gaps difficult to target under limited on-policy exploration.</p>
<p>RISE-RL uses repeatedly missed rubric criteria to elicit privileged trajectories that are difficult to discover through unguided exploration alone. It retains only trajectories whose complete-rubric reward exceeds the mean reward of natural rollouts, and re-evaluates them under the original prompt to emphasize behaviors that remain weakly supported by the natural policy.</p>
</div>
<div class="takeaways">
<article class="card"><div class="metric">Target</div><h3>Criterion-level gaps</h3><p>Priority is assigned to important rubric criteria that repeatedly fail across natural rollouts.</p></article>
<article class="card"><div class="metric">Filter</div><h3>Reward-improving guidance</h3><p>Only privileged trajectories that beat the natural-rollout mean are retained.</p></article>
<article class="card"><div class="metric">Remove</div><h3>Return to exploration</h3><p>Guidance is withdrawn after its reward advantage saturates, avoiding permanent dependence.</p></article>
</div>
</section>
<section id="method">
<div class="section-kicker">Method</div>
<h2>The RISE-RL framework</h2>
<p class="lead">RISE-RL augments a natural-prompt GRPO branch with a selective rubric-guidance branch used only during early training. Guided candidates are judged with the complete rubric, filtered by reward gain, and forwarded again under the natural prompt. Policy-support shaping then emphasizes useful tokens that still have low probability under the natural policy.</p>
<figure class="paper-figure wide-figure">
<!-- Export Figure 2 from your local PDF and replace this div with:
<img src="./static/images/framework.jpg" alt="Full RISE-RL algorithm framework">
-->
<img src="./static/images/framework.jpg"
alt="Full RISE-RL algorithm framework">
<figcaption><strong>Figure 2.</strong> Natural-prompt GRPO is paired with an early selective-guidance branch. The auxiliary guidance loss is removed after the reward advantage saturates.</figcaption>
</figure>
<div class="explain-grid">
<div class="explain"><strong>Criterion-level selective feedback</strong><p>Failure frequency and rubric importance jointly determine the highest-priority criteria for each prompt.</p></div>
<div class="explain"><strong>Quality-filtered guidance</strong><p>Guided responses must outperform the mean reward of natural rollouts under the complete original rubric.</p></div>
<div class="explain"><strong>Policy-support shaping</strong><p>Tokens with low support under the natural prompt receive stronger auxiliary weights, focusing learning on hard-to-discover behavior.</p></div>
<div class="explain"><strong>Decoupled optimization</strong><p>The guidance objective is separated from the natural GRPO group and switched off once its marginal benefit plateaus.</p></div>
</div>
</section>
<section id="results">
<div class="section-kicker">Main Results</div>
<h2>Strong gains across four open-ended domains</h2>
<p class="lead">Across Writing, Chat, Health, and Science, RISE-RL achieves the highest average score at both the Qwen3-4B and Qwen3-14B scales. The improvements extend beyond subjective open-ended evaluation to objectively scored medical and scientific benchmarks.</p>
<figure class="paper-figure wide-figure">
<!-- Export Table 1 from your local PDF and replace this div with:
<img src="./static/images/main-results.png" alt="Main results across four open-ended domains">
-->
<img src="./static/images/main-results.png"
alt="Main results across four open-ended domains">
<figcaption><strong>Table 1.</strong> Main results across four domains. Green values denote gains over the corresponding initial model; bold rows show RISE-RL.</figcaption>
</figure>
<div class="result-note">
<div><b>57.49</b>Qwen3-4B average, +10.25 over the initial model.</div>
<div><b>70.52</b>Qwen3-14B average, +13.40 over the initial model.</div>
<div><b>+6.02</b>CreativeWriting-V3 gain at 14B, where standard Rubric-RL is nearly flat.</div>
</div>
</section>
<section id="dynamics">
<div class="section-kicker">Training Dynamics</div>
<h2>Why guidance should be removed</h2>
<div class="dynamics">
<figure class="paper-figure">
<!-- Export Figure 3 from your local PDF and replace this div with:
<img src="./static/images/guidance-comparison.jpg" alt="Rubric-RL, persistent guidance, and RISE-RL comparison">
-->
<img src="./static/images/guidance-comparison.jpg"
alt="Rubric-RL, persistent guidance, and RISE-RL comparison">
<figcaption><strong>Figure 3.</strong> Rubric-RL, persistent guidance, and RISE-RL on Writing and Health. Dashed vertical lines indicate guidance removal.</figcaption>
</figure>
<div class="dynamics-copy">
<p class="lead">The comparison reveals a two-stage learning pattern:</p>
<ul>
<li><strong>Rubric-RL:</strong> entropy grows slowly, reflecting support-limited unguided exploration.</li>
<li><strong>Persistent guidance:</strong> entropy remains high and unstable; performance can plateau or degrade when the policy stays dependent on external rubrics.</li>
<li><strong>RISE-RL:</strong> early guidance expands the policy distribution, then removal consolidates the discovered behavior through autonomous exploration.</li>
</ul>
<p>After guidance is removed, benchmark performance continues to improve while entropy stabilizes above the Rubric-RL baseline. This supports the central design choice: use guidance to discover behavior, not as a permanent training crutch.</p>
</div>
</div>
</section>
<section id="citation">
<div class="section-kicker">Reference</div>
<h2>Citation</h2>
<pre><code>@article{hou2026riserl,
title = {RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning},
author = {Hou, Jinkun and Liu, Zhuo and Ren, Huimin and Xin, Hongsheng and Zhou, Pan and Zhan, Kun},
journal = {arXiv preprint arXiv:2608.09123},
year = {2026}
}</code></pre>
</section>
</main>
<footer>
<p><strong>RISE-RL</strong> · Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning</p>
<p>Figures and results are reproduced from <a href="https://arxiv.org/abs/2608.09123" target="_blank" rel="noopener">arXiv:2608.09123</a>.</p>
</footer>
</body>
</html>
|