Buckets:

hf-doc-build/doc-dev / openenv /pr_1028 /en /tutorials /opencode-agent-grpo.html
download
raw
13.1 kB
<meta charset="utf-8" /><meta name="hf:doc:metadata" content="{&quot;title&quot;:&quot;Coding Agent Training with TRL (OpenCode)&quot;,&quot;local&quot;:&quot;coding-agent-training-with-trl-opencode&quot;,&quot;sections&quot;:[{&quot;title&quot;:&quot;How It Works&quot;,&quot;local&quot;:&quot;how-it-works&quot;,&quot;sections&quot;:[],&quot;depth&quot;:2},{&quot;title&quot;:&quot;Full Recipe&quot;,&quot;local&quot;:&quot;full-recipe&quot;,&quot;sections&quot;:[],&quot;depth&quot;:2}],&quot;depth&quot;:1}"/>
<link href="/docs/openenv/pr_1028/en/_app/immutable/entry/start.Bs7kuoo8.js" rel="modulepreload">
<link href="/docs/openenv/pr_1028/en/_app/immutable/chunks/CNOAEpFk.js" rel="modulepreload">
<link href="/docs/openenv/pr_1028/en/_app/immutable/chunks/BSuxAqoA.js" rel="modulepreload">
<link href="/docs/openenv/pr_1028/en/_app/immutable/entry/app.D4PcSJ6S.js" rel="modulepreload">
<link href="/docs/openenv/pr_1028/en/_app/immutable/chunks/Ds_1Ru_R.js" rel="modulepreload">
<link href="/docs/openenv/pr_1028/en/_app/immutable/chunks/BlrbfE70.js" rel="modulepreload">
<link href="/docs/openenv/pr_1028/en/_app/immutable/chunks/DsnmJJEf.js" rel="modulepreload">
<link href="/docs/openenv/pr_1028/en/_app/immutable/chunks/BAUKCucN.js" rel="modulepreload">
<link href="/docs/openenv/pr_1028/en/_app/immutable/nodes/0.DQ3mQjnq.js" rel="modulepreload">
<link href="/docs/openenv/pr_1028/en/_app/immutable/chunks/CR5HsSXT.js" rel="modulepreload">
<link href="/docs/openenv/pr_1028/en/_app/immutable/nodes/65.CzwY00bL.js" rel="modulepreload">
<link href="/docs/openenv/pr_1028/en/_app/immutable/chunks/C07d6Hje.js" rel="modulepreload">
<!--ekqdbr--><meta name="hf:doc:metadata" content="{&quot;title&quot;:&quot;Coding Agent Training with TRL (OpenCode)&quot;,&quot;local&quot;:&quot;coding-agent-training-with-trl-opencode&quot;,&quot;sections&quot;:[{&quot;title&quot;:&quot;How It Works&quot;,&quot;local&quot;:&quot;how-it-works&quot;,&quot;sections&quot;:[],&quot;depth&quot;:2},{&quot;title&quot;:&quot;Full Recipe&quot;,&quot;local&quot;:&quot;full-recipe&quot;,&quot;sections&quot;:[],&quot;depth&quot;:2}],&quot;depth&quot;:1}"/><!---->
<link href="/docs/openenv/pr_1028/en/_app/immutable/assets/0.tn0RQdqM.css" rel="modulepreload"> <!--[--><!--[0--><!--[--><!--[0--><!--[--><p></p> <div class="items-center shrink-0 min-w-[100px] max-sm:min-w-[50px] justify-end ml-auto flex" style="float: right; margin-left: 10px; display: inline-flex; position: relative; z-index: 10;"><div class="inline-flex rounded-md max-sm:rounded-sm"><button class="inline-flex items-center gap-1 h-7 max-sm:h-7 px-2 max-sm:px-1.5 text-sm font-medium text-gray-800 border border-r-0 rounded-l-md max-sm:rounded-l-sm border-gray-200 bg-white hover:shadow-inner dark:border-gray-850 dark:bg-gray-950 dark:text-gray-200 dark:hover:bg-gray-800" aria-live="polite"><span class="inline-flex items-center justify-center rounded-md p-0.5 max-sm:p-0 hover:text-gray-800 dark:hover:text-gray-200"><svg class="sm:size-3.5 size-3" xmlns="http://www.w3.org/2000/svg" aria-hidden="true" fill="currentColor" focusable="false" role="img" width="1em" height="1em" preserveAspectRatio="xMidYMid meet" viewBox="0 0 32 32"><path d="M28,10V28H10V10H28m0-2H10a2,2,0,0,0-2,2V28a2,2,0,0,0,2,2H28a2,2,0,0,0,2-2V10a2,2,0,0,0-2-2Z" transform="translate(0)"></path><path d="M4,18H2V4A2,2,0,0,1,4,2H18V4H4Z" transform="translate(0)"></path><rect fill="none" width="32" height="32"></rect></svg><!----></span> <span>Copy page</span></button> <button class="inline-flex items-center justify-center w-6 max-sm:w-5 h-7 max-sm:h-7 disabled:pointer-events-none text-sm text-gray-500 hover:text-gray-700 dark:hover:text-white rounded-r-md max-sm:rounded-r-sm border border-l transition border-gray-200 bg-white hover:shadow-inner dark:border-gray-850 dark:bg-gray-950 dark:text-gray-200 dark:hover:bg-gray-800" aria-haspopup="menu" aria-expanded="false" aria-label="Open copy menu"><svg class="transition-transform text-gray-400 overflow-visible sm:size-3.5 size-3 rotate-0" width="1em" height="1em" viewBox="0 0 12 7" fill="none" xmlns="http://www.w3.org/2000/svg"><path d="M1 1L6 6L11 1" stroke="currentColor"></path></svg><!----></button></div> <!--[-1--><!--]--></div><!----> <!--[0--><h1 class="relative group"><a id="coding-agent-training-with-trl-opencode" class="header-link block pr-1.5 text-lg no-hover:hidden with-hover:absolute with-hover:p-1.5 with-hover:opacity-0 with-hover:group-hover:opacity-100 with-hover:right-full" href="#coding-agent-training-with-trl-opencode"><span><svg xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" aria-hidden="true" role="img" width="1em" height="1em" preserveAspectRatio="xMidYMid meet" viewBox="0 0 256 256"><path d="M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z" fill="currentColor"></path></svg><!----></span></a> <span>Coding Agent Training with TRL (OpenCode)</span></h1><!--]--><!----> <p>This tutorial covers the black-box training path: training the actual <a href="https://opencode.ai" rel="nofollow"><code>opencode</code></a> coding agent, with its own planner, tools,
context management, and stop condition, using TRL’s experimental <code>AsyncGRPOTrainer</code>. The agent owns its loop, and OpenEnv captures what it did.</p> <blockquote class="note"><p>Three GRPO patterns, three tutorials. For a standard <code>reset()</code> / <code>step()</code> flow where TRL drives the episode, see the <a href="wordle-grpo">Wordle GRPO tutorial</a>. For harness rollouts where the
trainer still generates each turn (white-box), see the <a href="browsergym-harness">BrowserGym harness tutorial</a>. Use this page when you
want to train a production agent as-is, without reimplementing its loop.</p></blockquote> <!--[1--><h2 class="relative group"><a id="how-it-works" class="header-link block pr-1.5 text-lg no-hover:hidden with-hover:absolute with-hover:p-1.5 with-hover:opacity-0 with-hover:group-hover:opacity-100 with-hover:right-full" href="#how-it-works"><span><svg xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" aria-hidden="true" role="img" width="1em" height="1em" preserveAspectRatio="xMidYMid meet" viewBox="0 0 256 256"><path d="M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z" fill="currentColor"></path></svg><!----></span></a> <span>How It Works</span></h2><!--]--><!----> <p>The full recipe lives in TRL. The moving pieces:</p> <ol><li>Each rollout runs the agent inside an OpenEnv session created by <code>OpenCodeSessionFactory</code> from <a href="https://github.com/huggingface/OpenEnv/tree/main/envs/opencode_env" rel="nofollow"><code>opencode_env</code></a>,
in <code>transparent_proxy</code> mode. A small proxy inside the sandbox forwards the
agent’s <code>/v1/chat/completions</code> calls to your vLLM server and records each
turn’s token ids and logprobs to a trace.</li> <li>When the agent stops, TRL’s <code>HarnessRolloutWorker</code> reads the trace, rebuilds
the per-turn training rows from the recorded ids, and scores the final
workspace with the session’s <code>verify()</code> method (a held-out verifier the
agent never sees).</li> <li><code>AsyncGRPOTrainer</code> trains on those rows, propagating the rollout reward to
every trained token through the group-relative advantage. NCCL weight sync
keeps the vLLM server on the current policy, so the agent always samples
from the model being trained.</li></ol> <p>Each rollout gets its own isolated session: one sandbox, one proxy port, one
agent process. Three small functions adapt the recipe to your task: <code>rollout_reward_fn</code> (outcome to scalar reward), <code>train_turn_fn</code> (which turns
receive gradient), and <code>agent_turn_fn</code> (which trace entries are real agent
turns rather than auxiliary calls like title generation). All three are
documented in <a href="https://huggingface.co/docs/trl/openenv#training-on-harnesses-training-a-real-coding-agent-opencode" rel="nofollow">TRL’s harness training guide</a>.</p> <!--[1--><h2 class="relative group"><a id="full-recipe" class="header-link block pr-1.5 text-lg no-hover:hidden with-hover:absolute with-hover:p-1.5 with-hover:opacity-0 with-hover:group-hover:opacity-100 with-hover:right-full" href="#full-recipe"><span><svg xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" aria-hidden="true" role="img" width="1em" height="1em" preserveAspectRatio="xMidYMid meet" viewBox="0 0 256 256"><path d="M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z" fill="currentColor"></path></svg><!----></span></a> <span>Full Recipe</span></h2><!--]--><!----> <p>The reference script trains on competitive-coding problems from <code>agentica-org/DeepCoder-Preview-Dataset</code>: the agent writes <code>solution.py</code>, and
the verifier runs it against held-out tests, returning the fraction passed. It
is self-contained, runs the agent in a local subprocess sandbox (no container
setup needed), needs two GPUs (one serving the policy with vLLM, one
training), and has been validated end to end on Qwen3. To scale
rollouts beyond a single node, a sibling script runs each rollout in its own
remote Hugging Face sandbox instead of a local subprocess.
Installation, the exact vLLM serving flags, and the run commands live next to
the recipe in TRL:</p> <ul><li><a href="https://huggingface.co/docs/trl/openenv#training-on-harnesses-training-a-real-coding-agent-opencode" rel="nofollow">Training on harnesses</a> in TRL’s OpenEnv docs: rollout semantics, the reward path, turn selection,
and the trace contract.</li> <li><a href="https://github.com/huggingface/trl/blob/main/examples/scripts/openenv/opencode.py" rel="nofollow"><code>examples/scripts/openenv/opencode.py</code></a> in TRL: the complete, runnable script.</li> <li><a href="https://github.com/huggingface/trl/blob/main/examples/scripts/openenv/opencode_hf_sandbox.py" rel="nofollow"><code>examples/scripts/openenv/opencode_hf_sandbox.py</code></a> in TRL: the same recipe, but each rollout runs in its own remote Hugging Face
sandbox, so rollouts scale out beyond one node.</li> <li><a href="https://github.com/huggingface/OpenEnv/tree/main/envs/opencode_env" rel="nofollow"><code>envs/opencode_env</code></a>:
the OpenEnv side, including the session factory, sandbox backends, and the
transparent interception proxy.</li></ul> <a class="!text-gray-400 !no-underline text-sm flex items-center not-prose mt-4" href="https://github.com/huggingface/openenv/blob/main/docs/source/tutorials/opencode-agent-grpo.md" target="_blank"><svg class="mr-1" xmlns="http://www.w3.org/2000/svg" aria-hidden="true" fill="currentColor" focusable="false" role="img" width="1em" height="1em" preserveAspectRatio="xMidYMid meet" viewBox="0 0 32 32"><path d="M31,16l-7,7l-1.41-1.41L28.17,16l-5.58-5.59L24,9l7,7z"></path><path d="M1,16l7-7l1.41,1.41L3.83,16l5.58,5.59L8,23l-7-7z"></path><path d="M12.419,25.484L17.639,6.552l1.932,0.518L14.351,26.002z"></path></svg><!----> <span><span class="underline">Update</span> on GitHub</span></a><!----> <p></p><!--]--><!----><!--]--><!--]--><!--]--> <!--[-1--><!--]--><!--]-->
<script>
{
__sveltekit_1yyva3b = {
base: "/docs/openenv/pr_1028/en",
assets: "/docs/openenv/pr_1028/en"
};
const element = document.currentScript.parentElement;
Promise.all([
import("/docs/openenv/pr_1028/en/_app/immutable/entry/start.Bs7kuoo8.js"),
import("/docs/openenv/pr_1028/en/_app/immutable/entry/app.D4PcSJ6S.js")
]).then(([kit, app]) => {
kit.start(app, element, {
node_ids: [0, 65],
data: [null,null],
form: null,
error: null
});
});
}
</script>

Xet Storage Details

Size:
13.1 kB
·
Xet hash:
db9716c00fa47b42b29bf3c9411f220dc2e03fc198d98e455071c328af9558b9

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.