Spaces:
Paused
HCM:21 Poster β Runbook for Human-Required Work
Everything that can be done without a GPU, network credentials, or a personal
account is already committed on branch reward-recalibration:
- β
Reward recalibration (
scoring.pyv2) β reward hacking fixed, 43/43 tests pass - β
Real baseline experiment + hero figure (
poster_experiment.py,results.json,poster_figure.png) - β
Figures 1, 2, 4 (
poster_diagrams.pyβfig1_*.png,fig2_*.png,fig4_*.png) - β
Hermes MCP tool server (
integrations/hcm21_mcp_server.py) β verified in-process - β
Poster draft (
POSTER_DRAFT.md)
The remaining items below need a human + GPU/accounts. Do them in order.
0. Prerequisites (5 min)
- A Hugging Face account with write access to
ParetoOptimal/hcm21(the Space). - A Google account for Colab (free T4 is enough).
ANTHROPIC_API_KEY(or any provider key) if you run the Claude/Hermes agent.- Merge or check out
reward-recalibration:gh pr create --base main --head reward-recalibration \ --title "Reward recalibration + poster assets" --fill # then review & merge
1. Redeploy the HF Space with scoring v2 β οΈ DO THIS FIRST
Why it's critical: the Colab notebook trains and evaluates against the live
Space (https://paretooptimal-hcm21.hf.space). Until you redeploy, the live env
still runs v1 scoring and GRPO will relearn the downsizing exploit. The fix
must be live before any training run.
# from the repo root, with scoring.py v2 in the tree:
git remote add space https://huggingface.co/spaces/ParetoOptimal/hcm21 # once
git push space reward-recalibration:main
Wait for the Space to rebuild (Docker), then sanity-check the floor is fixed:
# do-nothing should now score LOW (~0.25), not ~0.66
curl "https://paretooptimal-hcm21.hf.space/demo?seed=42&size=300" | python -m json.tool | grep final_score
Locally you can confirm the same separation any time:
pip install -e ".[server,poster]"
python poster_experiment.py # deterministic: null 0.247, random 0.262, heuristic 0.448
python -m pytest tests/ -q # 43 passing
2. Run GRPO training in Colab β produce Figure 5 (1β2 h on T4)
- Open
notebooks/hcm21_trl_training.ipynbin Colab (Runtime β T4 GPU). - Run cells top-to-bottom. The notebook already:
- loads
unsloth/Qwen3-0.6B+ LoRA (r=16, 4-bit), - runs the random and heuristic baselines against the live Space,
- configures
GRPOConfig(batch 2 Γ grad-accum 4, 4 generations, lr 5e-6, fp16), - trains, then plots the reward curve and a baseline-vs-trained comparison.
- loads
- Watch for these known gotchas:
- Space must be v2 (step 1) or the reward signal is hacked.
- If
from openenv.core import GenericEnvClienterrors, pinopenenv-coreto the version the Space was built with (pip install openenv-core==<ver>). - 0.6B cold-start: if completions rarely contain valid JSON actions and reward
stays ~0, either (a) do the Hermes SFT warm-start in step 4, or (b) raise
num_generationsand lowerlr.
- Export results for the poster β add this cell at the end of the notebook. The
schema is exactly what
poster_experiment.pyauto-detects:import json json.dump({ "trained": trained_scores, # per-seed final scores of the GRPO agent "reward_curve": [e["reward"] for e in trainer.state.log_history if "reward" in e], }, open("grpo_results.json", "w"), indent=2) - Download
grpo_results.jsonfrom Colab.
Acceptance: trained score > heuristic (β0.45), ideally β₯0.55, with a reward curve that trends up. If trained β heuristic, that's still a publishable result β report it honestly (small model, limited steps) rather than inflating.
3. Add the trained-agent bar to Figure 3 (1 min, local β paste-and-run)
poster_experiment.py already auto-detects the Colab output. Just drop
grpo_results.json (from step 2.5) into the repo root and re-run:
cp /path/to/grpo_results.json .
python poster_experiment.py
This regenerates poster_figure.png with a green Trained (GRPO) bar plus a
dashed "heuristic bar" reference line, and β if grpo_results.json includes a
reward_curve β also writes fig5_grpo_reward_curve.png (Figure 5). With no
file present the script is unchanged (baseline-only). No code edits needed.
4. (Optional, high-value) Hermes-Agent warm-start pipeline
This is the "agent that grows with you" enhancement. It both fixes the 0.6B cold-start and gives the poster a second experimental axis.
4a. Install Hermes and verify its real APIs
git clone https://github.com/nousresearch/hermes-agent && cd hermes-agent
# follow its README installer; then confirm the exact names of:
# - the batch trajectory generation script (README calls it "batch trajectory generation")
# - the trajectory compression / training-export utility
# - the skills directory + agentskills.io skill format
The public README does not print these script names. Confirm them in the cloned repo before scripting against them; update step 4c accordingly.
4b. Register the HCM:21 MCP server with Hermes
The server is built and tested (integrations/hcm21_mcp_server.py). Launch it:
pip install -e ".[server,mcp]"
python -m integrations.hcm21_mcp_server # stdio MCP transport
Add it to Hermes' MCP config (Hermes "Connect any MCP server"). Hermes can now
call hcm21_reset β hcm21_available_actions β hcm21_step ΓN β hcm21_score.
Author a Hermes skill that runs the scanβplanβproduceβcontrol loop and reads
the final score β this is the procedural-memory artifact the poster describes.
4c. Generate expert trajectories β SFT warm-start β GRPO
1. Run Hermes (strong model) over the 4 scenarios Γ several seeds via the MCP
server β use Hermes' batch trajectory generation to collect episodes.
2. Use Hermes' trajectory compression to export an SFT dataset of
(prompt, action-sequence) pairs from the high-scoring episodes.
3. SFT unsloth/Qwen3-0.6B on that dataset (1β2 epochs), THEN run the GRPO
notebook from that checkpoint instead of the base model.
Acceptance: SFT-then-GRPO reaches a higher score / converges faster than
GRPO-from-base. Report the three-regime table from POSTER_DRAFT.md Β§6
(GRPO-only vs Hermes-skills-only vs hybrid).
5. Assemble and submit the poster (by July 26, 11:59 PM PDT)
- Layout (A0 portrait or the conference's stated size β confirm on Sessionize):
- Top: title + Β§1 motivation.
- Left column: Β§2 environment + Fig 1 (architecture), Fig 2 (phase loop).
- Center (hero): Β§3 finding + Β§4 fix + Fig 3 (recalibration) β biggest panel.
- Right column: Β§5 training + Fig 4 (GRPO loop) + Fig 5 (reward curve), then Β§6 Hermes + the three-regime table.
- Footer: Β§7 "Try it" QR codes (Space, GitHub, Colab) + MIT license.
- Build the PDF in your tool of choice (PowerPoint/Keynote/Figma/LaTeX
tikzposterorbeamerposter). Embed the committed PNGs at 150+ dpi. - Submit on Sessionize (link from the CFP page). Max 2 speakers. Paste the
abstract from
POSTER_DRAFT.md(Title + Β§1). Upload the PDF if requested. - Diversity rule: if 3+ presenters, gender diversity is required β keep to β€2 or compose accordingly.
- Print & ship: presenters print and bring their own poster (CFP requirement). Order the print ~2 weeks before Oct 20 (San Jose).
CFP dates
- Submit: Sun July 26, 11:59 PM PDT
- Notifications: Mon Aug 17 Β· Schedule: Tue Aug 18 Β· Event: Oct 20β21, San Jose
- Accepted presenters get complimentary passes; slides due before the event.
6. Final verification checklist
- HF Space redeployed; live do-nothing score β 0.25 (not 0.66)
-
python -m pytest tests/ -qβ 43 passing -
python poster_experiment.pyβ null 0.247, random 0.262, heuristic 0.448 (deterministic) -
python poster_diagrams.pyβ Figs 1, 2, 4 regenerate - Colab GRPO run complete;
grpo_results.json+ reward curve saved (Fig 5) - Trained-agent bar added to Fig 3
- (Optional) Hermes warm-start results + three-regime table
- Poster PDF assembled with all 5 figures + QR codes
- Submitted on Sessionize before July 26
- Caveat resolved: Hermes script names verified against the actual repo