Spaces:
Paused
Paused
File size: 8,664 Bytes
670ccf0 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 | # HCM:21 Poster — Runbook for Human-Required Work
Everything that can be done without a GPU, network credentials, or a personal
account is already committed on branch `reward-recalibration`:
- ✅ Reward recalibration (`scoring.py` v2) — reward hacking fixed, 43/43 tests pass
- ✅ Real baseline experiment + hero figure (`poster_experiment.py`, `results.json`, `poster_figure.png`)
- ✅ Figures 1, 2, 4 (`poster_diagrams.py` → `fig1_*.png`, `fig2_*.png`, `fig4_*.png`)
- ✅ Hermes MCP tool server (`integrations/hcm21_mcp_server.py`) — verified in-process
- ✅ Poster draft (`POSTER_DRAFT.md`)
The remaining items below need **a human + GPU/accounts**. Do them in order.
---
## 0. Prerequisites (5 min)
- A Hugging Face account with write access to `ParetoOptimal/hcm21` (the Space).
- A Google account for Colab (free T4 is enough).
- `ANTHROPIC_API_KEY` (or any provider key) if you run the Claude/Hermes agent.
- Merge or check out `reward-recalibration`:
```bash
gh pr create --base main --head reward-recalibration \
--title "Reward recalibration + poster assets" --fill # then review & merge
```
---
## 1. Redeploy the HF Space with scoring v2 ⚠️ DO THIS FIRST
**Why it's critical:** the Colab notebook trains and evaluates against the *live*
Space (`https://paretooptimal-hcm21.hf.space`). Until you redeploy, the live env
still runs v1 scoring and **GRPO will relearn the downsizing exploit**. The fix
must be live before any training run.
```bash
# from the repo root, with scoring.py v2 in the tree:
git remote add space https://huggingface.co/spaces/ParetoOptimal/hcm21 # once
git push space reward-recalibration:main
```
Wait for the Space to rebuild (Docker), then sanity-check the floor is fixed:
```bash
# do-nothing should now score LOW (~0.25), not ~0.66
curl "https://paretooptimal-hcm21.hf.space/demo?seed=42&size=300" | python -m json.tool | grep final_score
```
Locally you can confirm the same separation any time:
```bash
pip install -e ".[server,poster]"
python poster_experiment.py # deterministic: null 0.247, random 0.262, heuristic 0.448
python -m pytest tests/ -q # 43 passing
```
---
## 2. Run GRPO training in Colab → produce Figure 5 (1–2 h on T4)
1. Open `notebooks/hcm21_trl_training.ipynb` in Colab (Runtime → T4 GPU).
2. Run cells top-to-bottom. The notebook already:
- loads `unsloth/Qwen3-0.6B` + LoRA (r=16, 4-bit),
- runs the **random** and **heuristic** baselines against the live Space,
- configures `GRPOConfig` (batch 2 × grad-accum 4, 4 generations, lr 5e-6, fp16),
- trains, then plots the reward curve and a baseline-vs-trained comparison.
3. Watch for these known gotchas:
- **Space must be v2** (step 1) or the reward signal is hacked.
- If `from openenv.core import GenericEnvClient` errors, pin `openenv-core`
to the version the Space was built with (`pip install openenv-core==<ver>`).
- 0.6B cold-start: if completions rarely contain valid JSON actions and reward
stays ~0, either (a) do the Hermes SFT warm-start in step 4, or (b) raise
`num_generations` and lower `lr`.
4. Export results for the poster — add this cell at the end of the notebook. The
schema is exactly what `poster_experiment.py` auto-detects:
```python
import json
json.dump({
"trained": trained_scores, # per-seed final scores of the GRPO agent
"reward_curve": [e["reward"] for e in trainer.state.log_history if "reward" in e],
}, open("grpo_results.json", "w"), indent=2)
```
5. Download `grpo_results.json` from Colab.
**Acceptance:** trained score > heuristic (≈0.45), ideally ≥0.55, with a reward
curve that trends up. If trained ≈ heuristic, that's still a publishable result —
report it honestly (small model, limited steps) rather than inflating.
---
## 3. Add the trained-agent bar to Figure 3 (1 min, local — paste-and-run)
`poster_experiment.py` already auto-detects the Colab output. Just drop
`grpo_results.json` (from step 2.5) into the repo root and re-run:
```bash
cp /path/to/grpo_results.json .
python poster_experiment.py
```
This regenerates `poster_figure.png` with a green **Trained (GRPO)** bar plus a
dashed "heuristic bar" reference line, and — if `grpo_results.json` includes a
`reward_curve` — also writes **`fig5_grpo_reward_curve.png`** (Figure 5). With no
file present the script is unchanged (baseline-only). No code edits needed.
---
## 4. (Optional, high-value) Hermes-Agent warm-start pipeline
This is the "agent that grows with you" enhancement. It both fixes the 0.6B
cold-start and gives the poster a second experimental axis.
### 4a. Install Hermes and verify its real APIs
```bash
git clone https://github.com/nousresearch/hermes-agent && cd hermes-agent
# follow its README installer; then confirm the exact names of:
# - the batch trajectory generation script (README calls it "batch trajectory generation")
# - the trajectory compression / training-export utility
# - the skills directory + agentskills.io skill format
```
> The public README does **not** print these script names. Confirm them in the
> cloned repo before scripting against them; update step 4c accordingly.
### 4b. Register the HCM:21 MCP server with Hermes
The server is built and tested (`integrations/hcm21_mcp_server.py`). Launch it:
```bash
pip install -e ".[server,mcp]"
python -m integrations.hcm21_mcp_server # stdio MCP transport
```
Add it to Hermes' MCP config (Hermes "Connect any MCP server"). Hermes can now
call `hcm21_reset → hcm21_available_actions → hcm21_step ×N → hcm21_score`.
Author a Hermes **skill** that runs the scan→plan→produce→control loop and reads
the final score — this is the procedural-memory artifact the poster describes.
### 4c. Generate expert trajectories → SFT warm-start → GRPO
```text
1. Run Hermes (strong model) over the 4 scenarios × several seeds via the MCP
server → use Hermes' batch trajectory generation to collect episodes.
2. Use Hermes' trajectory compression to export an SFT dataset of
(prompt, action-sequence) pairs from the high-scoring episodes.
3. SFT unsloth/Qwen3-0.6B on that dataset (1–2 epochs), THEN run the GRPO
notebook from that checkpoint instead of the base model.
```
**Acceptance:** SFT-then-GRPO reaches a higher score / converges faster than
GRPO-from-base. Report the three-regime table from `POSTER_DRAFT.md §6`
(GRPO-only vs Hermes-skills-only vs hybrid).
---
## 5. Assemble and submit the poster (by July 26, 11:59 PM PDT)
1. **Layout** (A0 portrait or the conference's stated size — confirm on Sessionize):
- Top: title + §1 motivation.
- Left column: §2 environment + Fig 1 (architecture), Fig 2 (phase loop).
- Center (hero): §3 finding + §4 fix + **Fig 3** (recalibration) — biggest panel.
- Right column: §5 training + Fig 4 (GRPO loop) + Fig 5 (reward curve),
then §6 Hermes + the three-regime table.
- Footer: §7 "Try it" QR codes (Space, GitHub, Colab) + MIT license.
2. **Build the PDF** in your tool of choice (PowerPoint/Keynote/Figma/LaTeX
`tikzposter` or `beamerposter`). Embed the committed PNGs at 150+ dpi.
3. **Submit on Sessionize** (link from the CFP page). Max 2 speakers. Paste the
abstract from `POSTER_DRAFT.md` (Title + §1). Upload the PDF if requested.
4. **Diversity rule:** if 3+ presenters, gender diversity is required — keep to ≤2
or compose accordingly.
5. **Print & ship:** presenters print and bring their own poster (CFP requirement).
Order the print ~2 weeks before Oct 20 (San Jose).
### CFP dates
- Submit: **Sun July 26, 11:59 PM PDT**
- Notifications: Mon Aug 17 · Schedule: Tue Aug 18 · Event: Oct 20–21, San Jose
- Accepted presenters get complimentary passes; slides due before the event.
---
## 6. Final verification checklist
- [ ] HF Space redeployed; live do-nothing score ≈ 0.25 (not 0.66)
- [ ] `python -m pytest tests/ -q` → 43 passing
- [ ] `python poster_experiment.py` → null 0.247, random 0.262, heuristic 0.448 (deterministic)
- [ ] `python poster_diagrams.py` → Figs 1, 2, 4 regenerate
- [ ] Colab GRPO run complete; `grpo_results.json` + reward curve saved (Fig 5)
- [ ] Trained-agent bar added to Fig 3
- [ ] (Optional) Hermes warm-start results + three-regime table
- [ ] Poster PDF assembled with all 5 figures + QR codes
- [ ] Submitted on Sessionize before July 26
- [ ] Caveat resolved: Hermes script names verified against the actual repo
|