File size: 8,664 Bytes
670ccf0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
# HCM:21 Poster — Runbook for Human-Required Work

Everything that can be done without a GPU, network credentials, or a personal
account is already committed on branch `reward-recalibration`:

- ✅ Reward recalibration (`scoring.py` v2) — reward hacking fixed, 43/43 tests pass
- ✅ Real baseline experiment + hero figure (`poster_experiment.py`, `results.json`, `poster_figure.png`)
- ✅ Figures 1, 2, 4 (`poster_diagrams.py``fig1_*.png`, `fig2_*.png`, `fig4_*.png`)
- ✅ Hermes MCP tool server (`integrations/hcm21_mcp_server.py`) — verified in-process
- ✅ Poster draft (`POSTER_DRAFT.md`)

The remaining items below need **a human + GPU/accounts**. Do them in order.

---

## 0. Prerequisites (5 min)

- A Hugging Face account with write access to `ParetoOptimal/hcm21` (the Space).
- A Google account for Colab (free T4 is enough).
- `ANTHROPIC_API_KEY` (or any provider key) if you run the Claude/Hermes agent.
- Merge or check out `reward-recalibration`:
  ```bash

  gh pr create --base main --head reward-recalibration \

    --title "Reward recalibration + poster assets" --fill   # then review & merge

  ```

---

## 1. Redeploy the HF Space with scoring v2  ⚠️ DO THIS FIRST

**Why it's critical:** the Colab notebook trains and evaluates against the *live*
Space (`https://paretooptimal-hcm21.hf.space`). Until you redeploy, the live env
still runs v1 scoring and **GRPO will relearn the downsizing exploit**. The fix
must be live before any training run.

```bash

# from the repo root, with scoring.py v2 in the tree:

git remote add space https://huggingface.co/spaces/ParetoOptimal/hcm21   # once

git push space reward-recalibration:main

```
Wait for the Space to rebuild (Docker), then sanity-check the floor is fixed:
```bash

# do-nothing should now score LOW (~0.25), not ~0.66

curl "https://paretooptimal-hcm21.hf.space/demo?seed=42&size=300" | python -m json.tool | grep final_score

```
Locally you can confirm the same separation any time:
```bash

pip install -e ".[server,poster]"

python poster_experiment.py     # deterministic: null 0.247, random 0.262, heuristic 0.448

python -m pytest tests/ -q      # 43 passing

```

---

## 2. Run GRPO training in Colab → produce Figure 5 (1–2 h on T4)

1. Open `notebooks/hcm21_trl_training.ipynb` in Colab (Runtime → T4 GPU).
2. Run cells top-to-bottom. The notebook already:
   - loads `unsloth/Qwen3-0.6B` + LoRA (r=16, 4-bit),
   - runs the **random** and **heuristic** baselines against the live Space,
   - configures `GRPOConfig` (batch 2 × grad-accum 4, 4 generations, lr 5e-6, fp16),
   - trains, then plots the reward curve and a baseline-vs-trained comparison.
3. Watch for these known gotchas:
   - **Space must be v2** (step 1) or the reward signal is hacked.
   - If `from openenv.core import GenericEnvClient` errors, pin `openenv-core`
     to the version the Space was built with (`pip install openenv-core==<ver>`).

   - 0.6B cold-start: if completions rarely contain valid JSON actions and reward

     stays ~0, either (a) do the Hermes SFT warm-start in step 4, or (b) raise

     `num_generations` and lower `lr`.

4. Export results for the poster — add this cell at the end of the notebook. The

   schema is exactly what `poster_experiment.py` auto-detects:

   ```python

   import json

   json.dump({

       "trained": trained_scores,   # per-seed final scores of the GRPO agent

       "reward_curve": [e["reward"] for e in trainer.state.log_history if "reward" in e],

   }, open("grpo_results.json", "w"), indent=2)

   ```

5. Download `grpo_results.json` from Colab.


**Acceptance:** trained score > heuristic (≈0.45), ideally ≥0.55, with a reward
curve that trends up. If trained ≈ heuristic, that's still a publishable result —
report it honestly (small model, limited steps) rather than inflating.

---

## 3. Add the trained-agent bar to Figure 3 (1 min, local — paste-and-run)

`poster_experiment.py` already auto-detects the Colab output. Just drop
`grpo_results.json` (from step 2.5) into the repo root and re-run:

```bash

cp /path/to/grpo_results.json .

python poster_experiment.py

```
This regenerates `poster_figure.png` with a green **Trained (GRPO)** bar plus a
dashed "heuristic bar" reference line, and — if `grpo_results.json` includes a
`reward_curve` — also writes **`fig5_grpo_reward_curve.png`** (Figure 5). With no

file present the script is unchanged (baseline-only). No code edits needed.



---



## 4. (Optional, high-value) Hermes-Agent warm-start pipeline



This is the "agent that grows with you" enhancement. It both fixes the 0.6B

cold-start and gives the poster a second experimental axis.



### 4a. Install Hermes and verify its real APIs

```bash

git clone https://github.com/nousresearch/hermes-agent && cd hermes-agent

# follow its README installer; then confirm the exact names of:

#   - the batch trajectory generation script (README calls it "batch trajectory generation")

#   - the trajectory compression / training-export utility

#   - the skills directory + agentskills.io skill format

```

> The public README does **not** print these script names. Confirm them in the

> cloned repo before scripting against them; update step 4c accordingly.



### 4b. Register the HCM:21 MCP server with Hermes

The server is built and tested (`integrations/hcm21_mcp_server.py`). Launch it:

```bash

pip install -e ".[server,mcp]"

python -m integrations.hcm21_mcp_server     # stdio MCP transport

```

Add it to Hermes' MCP config (Hermes "Connect any MCP server"). Hermes can now

call `hcm21_reset → hcm21_available_actions → hcm21_step ×N → hcm21_score`.

Author a Hermes **skill** that runs the scan→plan→produce→control loop and reads
the final score — this is the procedural-memory artifact the poster describes.

### 4c. Generate expert trajectories → SFT warm-start → GRPO
```text

1. Run Hermes (strong model) over the 4 scenarios × several seeds via the MCP

   server → use Hermes' batch trajectory generation to collect episodes.

2. Use Hermes' trajectory compression to export an SFT dataset of

   (prompt, action-sequence) pairs from the high-scoring episodes.

3. SFT unsloth/Qwen3-0.6B on that dataset (1–2 epochs), THEN run the GRPO

   notebook from that checkpoint instead of the base model.

```
**Acceptance:** SFT-then-GRPO reaches a higher score / converges faster than
GRPO-from-base. Report the three-regime table from `POSTER_DRAFT.md §6`
(GRPO-only vs Hermes-skills-only vs hybrid).

---

## 5. Assemble and submit the poster (by July 26, 11:59 PM PDT)

1. **Layout** (A0 portrait or the conference's stated size — confirm on Sessionize):
   - Top: title + §1 motivation.
   - Left column: §2 environment + Fig 1 (architecture), Fig 2 (phase loop).
   - Center (hero): §3 finding + §4 fix + **Fig 3** (recalibration) — biggest panel.
   - Right column: §5 training + Fig 4 (GRPO loop) + Fig 5 (reward curve),
     then §6 Hermes + the three-regime table.

   - Footer: §7 "Try it" QR codes (Space, GitHub, Colab) + MIT license.

2. **Build the PDF** in your tool of choice (PowerPoint/Keynote/Figma/LaTeX

   `tikzposter` or `beamerposter`). Embed the committed PNGs at 150+ dpi.

3. **Submit on Sessionize** (link from the CFP page). Max 2 speakers. Paste the

   abstract from `POSTER_DRAFT.md` (Title + §1). Upload the PDF if requested.

4. **Diversity rule:** if 3+ presenters, gender diversity is required — keep to ≤2

   or compose accordingly.

5. **Print & ship:** presenters print and bring their own poster (CFP requirement).

   Order the print ~2 weeks before Oct 20 (San Jose).


### CFP dates
- Submit: **Sun July 26, 11:59 PM PDT**
- Notifications: Mon Aug 17 · Schedule: Tue Aug 18 · Event: Oct 20–21, San Jose
- Accepted presenters get complimentary passes; slides due before the event.

---

## 6. Final verification checklist

- [ ] HF Space redeployed; live do-nothing score ≈ 0.25 (not 0.66)
- [ ] `python -m pytest tests/ -q` → 43 passing
- [ ] `python poster_experiment.py` → null 0.247, random 0.262, heuristic 0.448 (deterministic)
- [ ] `python poster_diagrams.py` → Figs 1, 2, 4 regenerate
- [ ] Colab GRPO run complete; `grpo_results.json` + reward curve saved (Fig 5)
- [ ] Trained-agent bar added to Fig 3
- [ ] (Optional) Hermes warm-start results + three-regime table
- [ ] Poster PDF assembled with all 5 figures + QR codes
- [ ] Submitted on Sessionize before July 26
- [ ] Caveat resolved: Hermes script names verified against the actual repo