hcm21 / RUNBOOK.md
ParetoOptimal's picture
Sync Space with HCM-21-Private main (72f5f30): reward recalibration, seeded RNG streams, MCP integration, poster artifacts
670ccf0 verified
|
Raw
History Blame Contribute Delete
8.66 kB

HCM:21 Poster β€” Runbook for Human-Required Work

Everything that can be done without a GPU, network credentials, or a personal account is already committed on branch reward-recalibration:

  • βœ… Reward recalibration (scoring.py v2) β€” reward hacking fixed, 43/43 tests pass
  • βœ… Real baseline experiment + hero figure (poster_experiment.py, results.json, poster_figure.png)
  • βœ… Figures 1, 2, 4 (poster_diagrams.py β†’ fig1_*.png, fig2_*.png, fig4_*.png)
  • βœ… Hermes MCP tool server (integrations/hcm21_mcp_server.py) β€” verified in-process
  • βœ… Poster draft (POSTER_DRAFT.md)

The remaining items below need a human + GPU/accounts. Do them in order.


0. Prerequisites (5 min)

  • A Hugging Face account with write access to ParetoOptimal/hcm21 (the Space).
  • A Google account for Colab (free T4 is enough).
  • ANTHROPIC_API_KEY (or any provider key) if you run the Claude/Hermes agent.
  • Merge or check out reward-recalibration:
    gh pr create --base main --head reward-recalibration \
      --title "Reward recalibration + poster assets" --fill   # then review & merge
    

1. Redeploy the HF Space with scoring v2 ⚠️ DO THIS FIRST

Why it's critical: the Colab notebook trains and evaluates against the live Space (https://paretooptimal-hcm21.hf.space). Until you redeploy, the live env still runs v1 scoring and GRPO will relearn the downsizing exploit. The fix must be live before any training run.

# from the repo root, with scoring.py v2 in the tree:
git remote add space https://huggingface.co/spaces/ParetoOptimal/hcm21   # once
git push space reward-recalibration:main

Wait for the Space to rebuild (Docker), then sanity-check the floor is fixed:

# do-nothing should now score LOW (~0.25), not ~0.66
curl "https://paretooptimal-hcm21.hf.space/demo?seed=42&size=300" | python -m json.tool | grep final_score

Locally you can confirm the same separation any time:

pip install -e ".[server,poster]"
python poster_experiment.py     # deterministic: null 0.247, random 0.262, heuristic 0.448
python -m pytest tests/ -q      # 43 passing

2. Run GRPO training in Colab β†’ produce Figure 5 (1–2 h on T4)

  1. Open notebooks/hcm21_trl_training.ipynb in Colab (Runtime β†’ T4 GPU).
  2. Run cells top-to-bottom. The notebook already:
    • loads unsloth/Qwen3-0.6B + LoRA (r=16, 4-bit),
    • runs the random and heuristic baselines against the live Space,
    • configures GRPOConfig (batch 2 Γ— grad-accum 4, 4 generations, lr 5e-6, fp16),
    • trains, then plots the reward curve and a baseline-vs-trained comparison.
  3. Watch for these known gotchas:
    • Space must be v2 (step 1) or the reward signal is hacked.
    • If from openenv.core import GenericEnvClient errors, pin openenv-core to the version the Space was built with (pip install openenv-core==<ver>).
    • 0.6B cold-start: if completions rarely contain valid JSON actions and reward stays ~0, either (a) do the Hermes SFT warm-start in step 4, or (b) raise num_generations and lower lr.
  4. Export results for the poster β€” add this cell at the end of the notebook. The schema is exactly what poster_experiment.py auto-detects:
    import json
    json.dump({
        "trained": trained_scores,   # per-seed final scores of the GRPO agent
        "reward_curve": [e["reward"] for e in trainer.state.log_history if "reward" in e],
    }, open("grpo_results.json", "w"), indent=2)
    
  5. Download grpo_results.json from Colab.

Acceptance: trained score > heuristic (β‰ˆ0.45), ideally β‰₯0.55, with a reward curve that trends up. If trained β‰ˆ heuristic, that's still a publishable result β€” report it honestly (small model, limited steps) rather than inflating.


3. Add the trained-agent bar to Figure 3 (1 min, local β€” paste-and-run)

poster_experiment.py already auto-detects the Colab output. Just drop grpo_results.json (from step 2.5) into the repo root and re-run:

cp /path/to/grpo_results.json .
python poster_experiment.py

This regenerates poster_figure.png with a green Trained (GRPO) bar plus a dashed "heuristic bar" reference line, and β€” if grpo_results.json includes a reward_curve β€” also writes fig5_grpo_reward_curve.png (Figure 5). With no file present the script is unchanged (baseline-only). No code edits needed.


4. (Optional, high-value) Hermes-Agent warm-start pipeline

This is the "agent that grows with you" enhancement. It both fixes the 0.6B cold-start and gives the poster a second experimental axis.

4a. Install Hermes and verify its real APIs

git clone https://github.com/nousresearch/hermes-agent && cd hermes-agent
# follow its README installer; then confirm the exact names of:
#   - the batch trajectory generation script (README calls it "batch trajectory generation")
#   - the trajectory compression / training-export utility
#   - the skills directory + agentskills.io skill format

The public README does not print these script names. Confirm them in the cloned repo before scripting against them; update step 4c accordingly.

4b. Register the HCM:21 MCP server with Hermes

The server is built and tested (integrations/hcm21_mcp_server.py). Launch it:

pip install -e ".[server,mcp]"
python -m integrations.hcm21_mcp_server     # stdio MCP transport

Add it to Hermes' MCP config (Hermes "Connect any MCP server"). Hermes can now call hcm21_reset → hcm21_available_actions → hcm21_step ×N → hcm21_score. Author a Hermes skill that runs the scan→plan→produce→control loop and reads the final score — this is the procedural-memory artifact the poster describes.

4c. Generate expert trajectories β†’ SFT warm-start β†’ GRPO

1. Run Hermes (strong model) over the 4 scenarios Γ— several seeds via the MCP
   server β†’ use Hermes' batch trajectory generation to collect episodes.
2. Use Hermes' trajectory compression to export an SFT dataset of
   (prompt, action-sequence) pairs from the high-scoring episodes.
3. SFT unsloth/Qwen3-0.6B on that dataset (1–2 epochs), THEN run the GRPO
   notebook from that checkpoint instead of the base model.

Acceptance: SFT-then-GRPO reaches a higher score / converges faster than GRPO-from-base. Report the three-regime table from POSTER_DRAFT.md Β§6 (GRPO-only vs Hermes-skills-only vs hybrid).


5. Assemble and submit the poster (by July 26, 11:59 PM PDT)

  1. Layout (A0 portrait or the conference's stated size β€” confirm on Sessionize):
    • Top: title + Β§1 motivation.
    • Left column: Β§2 environment + Fig 1 (architecture), Fig 2 (phase loop).
    • Center (hero): Β§3 finding + Β§4 fix + Fig 3 (recalibration) β€” biggest panel.
    • Right column: Β§5 training + Fig 4 (GRPO loop) + Fig 5 (reward curve), then Β§6 Hermes + the three-regime table.
    • Footer: Β§7 "Try it" QR codes (Space, GitHub, Colab) + MIT license.
  2. Build the PDF in your tool of choice (PowerPoint/Keynote/Figma/LaTeX tikzposter or beamerposter). Embed the committed PNGs at 150+ dpi.
  3. Submit on Sessionize (link from the CFP page). Max 2 speakers. Paste the abstract from POSTER_DRAFT.md (Title + Β§1). Upload the PDF if requested.
  4. Diversity rule: if 3+ presenters, gender diversity is required β€” keep to ≀2 or compose accordingly.
  5. Print & ship: presenters print and bring their own poster (CFP requirement). Order the print ~2 weeks before Oct 20 (San Jose).

CFP dates

  • Submit: Sun July 26, 11:59 PM PDT
  • Notifications: Mon Aug 17 Β· Schedule: Tue Aug 18 Β· Event: Oct 20–21, San Jose
  • Accepted presenters get complimentary passes; slides due before the event.

6. Final verification checklist

  • HF Space redeployed; live do-nothing score β‰ˆ 0.25 (not 0.66)
  • python -m pytest tests/ -q β†’ 43 passing
  • python poster_experiment.py β†’ null 0.247, random 0.262, heuristic 0.448 (deterministic)
  • python poster_diagrams.py β†’ Figs 1, 2, 4 regenerate
  • Colab GRPO run complete; grpo_results.json + reward curve saved (Fig 5)
  • Trained-agent bar added to Fig 3
  • (Optional) Hermes warm-start results + three-regime table
  • Poster PDF assembled with all 5 figures + QR codes
  • Submitted on Sessionize before July 26
  • Caveat resolved: Hermes script names verified against the actual repo