# HCM:21 — PyTorch Conference NA 2026 Poster Draft > Status: Path A complete — reward recalibrated, policy separation proven on real > data, 43/43 tests passing. Figure `poster_figure.png` is generated from > `poster_experiment.py` (raw numbers in `results.json`). Hermes-Agent > enhancement designed below. --- ## Title **Primary:** *HCM:21 — Reward Design for Long-Horizon LLM Agents in an OpenEnv Workforce-Planning Benchmark* **Subtitle:** *Why naïve efficiency rewards get hacked by downsizing, and a health-gated fix that restores policy separation — plus a skill-growing (Hermes-Agent) training pipeline.* Two-author max (per CFP). Submit on Sessionize by **July 26, 11:59 PM PDT**. --- ## 1. Motivation (top-left) - **Long-horizon planning + sparse/delayed rewards** is the hardest axis of agent evaluation. HCM:21 puts an LLM agent in the seat of a Chief Human Capital Officer steering a 200–500-person company across **6 quarters / 150–200 steps**, scored on validated Fitz-enz HR analytics (HCVA, HCROI, QIPS). - **A non-code domain.** Most agent benchmarks are code/math/games. Workforce strategy is a fresh, economically grounded long-horizon testbed. - **OpenEnv-native.** Built to the `meta-pytorch/OpenEnv` contract → plugs directly into the PyTorch RL-environments ecosystem (the conference's own turf). ## 2. The environment (left column) - **4-phase × 6-quarter state machine** (`phases.py`): Scanning → Planning → Producing → Controlling, with per-phase minimum-action gating. - **Company simulation** (`company.py`): 5 departments, Cobb-Douglas production linking headcount × skills × engagement → output; per-employee performance, engagement, flight-risk, promotability, salary, tenure. - **Fitz-enz metrics** (`metrics.py`): HCVA = (Rev − non-employment cost)/FTE; HCROI = Rev/employment cost; QIPS composite; Five Indexes of Change; Employee Value composite. - **Stochastic events** (`events.py`): market downturns, competitor poaching, exec departures, budget cuts — force re-planning. - **4 scenario tasks**: high-turnover, budget-cuts, rapid-growth, balanced-opt. - *Figure 1: architecture diagram. Figure 2: the phase × quarter loop.* ## 3. Finding: the reward was being hacked (center — the hook) Running three **non-LLM** policies through the env (5 seeds each) exposed a reward-hacking failure in the original scoring: | Policy | v1 score (sigmoid-centred) | What actually happened | |---|---|---| | Null (do-nothing) | **0.672** | workforce collapses 300→41; HCVA/HCROI *rise* +75% | | Random | 0.652 | — | | Heuristic (real strategy) | 0.649 | keeps 96 staff, +65% HCVA, QIPS maintained | **All three were indistinguishable.** Root cause: HCVA (per-FTE) and HCROI (per-employment-cost) are *ratios* that a do-nothing agent inflates by letting attrition shrink the denominator. The original scoring also mapped "no change" to 0.5 via a sigmoid, handing out free credit. The benchmark rewarded **gutting the company**. > This is the poster's intellectual hook: *efficiency ratios are reward-hackable > in long-horizon economic environments, and a do-nothing baseline is the cheapest > way to discover it.* ## 4. The fix: health-gated scoring (center-right) `scoring.py` v2 (this work): 1. **One-sided ramps** — "no improvement" scores **0**, not 0.5 (no free credit). 2. **Health gate** — HCVA/HCROI credit is multiplied by workforce retention; an agent that keeps <60% of staff cannot bank efficiency "gains." 3. **Sustainability as a first-class term (30%)** — headcount × engagement retention. Solvency is itself gated by health (a profitable shell ≠ healthy). Measured after recalibration (5 seeds, real runs, fully deterministic): | Policy | v2 score | ± | |---|---|---| | Null (do-nothing) | **0.247** | 0.017 | | Random | 0.262 | 0.028 | | Heuristic | **0.448** | 0.029 | **Separation restored:** strategy now scores ~1.8× no-strategy, with headroom to ~0.65+ for a trained agent. All 43 unit tests pass. Runs are bit-reproducible — all stochastic streams (turnover, hire generation) are seeded from the episode seed, so the same seed yields the same score every time. - *Figure 3 (hero): `poster_figure.png` — left: v1 collapse vs v2 separation; right: sparse quarterly-reward trajectory.* ## 5. Training pipeline (right column) - **GRPO** (TRL `GRPOTrainer`) + **Unsloth**, `unsloth/Qwen3-0.6B`, LoRA r=16, 4-bit, T4-friendly (batch 2 × grad-accum 4, 4 generations, lr 5e-6). - **Reward** = the env's health-gated quarterly reward; completions are parsed for JSON action sequences and executed in the live OpenEnv server (HF Space). - *Figure 4: GRPO loop diagram. Figure 5 (Path A follow-up): post-fix GRPO reward curve + trained-agent bar added to Fig 3.* ## 6. Skill-growing agents with Hermes-Agent (right column — the differentiator) A from-scratch GRPO'd 0.6B model struggles to even emit valid 150-step action sequences (cold start). We pair HCM:21 with **Nous Research's Hermes-Agent** ("the agent that grows with you") to attack this on three fronts: 1. **Expert rollouts → SFT warm-start, then GRPO.** Run Hermes (model-agnostic) as the CHCO against the HCM:21 OpenEnv server; use its **batch trajectory generation + trajectory compression** to build a dataset of successful 6-quarter episodes; SFT Qwen3-0.6B on it as a warm start *before* GRPO. Fixes the cold-start that makes raw GRPO on a small model fail. 2. **Skills as procedural memory for long-horizon competence.** Hermes' skills system (agentskills.io standard, autonomous skill creation/self-improvement) lets the agent distill reusable HR sub-routines — `diagnose_at_risk_depts`, `rebalance_hiring_to_offset_turnover`, `protect_engagement_under_budget_cut`. "Grows with you" = the agent accumulates HR strategy across episodes. This is a direct mechanism for the long-horizon theme. 3. **Memory for state beyond the context window.** By Q4 full state history exceeds 128K tokens. Hermes' agent-curated memory (`MEMORY.md`/`USER.md`) + session search/summarization lets the agent compress prior quarters into durable memory instead of carrying raw history — addressing the "beyond context memory limits" requirement head-on. **Experimental matrix the poster can show (learning in weight-space vs. memory/skill-space vs. hybrid):** | Regime | Mechanism | Hypothesis | |---|---|---| | GRPO-only (0.6B) | weight updates | strong but cold-start-limited | | Hermes skills-only | in-context skills + memory, no weight updates | strong sample-efficiency, larger model | | Hermes→SFT→GRPO (hybrid) | distill expert trajectories into a small model | best score-per-param | **Integration path:** expose HCM:21 actions as an **MCP tool server** (Hermes supports MCP natively) → Hermes calls `step()`/`reset()` as tools with zero glue. > Caveat to verify before final: exact Hermes script names for batch generation / > trajectory compression aren't in the public README — confirm against the repo > (`agent/`, `skills/`, any `batch_runner.py`) before printing command names. ## 7. Try it (footer) QR codes → HF Space (`paretooptimal-hcm21.hf.space`), GitHub repo, Colab notebook. MIT licensed. OpenEnv-compliant (`openenv validate`). --- ## Figure checklist | # | Figure | Status | Source | |---|---|---|---| | 1 | Architecture diagram | **DONE** | `poster_diagrams.py` → `fig1_architecture.png` | | 2 | Phase × quarter loop | **DONE** | `poster_diagrams.py` → `fig2_phase_loop.png` | | 3 | **v1→v2 recalibration + reward trajectory** | **DONE (real)** | `poster_experiment.py` → `poster_figure.png` | | 4 | GRPO training loop | **DONE** | `poster_diagrams.py` → `fig4_grpo_loop.png` | | 5 | Post-fix GRPO reward curve + trained-agent bar | NEEDS Colab run | auto-built by `poster_experiment.py` when `grpo_results.json` is present | Four of five figures are generated and committed. Fig 5 (and the trained-agent bar on Fig 3) is wired to auto-generate: drop the Colab `grpo_results.json` in the repo root and re-run `poster_experiment.py` — no code edits. See RUNBOOK.md §2–3. ## Hermes integration assets (in repo) `integrations/hcm21_mcp_server.py` — an MCP tool server exposing `hcm21_reset / hcm21_step / hcm21_available_actions / hcm21_state / hcm21_score`. Hermes (MCP client) drives the env natively. Verified end-to-end in-process. Setup in RUNBOOK.md step 5. ## Reproducibility appendix (for reviewers) ```bash pip install -e ".[server]" # faker, numpy, fastapi, pydantic python poster_experiment.py # writes results.json + poster_figure.png python -m pytest tests/ -q # 43 passing ``` Raw measured numbers live in `results.json`.