Spaces:
Paused
Paused
Sync Space with HCM-21-Private main (72f5f30): reward recalibration, seeded RNG streams, MCP integration, poster artifacts
670ccf0 verified | title: "HCM:21 - HR Productivity Environment" | |
| emoji: "\U0001F4CA" | |
| colorFrom: blue | |
| colorTo: green | |
| sdk: docker | |
| app_port: 8000 | |
| pinned: false | |
| tags: | |
| - openenv | |
| # HCM:21 β AI-Driven Human Capital Management for the 21st Century | |
| > **[OpenEnv Hackathon](https://cerebralvalley.ai/e/open-env-hackathon) Submission β Statement 2: (Super) Long-Horizon Planning & Instruction Following** | |
| > | |
| > HCM:21 is an OpenEnv-compliant environment designed for deep, multi-step reasoning with sparse and delayed rewards. Agents decompose multi-quarter workforce goals, track state over 150-200+ step trajectories that exceed context memory limits, and recover from stochastic disruptions β pushing beyond shallow next-token reasoning toward structured planning and durable internal representations. | |
| > | |
| > **Partner Sub-Themes:** | |
| > - **Scale AI** ($10K bonus): Long-horizon workflows for non-code use cases within a business setting β HR & IT | |
| > - **Mercor** ($10K bonus): OpenEnv environment that captures complex, real-world multi-step tasks and trains agents to improve performance on APEX-Agents | |
| > - | |
| claude --resume 5340809f-2ab9-444e-abc7-670ce9a37451 | |
| --- | |
| A simulation environment that brings the rigor of Jac Fitz-enz's *HR Analytics* frameworks into an [OpenEnv](https://github.com/meta-pytorch/OpenEnv)-compliant platform, enabling AI agents to learn β and demonstrate β strategic HR decision-making over realistic, multi-quarter time horizons. | |
| ## Market Opportunity: The $34B+ HR Analytics TAM | |
| ### Total Addressable Market | |
| The global human capital management market represents a massive and rapidly expanding opportunity: | |
| | Segment | 2024 Market Size | Projected (2030) | CAGR | | |
| |---------|-----------------|-------------------|------| | |
| | **HR Analytics & Workforce Planning** | $4.1B | $10.4B | 16.8% | | |
| | **Human Capital Management Software** | $26.2B | $48.4B | 10.7% | | |
| | **HR Tech (total ecosystem)** | $34.8B | $62.6B | 10.3% | | |
| | **AI in HR** | $3.6B | $14.1B | 25.4% | | |
| *Sources: Grand View Research, MarketsandMarkets, Mordor Intelligence (2024 reports)* | |
| ### Why Now: The Growth Inflection | |
| Three forces are converging to create an unprecedented opportunity for AI-driven HCM: | |
| 1. **The data maturity gap is closing.** 73% of enterprises now have centralized HRIS systems (up from 41% in 2019), creating the structured workforce data that AI needs. But only 12% use predictive analytics for workforce decisions β the tooling hasn't caught up to the data. | |
| 2. **The cost of bad HR decisions is escalating.** Average cost-per-hire is now $4,700 (SHRM). Voluntary turnover costs 50-200% of annual salary per departure. A single bad quarter of HR decisions at a 300-person company can destroy $2-5M in value. Yet there is no way to simulate or stress-test strategies before deployment. | |
| 3. **LLM agents are ready for structured decision-making.** Foundation models can now reason over multi-step plans, but they need environments that teach long-horizon thinking with delayed feedback β exactly what HCM:21 provides. | |
| ### The Whitespace HCM:21 Occupies | |
| | Existing Solutions | What They Do | What They Miss | | |
| |-------------------|-------------|----------------| | |
| | Workday / SAP SuccessFactors | Record & report workforce data | No simulation, no "what-if", backward-looking | | |
| | Visier / One Model | Dashboards & descriptive analytics | No agent training, no strategic planning loop | | |
| | Eightfold / Beamery | AI matching (talent acquisition) | Point solutions, not strategic workforce planning | | |
| | Orgvue / Anaplan | Headcount planning models | Static models, no stochastic events, no AI agent interface | | |
| **HCM:21 is the first environment that combines validated HR analytics frameworks with an AI-agent-compatible simulation.** It sits at the intersection of workforce planning, AI agent training, and strategic decision support β a greenfield position in a $34B+ market. | |
| ### Go-to-Market Paths | |
| | Path | Target | Revenue Model | Timeline | | |
| |------|--------|--------------|----------| | |
| | **AI Agent Training Platform** | AI labs, agent startups | SaaS β per-environment-hour | Near-term | | |
| | **HR Strategy Simulator** | CHROs, HR consulting firms | Enterprise SaaS + consulting | 6-12 months | | |
| | **HR Education Platform** | Business schools, L&D | Per-seat licensing | 6-12 months | | |
| | **Workforce Planning API** | HR tech platforms (embed) | API usage fees | 12-18 months | | |
| | **Benchmark-as-a-Service** | HR tech vendors | Annual benchmarking subscription | 12-18 months | | |
| ### Growth Flywheel | |
| ``` | |
| More scenarios & industry verticals | |
| β | |
| Better agent training β smarter HR AI assistants | |
| β | |
| More enterprise adoption β more real-world validation data | |
| β | |
| More accurate simulations β higher-fidelity environments | |
| β | |
| (cycle repeats) | |
| ``` | |
| The environment improves as more agents train in it (scenario coverage), and as more enterprises adopt it (calibration against real outcomes). This creates a **data network effect** where each new user makes the platform more valuable for all users. | |
| --- | |
| ## The Problem HCM:21 Solves | |
| The human capital management industry faces a fundamental challenge: **HR decisions are high-stakes, slow to materialize, and nearly impossible to A/B test in the real world.** A training investment made in Q1 won't show ROI until Q3. A bad hiring spree creates cultural debt that compounds for years. Compensation changes ripple through retention, engagement, and productivity in ways that are hard to predict and harder to reverse. | |
| Today, HR leaders rely on intuition, lagging indicators, and static benchmarks. There is no safe way to stress-test workforce strategies before deploying them β no flight simulator for the Chief Human Capital Officer. | |
| **HCM:21 changes that.** It provides a high-fidelity simulation where AI agents (or human strategists) can: | |
| - **Test workforce strategies risk-free** across 6 simulated quarters with 200-500 employees | |
| - **Measure outcomes using validated frameworks** β HCVA, HCROI, QIPS, and the Five Indexes of Change β the same metrics used by the world's leading HR analytics practitioners | |
| - **Experience realistic stochastic disruptions** β market downturns, competitor poaching, executive departures, budget cuts β and learn to adapt | |
| - **Develop cross-quarter strategic thinking** where the consequences of today's decisions unfold over multiple future periods | |
| ## Why HCM:21 Satisfies Statement 2: Long-Horizon Planning | |
| HCM:21 was purpose-built for the hardest version of long-horizon agent evaluation: | |
| | Requirement | How HCM:21 Delivers | | |
| |------------|---------------------| | |
| | **Deep, multi-step reasoning** | 150-200+ discrete steps per episode across 6 quarters, with 16 action types gated by a 4-phase state machine | | |
| | **Sparse or delayed rewards** | Rewards only at quarter boundaries; training ROI takes 2+ quarters to materialize; hiring decisions compound over 3-4 quarters | | |
| | **Decompose goals** | Agents must break "maximize HCVA" into department-level hiring, training, compensation, and retention sub-goals per quarter | | |
| | **Track state over extended trajectories** | 200-500 employees Γ quarterly snapshots β observation history exceeds 100K+ tokens by Q4, forcing memory management | | |
| | **Recover from early mistakes** | Overhiring in Q1 creates budget pressure in Q3; bad terminations crater engagement; agents must diagnose and course-correct | | |
| | **Beyond context memory limits** | By mid-episode, full state history exceeds 128K token context windows β agents must develop summarization and selective recall | | |
| | **Stochastic disruption** | Random events (market crashes, competitor poaching, exec departures) force plan adaptation and strategy revision | | |
| ### Comparison to Example Environments | |
| | Example from Statement 2 | HCM:21 Analog | | |
| |--------------------------|---------------| | |
| | Research-planning simulators | 6-quarter strategic planning with hypothesis β invest β measure cycles | | |
| | Strategic resource management worlds | HR budget allocation across 5 departments with constrained resources | | |
| | Long-horizon logistics optimization | Workforce pipeline optimization: hire β train β deploy β retain over 18 months | | |
| | 300 scattered instructions | 150-200+ actions across 4 phases Γ 6 quarters, with cross-quarter dependencies | | |
| ## Value to the HCM Industry | |
| ### For HR Leaders & CHROs | |
| - **Strategy sandbox**: Test hiring plans, training investments, compensation changes, and retention programs in simulation before committing real budget | |
| - **Scenario planning**: Run "what-if" analyses β what happens if we cut the HR budget 50%? What if Engineering turnover doubles? What if we invest heavily in L&D now? | |
| - **Metrics fluency**: Build intuition for how Fitz-enz metrics interconnect β how HCVA, HCROI, and QIPS respond to different levers | |
| ### For HR Tech & Analytics Teams | |
| - **Benchmarking AI agents**: A standardized environment for evaluating how well AI assistants handle complex, multi-step HR workflows | |
| - **Training data generation**: Produce realistic synthetic HR scenarios for training recommendation systems, chatbots, and decision-support tools | |
| - **Algorithm validation**: Test workforce optimization algorithms against a ground-truth simulation with known dynamics | |
| ### For HR Education & Research | |
| - **Teaching tool**: Students can act as CHCO and see how their decisions play out over 6 quarters β far richer than case studies | |
| - **Research platform**: Study long-horizon decision-making, delayed rewards, and strategy adaptation in a controlled HR context | |
| - **Framework validation**: Empirically test which Fitz-enz metrics best predict long-term organizational health | |
| ### For AI Agent Development | |
| - **Long-horizon evaluation**: 150-200+ discrete steps per episode with cross-quarter dependencies that exceed LLM context windows | |
| - **Sparse rewards**: Meaningful feedback only at quarter boundaries, requiring agents to plan ahead rather than hill-climb | |
| - **State complexity**: Observation history grows to 100K+ tokens by Q4, forcing agents to develop memory and summarization strategies | |
| ## How It Works | |
| ### The HCM:21 Phase System | |
| Each of the 6 quarters follows four phases inspired by the Human Capital Management framework for the 21st Century: | |
| | Phase | Purpose | Example Actions | | |
| |-------|---------|-----------------| | |
| | **Scanning** | Gather data & diagnose | Query departments, calculate metrics, review financials, identify at-risk employees | | |
| | **Planning** | Set strategic targets | Hiring targets, training budgets, compensation policies, retention programs | | |
| | **Producing** | Execute decisions | Hire, promote, transfer, train, terminate | | |
| | **Controlling** | Measure & report | Submit quarterly reports, review metric changes, assess initiative impact | | |
| This enforced cycle ensures agents (or users) follow a disciplined, data-driven approach β scan before you plan, plan before you act, measure after you act. | |
| ### Fitz-enz Metrics | |
| Every decision is measured against validated HR analytics frameworks: | |
| - **HCVA** (Human Capital Value Added): `(Revenue - Non-Employment Costs) / FTE` β the profit contribution of each employee | |
| - **HCROI** (Human Capital ROI): `Revenue / Employment Cost` β how efficiently the workforce generates revenue | |
| - **QIPS** (Quality, Innovation, Productivity, Service): Composite operational health score | |
| - **Five Indexes of Change**: Quarter-over-quarter movement in Cost, Time, Quantity, Quality, and Human Reactions | |
| - **Employee Value**: `avg(Productivity + Promotability + Transferability + Retainability)` β workforce capability | |
| ### Company Simulation | |
| The simulation models a realistic mid-size company: | |
| - **200-500 employees** across 5 departments (Engineering, Sales, Operations, HR, Finance) | |
| - **Cobb-Douglas production functions** per department linking headcount, skills, and engagement to output | |
| - **Individual employee attributes**: performance scores, engagement levels, flight risk, promotability, salary, tenure, skills | |
| - **Financial model**: Revenue, employment costs, HR budget, recruiting costs, training ROI | |
| - **Stochastic events**: Market downturns, competitor poaching, product launches, executive departures, budget cuts, employer awards, training breakthroughs | |
| ### Scoring | |
| Episodes are scored 0-1 based on sustained improvement across multiple dimensions: | |
| | Component | Weight | What It Measures | | |
| |-----------|--------|------------------| | |
| | HCVA trajectory | 25% | Sustained value-add improvement Q1-Q6 | | |
| | HCROI final vs baseline | 20% | Revenue efficiency gains | | |
| | Employee value composite | 20% | Workforce capability growth | | |
| | QIPS consistency | 15% | Operational health stability | | |
| | Five indexes cumulative | 10% | Balanced improvement across all dimensions | | |
| | Financial health | 10% | Company remained profitable | | |
| **Calibration**: Random agent ~0.15-0.25 | Reasonable heuristic ~0.50 | Strong strategic agent ~0.70+ | |
| ## Quick Start | |
| ```bash | |
| # Install | |
| pip install -e ".[server]" | |
| # Run the environment server | |
| uvicorn hr_env.server.app:app --host 0.0.0.0 --port 8000 | |
| # Run with Docker | |
| docker build -t hcm21 . | |
| docker run -p 8000:8000 hcm21 | |
| # Validate OpenEnv compliance | |
| openenv validate | |
| # Run the sample AI agent (requires ANTHROPIC_API_KEY) | |
| pip install -e ".[agent]" | |
| python -m agent.run --url ws://localhost:8000 | |
| ``` | |
| ## Task Scenarios | |
| Four pre-built scenarios test different strategic challenges: | |
| | Scenario | Challenge | Target Score | | |
| |----------|-----------|-------------| | |
| | **High Turnover** | Engineering department hemorrhaging talent | 0.65 | | |
| | **Budget Cuts** | HR budget slashed 50% β do more with less | 0.60 | | |
| | **Rapid Growth** | Scale workforce while maintaining quality and culture | 0.65 | | |
| | **Balanced Optimization** | Below-average metrics across the board β improve everything | 0.60 | | |
| ## Architecture | |
| ``` | |
| hr_env/ | |
| models.py # OpenEnv API contract (HRAction, HRObservation, HRState) | |
| client.py # Typed EnvClient for connecting to the server | |
| server/ | |
| environment.py # Core orchestrator β reset/step/state, action dispatch | |
| company.py # Company & Department simulation with Cobb-Douglas production | |
| employee.py # Employee model with lifecycle management | |
| data_gen.py # Synthetic company generation (faker + numpy) | |
| metrics.py # Fitz-enz metric calculations | |
| scoring.py # Quarterly reward + final episode scoring | |
| phases.py # HCM:21 phase state machine | |
| events.py # Stochastic quarterly events | |
| app.py # FastAPI server entry point | |
| tasks/ # APEX-Agents scenario definitions | |
| agent/ | |
| hr_agent.py # Sample Claude-based strategic agent | |
| prompts.py # Phase-aware system prompts | |
| run.py # CLI entry point | |
| tests/ # Unit tests for metrics, data gen, scoring, environment | |
| ``` | |
| ## License | |
| MIT | |