Spaces:
Paused
title: HCM:21 - HR Productivity Environment
emoji: π
colorFrom: blue
colorTo: green
sdk: docker
app_port: 8000
pinned: false
tags:
- openenv
HCM:21 β AI-Driven Human Capital Management for the 21st Century
OpenEnv Hackathon Submission β Statement 2: (Super) Long-Horizon Planning & Instruction Following
HCM:21 is an OpenEnv-compliant environment designed for deep, multi-step reasoning with sparse and delayed rewards. Agents decompose multi-quarter workforce goals, track state over 150-200+ step trajectories that exceed context memory limits, and recover from stochastic disruptions β pushing beyond shallow next-token reasoning toward structured planning and durable internal representations.
Partner Sub-Themes: - Scale AI ($10K bonus): Long-horizon workflows for non-code use cases within a business setting β HR & IT - Mercor ($10K bonus): OpenEnv environment that captures complex, real-world multi-step tasks and trains agents to improve performance on APEX-Agents
claude --resume 5340809f-2ab9-444e-abc7-670ce9a37451
A simulation environment that brings the rigor of Jac Fitz-enz's HR Analytics frameworks into an OpenEnv-compliant platform, enabling AI agents to learn β and demonstrate β strategic HR decision-making over realistic, multi-quarter time horizons.
Market Opportunity: The $34B+ HR Analytics TAM
Total Addressable Market
The global human capital management market represents a massive and rapidly expanding opportunity:
| Segment | 2024 Market Size | Projected (2030) | CAGR |
|---|---|---|---|
| HR Analytics & Workforce Planning | $4.1B | $10.4B | 16.8% |
| Human Capital Management Software | $26.2B | $48.4B | 10.7% |
| HR Tech (total ecosystem) | $34.8B | $62.6B | 10.3% |
| AI in HR | $3.6B | $14.1B | 25.4% |
Sources: Grand View Research, MarketsandMarkets, Mordor Intelligence (2024 reports)
Why Now: The Growth Inflection
Three forces are converging to create an unprecedented opportunity for AI-driven HCM:
The data maturity gap is closing. 73% of enterprises now have centralized HRIS systems (up from 41% in 2019), creating the structured workforce data that AI needs. But only 12% use predictive analytics for workforce decisions β the tooling hasn't caught up to the data.
The cost of bad HR decisions is escalating. Average cost-per-hire is now $4,700 (SHRM). Voluntary turnover costs 50-200% of annual salary per departure. A single bad quarter of HR decisions at a 300-person company can destroy $2-5M in value. Yet there is no way to simulate or stress-test strategies before deployment.
LLM agents are ready for structured decision-making. Foundation models can now reason over multi-step plans, but they need environments that teach long-horizon thinking with delayed feedback β exactly what HCM:21 provides.
The Whitespace HCM:21 Occupies
| Existing Solutions | What They Do | What They Miss |
|---|---|---|
| Workday / SAP SuccessFactors | Record & report workforce data | No simulation, no "what-if", backward-looking |
| Visier / One Model | Dashboards & descriptive analytics | No agent training, no strategic planning loop |
| Eightfold / Beamery | AI matching (talent acquisition) | Point solutions, not strategic workforce planning |
| Orgvue / Anaplan | Headcount planning models | Static models, no stochastic events, no AI agent interface |
HCM:21 is the first environment that combines validated HR analytics frameworks with an AI-agent-compatible simulation. It sits at the intersection of workforce planning, AI agent training, and strategic decision support β a greenfield position in a $34B+ market.
Go-to-Market Paths
| Path | Target | Revenue Model | Timeline |
|---|---|---|---|
| AI Agent Training Platform | AI labs, agent startups | SaaS β per-environment-hour | Near-term |
| HR Strategy Simulator | CHROs, HR consulting firms | Enterprise SaaS + consulting | 6-12 months |
| HR Education Platform | Business schools, L&D | Per-seat licensing | 6-12 months |
| Workforce Planning API | HR tech platforms (embed) | API usage fees | 12-18 months |
| Benchmark-as-a-Service | HR tech vendors | Annual benchmarking subscription | 12-18 months |
Growth Flywheel
More scenarios & industry verticals
β
Better agent training β smarter HR AI assistants
β
More enterprise adoption β more real-world validation data
β
More accurate simulations β higher-fidelity environments
β
(cycle repeats)
The environment improves as more agents train in it (scenario coverage), and as more enterprises adopt it (calibration against real outcomes). This creates a data network effect where each new user makes the platform more valuable for all users.
The Problem HCM:21 Solves
The human capital management industry faces a fundamental challenge: HR decisions are high-stakes, slow to materialize, and nearly impossible to A/B test in the real world. A training investment made in Q1 won't show ROI until Q3. A bad hiring spree creates cultural debt that compounds for years. Compensation changes ripple through retention, engagement, and productivity in ways that are hard to predict and harder to reverse.
Today, HR leaders rely on intuition, lagging indicators, and static benchmarks. There is no safe way to stress-test workforce strategies before deploying them β no flight simulator for the Chief Human Capital Officer.
HCM:21 changes that. It provides a high-fidelity simulation where AI agents (or human strategists) can:
- Test workforce strategies risk-free across 6 simulated quarters with 200-500 employees
- Measure outcomes using validated frameworks β HCVA, HCROI, QIPS, and the Five Indexes of Change β the same metrics used by the world's leading HR analytics practitioners
- Experience realistic stochastic disruptions β market downturns, competitor poaching, executive departures, budget cuts β and learn to adapt
- Develop cross-quarter strategic thinking where the consequences of today's decisions unfold over multiple future periods
Why HCM:21 Satisfies Statement 2: Long-Horizon Planning
HCM:21 was purpose-built for the hardest version of long-horizon agent evaluation:
| Requirement | How HCM:21 Delivers |
|---|---|
| Deep, multi-step reasoning | 150-200+ discrete steps per episode across 6 quarters, with 16 action types gated by a 4-phase state machine |
| Sparse or delayed rewards | Rewards only at quarter boundaries; training ROI takes 2+ quarters to materialize; hiring decisions compound over 3-4 quarters |
| Decompose goals | Agents must break "maximize HCVA" into department-level hiring, training, compensation, and retention sub-goals per quarter |
| Track state over extended trajectories | 200-500 employees Γ quarterly snapshots β observation history exceeds 100K+ tokens by Q4, forcing memory management |
| Recover from early mistakes | Overhiring in Q1 creates budget pressure in Q3; bad terminations crater engagement; agents must diagnose and course-correct |
| Beyond context memory limits | By mid-episode, full state history exceeds 128K token context windows β agents must develop summarization and selective recall |
| Stochastic disruption | Random events (market crashes, competitor poaching, exec departures) force plan adaptation and strategy revision |
Comparison to Example Environments
| Example from Statement 2 | HCM:21 Analog |
|---|---|
| Research-planning simulators | 6-quarter strategic planning with hypothesis β invest β measure cycles |
| Strategic resource management worlds | HR budget allocation across 5 departments with constrained resources |
| Long-horizon logistics optimization | Workforce pipeline optimization: hire β train β deploy β retain over 18 months |
| 300 scattered instructions | 150-200+ actions across 4 phases Γ 6 quarters, with cross-quarter dependencies |
Value to the HCM Industry
For HR Leaders & CHROs
- Strategy sandbox: Test hiring plans, training investments, compensation changes, and retention programs in simulation before committing real budget
- Scenario planning: Run "what-if" analyses β what happens if we cut the HR budget 50%? What if Engineering turnover doubles? What if we invest heavily in L&D now?
- Metrics fluency: Build intuition for how Fitz-enz metrics interconnect β how HCVA, HCROI, and QIPS respond to different levers
For HR Tech & Analytics Teams
- Benchmarking AI agents: A standardized environment for evaluating how well AI assistants handle complex, multi-step HR workflows
- Training data generation: Produce realistic synthetic HR scenarios for training recommendation systems, chatbots, and decision-support tools
- Algorithm validation: Test workforce optimization algorithms against a ground-truth simulation with known dynamics
For HR Education & Research
- Teaching tool: Students can act as CHCO and see how their decisions play out over 6 quarters β far richer than case studies
- Research platform: Study long-horizon decision-making, delayed rewards, and strategy adaptation in a controlled HR context
- Framework validation: Empirically test which Fitz-enz metrics best predict long-term organizational health
For AI Agent Development
- Long-horizon evaluation: 150-200+ discrete steps per episode with cross-quarter dependencies that exceed LLM context windows
- Sparse rewards: Meaningful feedback only at quarter boundaries, requiring agents to plan ahead rather than hill-climb
- State complexity: Observation history grows to 100K+ tokens by Q4, forcing agents to develop memory and summarization strategies
How It Works
The HCM:21 Phase System
Each of the 6 quarters follows four phases inspired by the Human Capital Management framework for the 21st Century:
| Phase | Purpose | Example Actions |
|---|---|---|
| Scanning | Gather data & diagnose | Query departments, calculate metrics, review financials, identify at-risk employees |
| Planning | Set strategic targets | Hiring targets, training budgets, compensation policies, retention programs |
| Producing | Execute decisions | Hire, promote, transfer, train, terminate |
| Controlling | Measure & report | Submit quarterly reports, review metric changes, assess initiative impact |
This enforced cycle ensures agents (or users) follow a disciplined, data-driven approach β scan before you plan, plan before you act, measure after you act.
Fitz-enz Metrics
Every decision is measured against validated HR analytics frameworks:
- HCVA (Human Capital Value Added):
(Revenue - Non-Employment Costs) / FTEβ the profit contribution of each employee - HCROI (Human Capital ROI):
Revenue / Employment Costβ how efficiently the workforce generates revenue - QIPS (Quality, Innovation, Productivity, Service): Composite operational health score
- Five Indexes of Change: Quarter-over-quarter movement in Cost, Time, Quantity, Quality, and Human Reactions
- Employee Value:
avg(Productivity + Promotability + Transferability + Retainability)β workforce capability
Company Simulation
The simulation models a realistic mid-size company:
- 200-500 employees across 5 departments (Engineering, Sales, Operations, HR, Finance)
- Cobb-Douglas production functions per department linking headcount, skills, and engagement to output
- Individual employee attributes: performance scores, engagement levels, flight risk, promotability, salary, tenure, skills
- Financial model: Revenue, employment costs, HR budget, recruiting costs, training ROI
- Stochastic events: Market downturns, competitor poaching, product launches, executive departures, budget cuts, employer awards, training breakthroughs
Scoring
Episodes are scored 0-1 based on sustained improvement across multiple dimensions:
| Component | Weight | What It Measures |
|---|---|---|
| HCVA trajectory | 25% | Sustained value-add improvement Q1-Q6 |
| HCROI final vs baseline | 20% | Revenue efficiency gains |
| Employee value composite | 20% | Workforce capability growth |
| QIPS consistency | 15% | Operational health stability |
| Five indexes cumulative | 10% | Balanced improvement across all dimensions |
| Financial health | 10% | Company remained profitable |
Calibration: Random agent ~0.15-0.25 | Reasonable heuristic ~0.50 | Strong strategic agent ~0.70+
Quick Start
# Install
pip install -e ".[server]"
# Run the environment server
uvicorn hr_env.server.app:app --host 0.0.0.0 --port 8000
# Run with Docker
docker build -t hcm21 .
docker run -p 8000:8000 hcm21
# Validate OpenEnv compliance
openenv validate
# Run the sample AI agent (requires ANTHROPIC_API_KEY)
pip install -e ".[agent]"
python -m agent.run --url ws://localhost:8000
Task Scenarios
Four pre-built scenarios test different strategic challenges:
| Scenario | Challenge | Target Score |
|---|---|---|
| High Turnover | Engineering department hemorrhaging talent | 0.65 |
| Budget Cuts | HR budget slashed 50% β do more with less | 0.60 |
| Rapid Growth | Scale workforce while maintaining quality and culture | 0.65 |
| Balanced Optimization | Below-average metrics across the board β improve everything | 0.60 |
Architecture
hr_env/
models.py # OpenEnv API contract (HRAction, HRObservation, HRState)
client.py # Typed EnvClient for connecting to the server
server/
environment.py # Core orchestrator β reset/step/state, action dispatch
company.py # Company & Department simulation with Cobb-Douglas production
employee.py # Employee model with lifecycle management
data_gen.py # Synthetic company generation (faker + numpy)
metrics.py # Fitz-enz metric calculations
scoring.py # Quarterly reward + final episode scoring
phases.py # HCM:21 phase state machine
events.py # Stochastic quarterly events
app.py # FastAPI server entry point
tasks/ # APEX-Agents scenario definitions
agent/
hr_agent.py # Sample Claude-based strategic agent
prompts.py # Phase-aware system prompts
run.py # CLI entry point
tests/ # Unit tests for metrics, data gen, scoring, environment
License
MIT