Spaces:
Sleeping
Sleeping
OpenEnv Hackathon Requirements (Round 1)
1. Project Objective
Build a complete, real-world OpenEnv environment that an AI agent can learn from through the standard step(), reset(), and state() API.
2. Key Environment Requirements
- Real-World Utility: Must simulate a genuine, practical task (e.g., email triage, code review, scheduling) rather than an abstract game or toy.
- Spec Compliance: Fully implement the OpenEnv specification, including:
- Strongly typed models parsing action/observation boundaries.
- Required interface endpoints (
step(),reset(),state()). - A valid
openenv.yamlconfiguration file.
- Tasks & Grading:
- Must define a minimum of 3 tasks (scaled as easy, medium, and hard).
- Must use programmatic agent graders for each task.
- Graders must produce deterministic scores mathematically bounded between
0.0and1.0.
- Reward Function: Must implement a meaningful, structured reward function featuring partial progress signals and strict penalties for undesirable agent behavior.
- Space Deployment: The final environment must be deployed as a containerized Hugging Face Space and be tagged with
openenv. It must also contain a robust, functioning Dockerfile. - Documentation: Root
README.mdmust clearly document:- Substantive environment description.
- Definitions of the action and observation spaces.
- Setup and usage instructions.
3. Mandatory Technical implementation & Inference Rules
- Environment Variables: Must utilize the following strictly named variables in your configuration:
API_BASE_URL: The API endpoint for the LLM.MODEL_NAME: The model identifier to use for inference.HF_TOKEN: Hugging Face / API Key.
- Inference Client: Participants must use the OpenAI Client for all LLM calls referencing the variables mapped above.
- Inference Script: A functional baseline inference script named exactly
inference.pymust be present in the root directory.
4. Infrastructure Restrictions
- Timeout Limit: The runtime of the
inference.pybaseline evaluation shouldn't exceed 20 minutes. - Resource Constraints: The environment and baseline inference script must be capable of running cleanly on a machine bounded by 2 vCPUs and 8GB memory.
5. Pre-Submission Checklist & Automated Validation
Before submitting, use the openenv validate command or the pre-submission validator script. An automated evaluation gate will require your project to pass the following:
- HF Space Pings: An automated ping to the deployed HF Space URL must return HTTP
200and respond correctly toreset(). - Schema & Spec Audits: Validate
openenv.yamland typed models. - Docker Audits: The Dockerfile must build natively from the submitted project repository without errors.
- Baseline Execution: The root
inference.pyruns start-to-finish without failures and outputs reproducible evaluation scores. - Grader Verification: All 3+ tasks must be enumerated automatically, and the grader must output validated ranges (0.0 to 1.0) for every task.
6. Official Evaluation & Scoring Criteria
- 30% - Real-World Utility: Does the environment model a practical necessity for the RL/agent community?
- 25% - Task & Grader Quality: Are there multiple scalable tasks constructed with clear, non-cheatable, logic-bound grading criteria?
- 20% - Environment Design: Are the action and observation spaces thoughtfully designed? Does the reward logic establish useful agent training curves?
- 15% - Code Quality & Spec Compliance: Does the environment adhere flawlessly to the OpenEnv API endpoints? Does the HF deployment pipeline and Docker config build correctly under automated scrutiny?
- 10% - Creativity & Novelty: Does the project pioneer fresh mechanics or introduce a novel problem domain?
7. Immediate Disqualification Triggers
- Plagiarized or superficially modified variants of existing environments.
- Graders producing static, non-deterministic, or hardcoded values.
- Environments failing deployment orchestration or returning dead pings over standard OpenEnv endpoints.
- Absence of a functional
inference.pyscript bridging the agent.