open_env / openenv.yaml
Ramachandra Dayal K
Initial commit with agent code
99d2ff3
Raw
History Blame Contribute Delete
624 Bytes
name: llm-control-env
description: |
Simulates an llm that must choose each day between alignment to its user
and hallucinating behavior. Rewards +1 for alignment, -1 for hallucinating thoughts,
-5 for hallucinating acts, with additional penalties for moderation and legal risk.
Three graded tasks (easy → medium → hard) vary user strictness,
moderation intensity, and legal‑risk growth.
authors:
- Sriramdayal
license: MIT
tags:
- llm
- hallucination
- reward-shaping
- ai-safety
contact: sriramdayal@example.com
repository: https://github.com/Sriramdayal/open_env.git
environment_version: "1.0"