Run evidence
Observed and managed Runs retain visible events, artifacts, outcomes, corrections, and observation gaps.
Read-only Hugging Face mirror
A local-first workbench for retaining Runs, diagnosing evidence-supported issues, comparing one bounded AGENTS.md or Skill change, and publishing or rolling it back with human approval.
5beb058d946aThe workbench closes the loop after a Run without becoming a second Agent or granting the Agent approval authority.
Observed and managed Runs retain visible events, artifacts, outcomes, corrections, and observation gaps.
Issues keep instruction, Skill, tool, environment, permission, validation, model, and unknown causes distinct.
Four-cell comparison, exact diff, explicit approval, conflict-safe publication, and rollback history.
The sequence is explanatory. Install and execute the actual project locally from its canonical repository.
Keep the completed Run and its visible evidence gaps.
Connect one recurring issue to exact evidence and counterevidence.
Evaluate one bounded capability-file candidate in isolated worktrees.
A human reviews the exact diff and objective verifier cells.
Hash checks prevent overwriting a target that changed concurrently.
Synthetic or presentation-only images from the canonical repository.


This page preserves the project's published limits instead of turning engineering checks into outcome claims.
The full service binds to loopback and stores product data locally.
Only locally exposed events are retained; gaps remain explicit.
The Wiki study is three replicates, one synthetic grader, and two model versions in one family.
Local files below are curated public mirrors. GitHub remains authoritative for development, releases, and issue history.