agent-harness / docs /STUDY2_PREFLIGHT.md
cuber12's picture
Publish agent harness research code and paper artifacts
d61821a verified
|
Raw
History Blame Contribute Delete
1.86 kB

Study 2 preflight and execution

Run these commands only from a clean committed research worktree. No GUI model control is part of the protocol.

PYTHONPATH=src MPLCONFIGDIR=/tmp/agent-harness-mpl .venv/bin/python -m pytest -q
PYTHONPATH=src .venv/bin/python -m agent_harness.cli validate
PYTHONPATH=src .venv/bin/python scripts/audit_study2_design.py
PYTHONPATH=src .venv/bin/python scripts/preflight_study2.py

The outcome-blind preflight uses lms server start/status/stop for lifecycle and LM Studio's native REST API for exclusive residency. It verifies pinned repository heads, dependency versions, disk headroom, exact agent model variants/contexts/reasoning defaults, tokenizer hashes, custom tool calling, and embedding dimensions/normalization. It writes results/reports/study2_preflight.json and unloads all models before stopping the server.

The main and reliability executions are resumable:

PYTHONPATH=src MPLCONFIGDIR=/tmp/agent-harness-mpl .venv/bin/python -m agent_harness.cli run-study2
PYTHONPATH=src MPLCONFIGDIR=/tmp/agent-harness-mpl .venv/bin/python -m agent_harness.cli run-study2-reliability

The main command must yield exactly 840 completed deterministic cells. The reliability command must yield exactly 72 stochastic sensitivity cells. Any incomplete infrastructure attempt is retained under results/infrastructure_attempts/; it is not a scored model outcome.

Completion record

The outcome-blind preflight passed. Execution revision 58933d6fa8af09fcc5a832fb3b523ffb4182bc50 produced exactly 840 main and 72 reliability cells with no missing identity and no scored infrastructure failure. Both run reports end with an empty model-residency set and a stopped LM Studio server. The deterministic analysis at revision bab257bfc30d9959e31ffdd2ba2c9ebf2c5e575b is under results/derived/study2.