Spaces:
Sleeping
An execution trace of autocodabench validate
This document presents a complete execution trace of the command
autocodabench validate benchmark/autocodabench_create_bench/competitions/style-trans-fair/ground_truth/bundle
from shell invocation to process exit code. It is written for a reader who
has not previously studied this codebase. Each stage names the relevant
file and function so that the trace can be followed in an editor, or
stepped through interactively: set a breakpoint at each stage's anchor and
use the validate: configurations in .vscode/launch.json (F5 to launch,
F11 to step into a call, and the Call Stack panel to observe the current
position).
Line numbers cited below are anchors as of the commit that added this document; they drift as code changes, but the function names will not.
Overview: the call graph
shell: autocodabench validate <bundle>
ββ cli/main.py: main() β console-script entry point
ββ auth.load_dotenv() β read <cwd>/.env (never overrides real env)
ββ argparse: _build_parser() β parse subcommand + flags
ββ _cmd_validate()
ββ [--judged only] auth preflight
ββ checks/api.py: validate_bundle_path()
ββ resolve dir-or-zip
ββ facts.py: CompetitionFacts.discover()
ββ base.py: CheckContext.from_bundle_dir() β parse competition.yaml ONCE
ββ base.py: run_checks() β loop the check registry
β ββ deterministic.py: 11 checks (gates + findings)
β β ββ SchemaLint β core/bundle_io.py: validate_bundle()
β ββ attestations.py: 5 human-only boxes
ββ [--judged only] judged.py: run_judged_checks()
ββ report.py: ValidationReport
ββ print markdown or JSON β exit 0 (ok) / 1 (gated)
Two design properties organize everything that follows:
competition.yamlis parsed exactly once into a plain dictionary, and every check reads from that dictionary (together with the bundle directory, for questions of file existence). There is no other shared state.- Checks occupy three epistemic tiers, and the tiers never mix: deterministic (code computes PASS/FAIL β the only tier that can gate), judged (an LLM grades a rubric β advisory FINDINGs only), and attestation (criteria only a human can verify β surfaced as unchecked boxes).
1. Stage 0 β from shell to Python
autocodabench is not a shell script. pyproject.toml declares:
[project.scripts]
autocodabench = "autocodabench.cli.main:main"
When the package is installed with pip install -e ., pip writes a small
launcher onto the PATH that imports autocodabench.cli.main and calls
main(). The first line of project code that executes is therefore
main in src/autocodabench/cli/main.py.
(validate is a subcommand of the single autocodabench entry
point, so the validator behaves as a standalone tool that can be pointed
at any bundle, whether hand-written or generated. validate is retained
as a back-compatible alias for the same subcommand.)
2. Stage 1 β CLI: argument parsing, .env loading, and dispatch
File: src/autocodabench/cli/main.py
def main(argv=None) -> int:
from ..auth import load_dotenv
load_dotenv() # <cwd>/.env, if present
args = _build_parser().parse_args(argv)
return args.func(args) # β _cmd_validate for the validate subcommand
_build_parser() registers the subcommand and its flag surface:
p = sub.add_parser("validate", aliases=["validate"], ...)
_add_validate_args(p) # bundle, --facts, --judged, --backend, --model, --json
p.set_defaults(func=_cmd_validate)
_add_validate_argsdefines the flag surface. With no flags, the invocation follows the fully keyless path: deterministic and attestation tiers only, with no LLM and no network access.load_dotenv()(auth.py) is a small (~20-line) stdlib-only parser: it readsKEY=VALUElines from./.env, toleratingexportprefixes, quotes, and comments, and never overrides a variable already set in the real environment. This is the mechanism by which anANTHROPIC_API_KEY=...entry in a.envfile becomes visible to the--judgedpath.
Control then passes to _cmd_validate(args):
def _cmd_validate(args) -> int:
from ..checks import validate_bundle_path
if args.judged and not _require_live_claude_auth(args.backend):
return 2 # refused before any work
backend = None
if args.judged and (args.backend or args.model):
from ..backends import resolve_backend
backend = resolve_backend(args.backend, model=args.model)
report = validate_bundle_path(args.bundle, facts_path=args.facts,
judged=args.judged, backend=backend)
print(report.to_markdown()) # or json.dumps(report.to_dict()) with --json
return 0 if report.ok else 1
Three details merit attention:
- The auth preflight fires only for
--judgedon the Claude backend.ensure_live_auth()(auth.py) checks for a subscription login or anANTHROPIC_API_KEY; if neither exists and the process is attached to an interactive terminal, it guides the user through logging in or pasting a key (input hidden, with an optional save to./.env). Non-interactive contexts receive an explicit refusal with guidance and exit code 2, rather than an opaque SDK error well into the run. --backend ollama:llama3.1(oropenai:β¦, or a URL) routes the judged tier throughbackends/openai_compat.pyinstead of Claude. The backend is treated as a measured variable, not a hard binding.- Imports are deliberately lazy (placed inside the function): the keyless path never imports the SDK, which is what keeps continuous-integration and reviewer machines free of credentials.
3. Stage 2 β entering the check framework
File: src/autocodabench/checks/api.py
validate_bundle_path is a one-line synchronous wrapper:
def validate_bundle_path(bundle, *, facts_path=None, judged=False, backend=None):
return asyncio.run(validate_bundle_path_async(...))
(The implementation is asynchronous underneath because judged checks await an LLM; the deterministic path runs through the event loop without ever yielding.)
validate_bundle_path_async begins with the module's most important
non-obvious lines β the imports that register the checks:
from . import attestations as _attestations # noqa: F401
from . import deterministic as _deterministic # noqa: F401
from . import judged as _judged # noqa: F401
Importing those modules executes every @register decorator
(base.py:146), which instantiates each check class and stores the
instance in the module-global REGISTRY dictionary, keyed by check id.
By the time any function in api.py runs, the registry is fully
populated. No registration happens at validation time; when stepping
through with a debugger, this has already occurred during import.
The bundle argument is then resolved (api.py:30-49):
- a directory is used as-is;
- a
.zipis extracted into atempfile.TemporaryDirectory(cleaned up in thefinally:block at the end), and_locate_bundle_roothandles a common packaging error β zipping the containing folder: ifcompetition.yamlis not at the zip root but exactly one subdirectory contains it, that subdirectory is used.
4. Stage 3 β constructing the context (the only parse)
Files: checks/facts.py, then checks/base.py
facts = CompetitionFacts.discover(bundle_dir, facts_path)
ctx = CheckContext.from_bundle_dir(bundle_dir, facts=facts)
CompetitionFacts (facts.py) implements the declare-then-verify
side channel: a six-field dataclass (anticipated_error_rate,
test_set_size, unit_of_generalization, external_data_allowed,
prizes, task_type) loaded from --facts <file>, else from
<bundle>/competition_facts.yaml, else left empty. Some checklist items
are unverifiable from the bundle alone β rather than guessing, checks
that need such a fact consume a declaration and verify against it.
Unknown keys in the YAML are a hard error, so typographical mistakes
cannot pass silently.
CheckContext.from_bundle_dir (base.py:74-84) is where the bundle's
configuration becomes data:
comp = None
yaml_path = bundle_dir / "competition.yaml"
if yaml_path.is_file():
try:
loaded = yaml.safe_load(yaml_path.read_text(encoding="utf-8"))
comp = loaded if isinstance(loaded, dict) else None
except yaml.YAMLError:
comp = None
return cls(bundle_dir=bundle_dir, comp=comp, facts=facts or CompetitionFacts())
The entire shared state of a validation run is these three fields:
bundle_dir (a Path), comp (the parsed YAML dictionary, or None
when the file is missing or corrupt β every check must tolerate that
case), and facts. The helper ctx.phases() (base.py:86-89) returns
comp["phases"] filtered to dictionaries, or [].
Observation: placing a breakpoint on the
returnline offrom_bundle_dirand inspectingcompin the Variables panel shows the exact dictionary that every subsequent check reads.
5. Stage 4 β the dispatch loop
File: checks/base.py:162 β run_checks(ctx)
for check in checks_for(tiers): # registry, sorted by (tier, id)
if check.tier == Tier.JUDGED:
continue # judged is async; dispatched separately
missing = check.missing_facts(ctx)
if missing:
results.append(check.skipped(f"requires facts not provided: ..."))
continue
results.extend(check.run(ctx))
Three rules are visible directly in the loop:
- Judged checks are skipped here. They require
await, soapi.pydispatches them after this loop (Stage 6). - A check whose
requires_factsare not declared reports SKIPPED with instructions; a missing fact is never a silent pass. - Each check returns a list of
CheckResults, since some checks emit one result per phase or per leaderboard column.
6. Stage 5 β inside a deterministic check (two worked examples)
File: checks/deterministic.py
Example A: DevPhaseDuration (line 78) β pure dictionary arithmetic
def run(self, ctx):
phases = ctx.phases()
if not phases:
return [self.skipped("no phases declared")]
first = phases[0]
start, end = _parse_date(first.get("start")), _parse_date(first.get("end"))
if start is None or end is None:
return [self.skipped("first phase start/end not parseable as dates", where="phases[0]")]
days = (end - start).days
if days >= 40:
return [self.passed(f"development phase runs {days} days", where="phases[0]")]
return [self.finding(f"development phase runs only {days} days β ...", where="phases[0]")]
The control flow proceeds as follows: obtain the phase list from the
shared dictionary; an absent phase list yields SKIP, since the schema
check owns that failure and checks do not double-report; parse the first
phase's start and end (_parse_date tries %Y-%m-%d %H:%M:%S, then
%Y-%m-%d, and also accepts values PyYAML has already converted to
datetime); unparseable dates are SKIPPED, never guessed; otherwise the
check is plain subtraction against a threshold drawn from the literature
(the citation on the class, PavΓ£o et al. (Ch. 13), is stamped onto
every result it emits).
The verdict vocabulary is significant: a too-short phase is a
finding, not a FAIL β such a configuration is legal, merely a
documented risk. The same pattern holds for TwoPhaseStructure
(line 59: len(ctx.phases()) >= 2), DailySubmissionCap (line 105:
reads max_submissions_per_day on every non-final phase),
FinalPhaseSubmissionLimit (reads max_submissions on the last phase),
and LeaderboardSortingDeclared (walks
comp["leaderboards"][i]["columns"][j]["sorting"]) β approximately 20
lines each, taking a dictionary in and producing cited results out.
Example B: SchemaLint (line 33) β the check that gates
def run(self, ctx):
report = validate_bundle(ctx.bundle_dir.name, str(ctx.bundle_dir.parent))
issues = report.get("issues") or []
if not issues:
return [self.passed("competition.yaml parses; all referenced files exist; ...")]
out = []
for issue in issues:
status = Status.FAIL if issue.get("severity") == "error" else Status.FINDING
...
This is the only deterministic check that leaves the YAML dictionary and
walks the filesystem β and it does so by delegating to the core layer:
core/bundle_io.py:validate_bundle (line 351), a pure-I/O lint that
checks, in order:
competition.yamlparses and contains all required top-level keys;- every file the YAML references actually exists (logo, terms, every page, every program directory, data directories, solutions);
- every leaderboard
column.keyappears as a key the scoring program writes (a static scan ofscore.pyforjson.dumpkeys β where certainty is unavailable, the lint warns rather than fails); - each scoring and ingestion program carries runnable metadata β
metadata.yamlor the legacy extensionlessmetadatafilename (production Codabench accepts both; the validator incorrectly gated the legacy form until the STYLE-TRANS-FAIR ground-truth bundle demonstrated otherwise β see CHANGELOG); - phases are sorted, non-overlapping, and reference declared tasks;
- tasks referenced from phases and solutions exist by index.
Error-severity issues become the report's gate failures; warnings become findings. This is why a structurally broken bundle fails validation while a merely risky design only accumulates advisories.
After the deterministic checks, the same loop runs
checks/attestations.py β five checks whose run() unconditionally
returns ATTESTATION_REQUIRED results ("external review of the task
design happened", "legal signed off on the prize terms", and so on).
Code cannot know these facts; the report surfaces them as unchecked
boxes rather than asserting them.
7. Stage 6 β the judged tier (only with --judged)
File: checks/judged.py
Returning to api.py:
if judged:
if backend is None:
from ..backends import get_claude_backend
backend = get_claude_backend()
results.extend(await run_judged_checks(ctx, backend))
run_judged_checks iterates over the same REGISTRY, selecting the
JudgedCheck instances. Each one:
- builds one rubric prompt (
build_prompt) β for example,DocsConfigConsistencyembeds the rawcompetition.yamltext (capped at 8k characters) and allpages/*.md(capped at 16k) and asks for contradictions only, quoted from both sides, as strict JSON; - runs it as a tool-less backend session:
await backend.run(AgentTask(prompt=prompt, allowed_tools=[]))β the judge cannot browse and cannot run code; it reads only what is in the prompt; - parses the verdict with
_extract_json(a fencedjsonblock or the first{...}in the reply). Unparseable output yields SKIPPED ("judge returned no parseable JSON verdict"); a failed session yields SKIPPED with the error. A pass is never fabricated.
The constructed results are FINDINGs (or an advisory pass). By
construction there is no code path from a judged check to Status.FAIL
β an LLM's opinion can warn, but it cannot gate. This is the falsifiable
answer to the objection that the system merely asks the model: the
property called valid is decided entirely by Stage 5.
8. Stage 7 β the report and the exit code
File: checks/report.py
Every result funnels into one dataclass:
@dataclass
class ValidationReport:
bundle_dir: Path
results: list[CheckResult]
facts: CompetitionFacts
@property
def ok(self) -> bool:
return not any(r.status == Status.FAIL for r in self.results)
The ok property expresses the entire gating policy in one line: only
deterministic FAILs block. to_markdown() renders five sections in
severity order β Gate failures (which must be fixed), Findings
(advisory), Attestations required (where an unchecked box is not
equivalent to completion), Skipped, and Passed β each line carrying
its check id, locator (where), message, and citation. to_dict()
produces the same content for --json.
Control returns to _cmd_validate, which prints the rendering and maps
report.ok to the process exit code: 0 indicates that no gates
failed, 1 indicates that at least one gate failed, and 2
indicates that --judged was refused for lack of authentication. That
exit code is what continuous-integration systems consume.
Following the trace live: breakpoints per stage
The following table lists one breakpoint per stage, together with the state worth inspecting at each.
| Stage | Breakpoint | What to inspect there |
|---|---|---|
| 1 CLI | cli/main.py β _cmd_validate first line |
args |
| 2 entry | checks/api.py β validate_bundle_path_async first line |
path, zip-vs-dir branch |
| 3 context | checks/base.py:84 (from_bundle_dir return) |
comp β the parsed YAML dict |
| 4 dispatch | checks/base.py:162 (run_checks loop) |
conditional bp: check.id == "dev-phase-duration" |
| 5 a check | checks/deterministic.py:86 (DevPhaseDuration.run) |
phases, days |
| 5 core lint | core/bundle_io.py:351 (validate_bundle) |
issues accumulating |
| 6 judged | checks/judged.py β parse_verdict |
text β the raw LLM reply |
| 7 report | checks/report.py β ok property |
self.results |
Launch with the "validate: ground-truth bundle (keyless)"
configuration in .vscode/launch.json, step into each call with F11,
and observe the Call Stack panel reproduce the call graph at the top of
this page.
Reproducing on real artifacts
# the human-built production bundle (passes the gates; 4 advisory findings,
# incl. uncapped submissions β which --judged then catches contradicting its own pages)
autocodabench validate benchmark/autocodabench_create_bench/competitions/style-trans-fair/ground_truth/bundle
# machine-readable
autocodabench validate <bundle> --json | python -m json.tool
# add the LLM-judged advisory tier (prompts for auth if you have none)
autocodabench validate <bundle> --judged
# the full check inventory, by tier, with citations
autocodabench checks list
Related reading
docs/architecture.mdβ the maintainer's map of all layers (this document covers thechecks/+core/slice).docs/scientific-validation.mdβ why the tiers exist and how each check's decision rule is justified.tests/test_checks.pyβ every behavior above as a minimal runnable example; "Debug Test" on any case drops you straight into a check.