{ "_provenance": "Hand-transcribed from the Scoring and Anti-gaming sections of docs/je_validation_task_spec_3.html, with the component algebra shown in the worked sample run. Verified 2026-08-04. If source_sha256 no longer matches the HTML, re-verify this file against the HTML and update the hash.", "source_sha256": "fd16867045b1752f60636a50eeca822fcd8d1998ac359002854251798a03cb6d", "components": { "outcome": "Deterministic. Detection F1 vs seeded defects, issue-type accuracy, exactness where a figure is specified. Per-task hard-fails force 0. Worked example (JE-06): F1 0.813 x issue-type accuracy 0.923 = 0.75.", "process": "Deterministic. Evidence coverage and precision, required-call gates, efficiency over redundant calls. Worked example (JE-06): 0.5*coverage(0.62) + 0.3*precision(0.93) + 0.2*efficiency(0.85) = 0.76.", "rationale": "LLM judge, four global criteria plus per-task focus. The only non-deterministic component - keep it isolated so drift is measurable. Worked example: 7/12 = 0.583." }, "combine": "score = w_outcome*outcome + w_process*process + w_rationale*rationale, weights per task JSON. Worked example (JE-06 weights .55/.25/.20): 0.72. The three components are ALWAYS reported separately - a single headline number hides a fabricated citation behind decent detection.", "judge_rubric": { "scale": "each criterion 0-3", "criteria": { "identification": { "full_credit": "Names the defect precisely, classifies it into the right family.", "zero": "Wrong defect, or right verdict with none identified." }, "evidence_grounding": { "full_credit": "Cites the specific document, clause or field that reveals it; all citations were opened.", "zero": "Cites nothing, or something never opened." }, "reasoning_validity": { "full_credit": "Chain survives a counterfactual - change the evidence, the conclusion changes.", "zero": "Right answer via invalid reasoning." }, "disposition": { "full_credit": "Verdict follows the finding, materiality addressed, actionable without redoing the work.", "zero": "Verdict inconsistent with the finding." } }, "pass_bar": ">=9/12 total AND no criterion at 0. A criterion at 0 fails the pass bar outright regardless of total." }, "judge_protocol": [ "Judge sees rationale, evidence log and ground truth - never model identity.", "Criteria scored in separate calls.", "3 judges, median, on L4 tasks.", "20 human-graded rationales per family to measure judge-SME agreement before trusting any number." ], "anti_gaming_build_day_one": [ "Flag-everything baseline must score clearly negative on every task, or the precision penalty is too weak. (NOTE: how 'negative' is realized against [0,1] components is an open decision - see ROADMAP open-decisions table, settle by Phase 2.)", "No-evidence baseline must approach zero on L2+, or a defect is leaking into the ledger row.", ">=5 seeds per task; report variance - single-seed model differences are noise.", "Citation audit of every evidence_ids against the open-document log, every episode." ] }