DAO_kdd26 / docs /investigation /evaluation_consistency_audit.md
sipe5001's picture
Add Hugging Face Docker Space configuration
d3d0e0e
|
Raw
History Blame Contribute Delete
6.54 kB

Evaluation Consistency Audit (Phase-2 AAT Migration)

Date: 2026-06-14 Run validated: 20260614T044356Z

Scope

Deep consistency audit across:

  • AAT phase aggregation
  • Legacy stage aggregation
  • Agent resource accounting
  • Confidence semantics
  • Verification semantics
  • Conservation checks (token/time)

Check 1: AAT Phase Tokens vs Agent Tokens

  • Severity: High
  • Root cause: AAT phase token display mixed incompatible definitions (AAT-phase-only vs legacy-stage totals), so PLAN looked artificially small versus Planner + Plan Critic agent totals.
  • Files: src/data_agent_baseline/langgraph_agent/eval_v2_viz.py
  • Current formula (before):
    • AAT phase tokens: understanding_tokens/planning_tokens/... (from AAT observability)
    • Legacy tokens shown separately in another table
    • No reconciliation in primary phase table
  • Expected formula:
    • Primary phase table should show reconciled totals from both views, with duplicate-safe merge.
  • Fix:
    • Added reconciliation logic in render_per_stage_timing.
    • Displayed phase tokens now merge AAT + legacy components with duplicate detection tolerance.
  • Validation:
    • PLAN now displays 70,747 tokens, reconcilable with Planner + Plan Critic + coordinator planning contribution.

Check 2: Phase Time vs Agent Time

  • Severity: High
  • Root cause: One time notion was displayed without distinguishing cumulative compute vs displayed phase time, causing confusion under parallelism/concurrency.
  • Files: src/data_agent_baseline/langgraph_agent/eval_v2_viz.py
  • Current formula (before):
    • Single time column in AAT Execution Phases, no reconciliation table.
  • Expected formula:
    • Explicitly report both displayed phase time and cumulative compute time.
  • Fix:
    • Added Phase Time Reconciliation table.
    • Shows Displayed Time vs Cumulative Compute with interpretation guidance.
  • Validation:
    • UNDERSTAND now clearly shows compute can exceed displayed time when work overlaps.

Check 3: EXECUTE Confidence = 0

  • Severity: Medium
  • Root cause: execution_confidence was hardcoded to 0.0 when executor confidence is unavailable.
  • Files: src/data_agent_baseline/langgraph_agent/eval_v2.py
  • Current formula (before): execution_confidence = 0.0
  • Expected formula: show unavailable as N/A, not numeric zero.
  • Fix:
    • Set execution_confidence to NaN.
    • Rendering already maps NaN to N/A.
  • Validation:
    • EXECUTE confidence now renders as N/A.

Check 4: Verification Timeline Logic/Terminology

  • Severity: Medium
  • Root cause: "Failed Checks" wording implied hard failure even when task outcome was PASS after recovery.
  • Files: src/data_agent_baseline/langgraph_agent/eval_v2_viz.py
  • Fix:
    • Renamed columns:
      • Failed Checks -> Detected Issues
      • Warnings -> Tool Failures
  • Validation:
    • task_11 PASS with Detected Issues=1 is now semantically aligned with recoverable/retried behavior.

Check 5: Zero Variance Metrics (Std=0)

  • Severity: High
  • Root cause: research std-dev columns were derived by label-to-column string transformation; this broke for Data Understanding and other labels.
  • Files: src/data_agent_baseline/langgraph_agent/eval_v2_viz.py
  • Current formula (before): col_name derived from label text (brittle).
  • Expected formula: explicit label->column mapping.
  • Fix:
    • Replaced string transform with explicit metric mapping list.
  • Validation:
    • Research output now shows Mean Data Understanding std = 0.760 (not 0.000).

Check 6: Token Conservation Audit

  • Severity: High
  • Root cause: no explicit token conservation validator existed.
  • Files: src/data_agent_baseline/langgraph_agent/eval_v2_validator.py, tests/test_eval_v2_validator.py
  • Fix:
    • Added token_conservation_mismatch warning check.
    • Added phase_token_reconciliation_mismatch warning check.
  • Validation:
    • Run now emits warnings when phase/component token totals diverge beyond threshold.

Check 7: Time Conservation Audit

  • Severity: Medium
  • Root cause: no explicit time conservation checks.
  • Files: src/data_agent_baseline/langgraph_agent/eval_v2_validator.py
  • Fix:
    • Added time_conservation_mismatch warning check.
    • Added negative_duration error check.
  • Validation:
    • Validator now checks impossible/implausible timing relationships while allowing concurrency.

Check 8: Resource Accounting Audit

  • Severity: High
  • Root cause: validator lacked a direct failures <= calls invariant.
  • Files: src/data_agent_baseline/langgraph_agent/eval_v2_validator.py, tests/test_eval_v2_validator.py
  • Fix:
    • Added failures_exceed_calls error check across planner/critic/execute/coordinator/verifier/synthesizer/summary.
  • Validation:
    • Added unit test test_validator_catches_failures_exceed_calls.

Check 9: Legacy vs AAT Consistency

  • Severity: Medium
  • Root cause: no explicit mismatch checks between legacy activity and AAT phase fields.
  • Files: src/data_agent_baseline/langgraph_agent/eval_v2_validator.py
  • Fix:
    • Added:
      • legacy_aat_plan_mismatch
      • legacy_aat_verify_mismatch
  • Validation:
    • Validator now flags zero phase values when legacy activity exists.

Check 10: Audit Report

  • Severity: Informational
  • Output: this file documents root causes, formulas, fixes, and validation outcomes.

Added/Updated Validation Rules

New validator checks introduced:

  • negative_duration (error)
  • failures_exceed_calls (error)
  • token_conservation_mismatch (warning)
  • time_conservation_mismatch (warning)
  • legacy_aat_plan_mismatch (error)
  • legacy_aat_verify_mismatch (error)
  • phase_token_reconciliation_mismatch (warning)

Unit Test Coverage

  • Existing tests retained and passing.
  • Added:
    • test_validator_catches_failures_exceed_calls
    • test_validator_catches_phase_token_reconciliation_mismatch

Test run:

  • 10 passed in tests/test_eval_v2_validator.py + tests/test_eval_v2_aat_observability.py

Real Run Validation

Command:

  • dabench eval-v2 20260614T044356Z --mode verbose

Observed:

  • No hard consistency errors.
  • 3 warnings for phase_token_reconciliation_mismatch (expected under mixed legacy/AAT trace sources).
  • Phase table now reconciles PLAN/VERIFY token magnitudes with agent usage.
  • EXECUTE confidence shown as N/A.
  • Verification timeline terminology aligned with behavior.

Notes on Remaining Warnings

Current warnings indicate residual divergence between phase-level and component-level token accounting under mixed trace sources. This is now surfaced explicitly (instead of silently misreporting), and is suitable for follow-up instrumentation tightening at trace source level.