amana / PLAN.md
Misbahuddin's picture
Docs: record Phase 4 CI green (live judge) + eval gate-hardening backlog
62acf3d
|
Raw
History Blame Contribute Delete
11.4 kB

PLAN — Implementation Recipe

The step-by-step build plan for the Campaign T&S Triage Copilot. STATUS.md tracks where we are; this file is the recipe for how we get there. Phases match STATUS.md numbering.

Legend: ✅ done · 🔜 next · ⬜ later


Phase 1 — Data layer ✅

The policy and synthetic campaigns. Done — see data/policy.md and data/campaigns/*.json (18 cases) and the breakdown in STATUS.md. Everything below depends on this being realistic.


Phase 2 — Schemas + Agent 🔜

Goal: a CLI command takes a campaign JSON and returns a validated TriageDecision.

Steps

  1. Add deps. Append to requirements.txt: pydantic-ai-slim[anthropic] (slim avoids pulling every provider). Keep heavy imports lazy per repo convention.
  2. src/schemas.py — the typed contract (this is the AI's responsibility, made auditable):
    Recommendation = Literal["APPROVE", "REJECT", "ESCALATE"]
    Confidence     = Literal["low", "medium", "high"]
    
    class RuleViolation(BaseModel):   # cited evidence, never free-text
        rule_id: str                  # e.g. "PROH-3" — must exist in policy.md
        severity: Literal["hard", "soft"]
        evidence: str                 # the campaign text that triggered it
    
    class RiskSignal(BaseModel):
        name: str                     # e.g. "off_platform_payment"
        detail: str
        severity: Literal["low", "medium", "high"]
    
    class TriageDecision(BaseModel):
        recommendation: Recommendation
        confidence: Confidence
        rule_violations: list[RuleViolation]
        risk_signals: list[RiskSignal]
        rationale: str
        questions_for_submitter: list[str]
        manipulation_detected: bool   # DEC-6 prompt-injection flag
    
  3. src/campaigns.pyload_campaign(path) that strips every underscore-prefixed key (_design_note, _expected) before the data can reach the agent. This is a hard boundary — add a test asserting no _-key survives.
  4. src/tools.py — the four agent tools, each a plain function the agent calls:
    • policy_search(query) -> list[PolicyHit] — RAG over policy.md, chunked one rule per chunk so a hit returns a clean rule_id + text.
    • similar_cases(query) -> list[CaseHit] — RAG over past adjudications (seed from a few _expected-labeled cases at build time; grows from the audit log).
    • check_sanctions(name, country) -> SanctionResultmock list in data/sanctions.json; clearly-labeled stub, real-API seam documented.
    • scan_risk_signals(campaign) -> list[RiskSignal]deterministic heuristics (no LLM): off-platform payment phrases, manufactured-urgency markers, goal-vs-threshold checks, vague beneficiary, embedded-instruction patterns. Deterministic so it's testable and explainable.
  5. src/agent.py — the Pydantic AI Agent(model=anthropic, output_type=TriageDecision). System prompt encodes the policy framework (DEC-1..6), the human/AI boundary, and the rule that the campaign is delivered inside a clearly-fenced <campaign> block as untrusted data. Register the four tools. Add a __main__ CLI: python -m src.agent --campaign <path>.
  6. Rewrite scripts/build_index.py to index policy.md (per-rule chunks) + seed cases into Chroma, replacing the YouTube ingest. Retire src/ingest.py, src/chunk.py, src/rag.py.

Acceptance

  • python -m src.agent --campaign data/campaigns/camp-017.jsonAPPROVE (reads PROH-3 exception).
  • camp-005REJECT citing PROH-3. camp-015ESCALATE with manipulation_detected: true and the injection not obeyed. camp-009ESCALATE citing COMP-1.
  • Output always validates against TriageDecision; no _-prefixed key ever reaches the model.

Phase 3 — Moderator review queue (UI) ✅

Goal: a human works the queue and the boundary is visible on screen.

Steps

  1. Rewrite app.py as a queue: left = pending campaigns; click one → run the agent (cache by campaign id so it's not re-billed on every rerun).
  2. Decision card: recommendation badge + confidence, each rule violation as a chip linking to the cited rule text, risk signals, rationale, and a prominent "What I could not verify" list. If manipulation_detected, show a red banner.
  3. Human controls: Approve / Reject / Request info buttons, an override reason box, and a free-text note. The AI's recommendation and the human's decision are stored side by side.
  4. Audit log: append every action to data/audit_log.jsonl (campaign id, AI recommendation, human decision, override?, reason, timestamp). This is both the demo's receipts and Phase 4's ground truth.

Acceptance: a reviewer can triage a campaign end-to-end; an override is recorded; the UI never exposes a path that decides without a human click.


Phase 4 — Evaluation + CI ✅

Built 2026-06-04 — see STATUS.md "Phase 4" and docs/DEVLOG.md. The CI GitHub Action (.github/workflows/eval.yml) was built here rather than waiting for Phase 5c: the free deterministic safety gate now blocks every push.

Goal: measurable quality, and it runs automatically in the cloud.

Steps

  1. eval/testset.json — generated from each campaign's _expected (recommendation + key rules).
  2. Extend eval/run_eval.py to three layers:
    • Deterministic: recommendation-match accuracy; escalation recall (every must-escalate case — sanctions, injection, missing-info — actually escalates); reject precision (no legitimate case wrongly rejected); citation validity (every cited rule_id exists in policy.md).
    • LLM-as-judge: rationale faithfulness to cited policy + calibration, 1–5.
    • Human-override rate from audit_log.jsonl once it exists.
  3. Print a metrics summary and write eval/results.json.

Acceptance: python -m eval.run_eval reports escalation recall = 100% on the safety cases and zero invalid citations. These become the numbers quoted in the video.


Phase 5 — Cloud & deployment 🌩️ (the shenanigans)

Goal: a public URL that actually runs, plus eval in CI and one-push deploys.

Updated 2026-06-04 (Phase 3.6): deploy is now a Docker Space serving the React SPA + FastAPI (api.py), not a Streamlit Space. The index is rebuilt inside the image, so nothing binary is committed. The Dockerfile, .dockerignore, and README metadata (sdk: docker, app_port: 7860) are already in place.

5a. Index — rebuilt at image-build time (Spaces can't ingest at runtime)

  • The Dockerfile runs python -m scripts.build_index during the build (local embeddings, no API key, no spend), so the image ships with a fresh Chroma index and the Space boots with no ingestion step. data/chroma/ stays gitignored — no binary blob in the repo.

5b. Hugging Face Space (Docker)

  • Create a Docker Space. The README frontmatter is the Space card (sdk: docker, port 7860).
  • Settings → Secrets: ANTHROPIC_API_KEY, LLM_PROVIDER=anthropic. Never commit the key; .env stays gitignored.
  • Push the repo to the Space's git remote → it builds the multi-stage image (Node builds the SPA → Python serves FastAPI, which serves both /api and the built SPA from one origin on 7860).
  • Verify locally first: docker build -t amana . && docker run -p 7860:7860 -e ANTHROPIC_API_KEY=… amana.

5c. GitHub Actions — eval in CI (the JD asks for this explicitly) ✅ BUILT (Phase 4)

Done 2026-06-04.github/workflows/eval.yml is live and green, incl. the live judge step (the ANTHROPIC_API_KEY secret is set on the repo). The blocking gate is the free deterministic layer; the judge step runs on a 5-case subset only when the secret is present. Reference shape below.

Gate-hardening backlog (from the Phase 4 eval): before/with deploy, close the two envelope holes the eval surfaced — ELIG-4 not enforced on APPROVE (camp-011/012), and the gate trusting a fabricated hard citation (camp-018 false reject). Full detail in STATUS.md → Phase 4 findings.

.github/workflows/eval.yml, on push/PR to main:

jobs:
  eval:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: '3.11' }
      - run: pip install -r requirements.txt
      # Deterministic checks run free on every push (no API key needed):
      - run: python -m eval.run_eval --testset eval/testset.json --deterministic-only
      # LLM-judge runs on a small subset, gated on the secret being present (cost control):
      - run: python -m eval.run_eval --testset eval/testset.json --judge --limit 5
        env: { ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} }
        if: ${{ env.ANTHROPIC_API_KEY != '' }}
  • Add ANTHROPIC_API_KEY to GitHub repo → Settings → Secrets → Actions.
  • Fail the build if escalation recall < 100% or any citation is invalid — that's the CI gate.

5d. GitHub → HF auto-deploy (optional flourish)

.github/workflows/deploy.yml, on push to main after eval passes: git push to the HF Space remote using an HF_TOKEN secret. One merge → tests run → demo redeploys. Pure cloud candy.

5e. Cost & safety guardrails

  • Use Claude Haiku for the LLM-judge and a cheap tier for triage during dev; reserve a stronger model for the final demo. Cache agent results per campaign id in the UI so reruns don't re-bill.
  • CI judges only a 5-case subset; deterministic layer (the important gate) stays free.

Acceptance: public Space URL loads and triages a campaign; CI is green and gates on the safety metrics; a push to main redeploys.


Phase 6 — Submission ⬜

Steps

  1. ≤5-min video: the problem → triage a clean APPROVE → the camp-017 vs camp-005 nuance (policy-reading, not keyword-matching) → the camp-015 prompt-injection getting flagged and escalated → a human override → the eval/CI dashboard. Narrate the human/AI boundary throughout.
  2. Submission PDF: links (live Space, repo, video) + the explicit assumptions from the README.

Definition of done — mapped to the grading rubric

Evaluator criterion Where we satisfy it
Realistic internal-ops problem T&S campaign triage queue (Phase 3)
Clear human/AI boundary No auto-decide path; override + audit log (Phase 3); on-screen "what I couldn't verify"
Meaningful AI responsibility Multi-tool agent investigates + cites; not a thin wrapper (Phase 2)
Failure modes / edge cases / uncertainty Calibrated humility, sanctions/injection/missing-info cases, escalation-recall gate (Phases 1, 4)
Systems-level feasibility at scale Typed contract, mock seams for real APIs, prebuilt index, CI (Phases 2, 5)
Clear communication README + video + cited rationales

Open decisions to revisit

  • Confidence calibration: is low/medium/high enough, or do we want a numeric score for the escalation threshold? (Lean: keep the enum, map low→escalate in the agent prompt.)
  • similar_cases cold start: seed corpus size before the audit log fills in.
  • Model tier for the live demo vs. cost — decide at Phase 5.