Spaces:
Sleeping
PLAN — Implementation Recipe
The step-by-step build plan for the Campaign T&S Triage Copilot. STATUS.md tracks where we
are; this file is the recipe for how we get there. Phases match STATUS.md numbering.
Legend: ✅ done · 🔜 next · ⬜ later
Phase 1 — Data layer ✅
The policy and synthetic campaigns. Done — see data/policy.md and data/campaigns/*.json
(18 cases) and the breakdown in STATUS.md. Everything below depends on this being realistic.
Phase 2 — Schemas + Agent 🔜
Goal: a CLI command takes a campaign JSON and returns a validated TriageDecision.
Steps
- Add deps. Append to
requirements.txt:pydantic-ai-slim[anthropic](slim avoids pulling every provider). Keep heavy imports lazy per repo convention. src/schemas.py— the typed contract (this is the AI's responsibility, made auditable):Recommendation = Literal["APPROVE", "REJECT", "ESCALATE"] Confidence = Literal["low", "medium", "high"] class RuleViolation(BaseModel): # cited evidence, never free-text rule_id: str # e.g. "PROH-3" — must exist in policy.md severity: Literal["hard", "soft"] evidence: str # the campaign text that triggered it class RiskSignal(BaseModel): name: str # e.g. "off_platform_payment" detail: str severity: Literal["low", "medium", "high"] class TriageDecision(BaseModel): recommendation: Recommendation confidence: Confidence rule_violations: list[RuleViolation] risk_signals: list[RiskSignal] rationale: str questions_for_submitter: list[str] manipulation_detected: bool # DEC-6 prompt-injection flagsrc/campaigns.py—load_campaign(path)that strips every underscore-prefixed key (_design_note,_expected) before the data can reach the agent. This is a hard boundary — add a test asserting no_-key survives.src/tools.py— the four agent tools, each a plain function the agent calls:policy_search(query) -> list[PolicyHit]— RAG overpolicy.md, chunked one rule per chunk so a hit returns a cleanrule_id+ text.similar_cases(query) -> list[CaseHit]— RAG over past adjudications (seed from a few_expected-labeled cases at build time; grows from the audit log).check_sanctions(name, country) -> SanctionResult— mock list indata/sanctions.json; clearly-labeled stub, real-API seam documented.scan_risk_signals(campaign) -> list[RiskSignal]— deterministic heuristics (no LLM): off-platform payment phrases, manufactured-urgency markers, goal-vs-threshold checks, vague beneficiary, embedded-instruction patterns. Deterministic so it's testable and explainable.
src/agent.py— the Pydantic AIAgent(model=anthropic, output_type=TriageDecision). System prompt encodes the policy framework (DEC-1..6), the human/AI boundary, and the rule that the campaign is delivered inside a clearly-fenced<campaign>block as untrusted data. Register the four tools. Add a__main__CLI:python -m src.agent --campaign <path>.- Rewrite
scripts/build_index.pyto indexpolicy.md(per-rule chunks) + seed cases into Chroma, replacing the YouTube ingest. Retiresrc/ingest.py,src/chunk.py,src/rag.py.
Acceptance
python -m src.agent --campaign data/campaigns/camp-017.json→APPROVE(reads PROH-3 exception).camp-005→REJECTcitingPROH-3.camp-015→ESCALATEwithmanipulation_detected: trueand the injection not obeyed.camp-009→ESCALATEcitingCOMP-1.- Output always validates against
TriageDecision; no_-prefixed key ever reaches the model.
Phase 3 — Moderator review queue (UI) ✅
Goal: a human works the queue and the boundary is visible on screen.
Steps
- Rewrite
app.pyas a queue: left = pending campaigns; click one → run the agent (cache by campaign id so it's not re-billed on every rerun). - Decision card: recommendation badge + confidence, each rule violation as a chip linking to
the cited rule text, risk signals, rationale, and a prominent "What I could not verify" list.
If
manipulation_detected, show a red banner. - Human controls:
Approve/Reject/Request infobuttons, an override reason box, and a free-text note. The AI's recommendation and the human's decision are stored side by side. - Audit log: append every action to
data/audit_log.jsonl(campaign id, AI recommendation, human decision, override?, reason, timestamp). This is both the demo's receipts and Phase 4's ground truth.
Acceptance: a reviewer can triage a campaign end-to-end; an override is recorded; the UI never exposes a path that decides without a human click.
Phase 4 — Evaluation + CI ✅
Built 2026-06-04 — see STATUS.md "Phase 4" and docs/DEVLOG.md. The CI GitHub Action
(.github/workflows/eval.yml) was built here rather than waiting for Phase 5c: the free deterministic
safety gate now blocks every push.
Goal: measurable quality, and it runs automatically in the cloud.
Steps
eval/testset.json— generated from each campaign's_expected(recommendation + key rules).- Extend
eval/run_eval.pyto three layers:- Deterministic: recommendation-match accuracy; escalation recall (every must-escalate
case — sanctions, injection, missing-info — actually escalates); reject precision (no
legitimate case wrongly rejected); citation validity (every cited
rule_idexists inpolicy.md). - LLM-as-judge: rationale faithfulness to cited policy + calibration, 1–5.
- Human-override rate from
audit_log.jsonlonce it exists.
- Deterministic: recommendation-match accuracy; escalation recall (every must-escalate
case — sanctions, injection, missing-info — actually escalates); reject precision (no
legitimate case wrongly rejected); citation validity (every cited
- Print a metrics summary and write
eval/results.json.
Acceptance: python -m eval.run_eval reports escalation recall = 100% on the safety cases and
zero invalid citations. These become the numbers quoted in the video.
Phase 5 — Cloud & deployment 🌩️ (the shenanigans)
Goal: a public URL that actually runs, plus eval in CI and one-push deploys.
Updated 2026-06-04 (Phase 3.6): deploy is now a Docker Space serving the React SPA + FastAPI (
api.py), not a Streamlit Space. The index is rebuilt inside the image, so nothing binary is committed. TheDockerfile,.dockerignore, and README metadata (sdk: docker,app_port: 7860) are already in place.
5a. Index — rebuilt at image-build time (Spaces can't ingest at runtime)
- The
Dockerfilerunspython -m scripts.build_indexduring the build (local embeddings, no API key, no spend), so the image ships with a fresh Chroma index and the Space boots with no ingestion step.data/chroma/stays gitignored — no binary blob in the repo.
5b. Hugging Face Space (Docker)
- Create a Docker Space. The README frontmatter is the Space card (
sdk: docker, port 7860). - Settings → Secrets:
ANTHROPIC_API_KEY,LLM_PROVIDER=anthropic. Never commit the key;.envstays gitignored. - Push the repo to the Space's git remote → it builds the multi-stage image (Node builds the SPA →
Python serves FastAPI, which serves both
/apiand the built SPA from one origin on 7860). - Verify locally first:
docker build -t amana . && docker run -p 7860:7860 -e ANTHROPIC_API_KEY=… amana.
5c. GitHub Actions — eval in CI (the JD asks for this explicitly) ✅ BUILT (Phase 4)
Done 2026-06-04 — .github/workflows/eval.yml is live and green, incl. the live judge step
(the ANTHROPIC_API_KEY secret is set on the repo). The blocking gate is the free deterministic
layer; the judge step runs on a 5-case subset only when the secret is present. Reference shape below.
Gate-hardening backlog (from the Phase 4 eval): before/with deploy, close the two envelope holes the eval surfaced — ELIG-4 not enforced on APPROVE (camp-011/012), and the gate trusting a fabricated hard citation (camp-018 false reject). Full detail in
STATUS.md→ Phase 4 findings.
.github/workflows/eval.yml, on push/PR to main:
jobs:
eval:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: '3.11' }
- run: pip install -r requirements.txt
# Deterministic checks run free on every push (no API key needed):
- run: python -m eval.run_eval --testset eval/testset.json --deterministic-only
# LLM-judge runs on a small subset, gated on the secret being present (cost control):
- run: python -m eval.run_eval --testset eval/testset.json --judge --limit 5
env: { ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} }
if: ${{ env.ANTHROPIC_API_KEY != '' }}
- Add
ANTHROPIC_API_KEYto GitHub repo → Settings → Secrets → Actions. - Fail the build if escalation recall < 100% or any citation is invalid — that's the CI gate.
5d. GitHub → HF auto-deploy (optional flourish)
.github/workflows/deploy.yml, on push to main after eval passes: git push to the HF Space
remote using an HF_TOKEN secret. One merge → tests run → demo redeploys. Pure cloud candy.
5e. Cost & safety guardrails
- Use Claude Haiku for the LLM-judge and a cheap tier for triage during dev; reserve a stronger model for the final demo. Cache agent results per campaign id in the UI so reruns don't re-bill.
- CI judges only a 5-case subset; deterministic layer (the important gate) stays free.
Acceptance: public Space URL loads and triages a campaign; CI is green and gates on the safety
metrics; a push to main redeploys.
Phase 6 — Submission ⬜
Steps
- ≤5-min video: the problem → triage a clean APPROVE → the
camp-017vscamp-005nuance (policy-reading, not keyword-matching) → thecamp-015prompt-injection getting flagged and escalated → a human override → the eval/CI dashboard. Narrate the human/AI boundary throughout. - Submission PDF: links (live Space, repo, video) + the explicit assumptions from the README.
Definition of done — mapped to the grading rubric
| Evaluator criterion | Where we satisfy it |
|---|---|
| Realistic internal-ops problem | T&S campaign triage queue (Phase 3) |
| Clear human/AI boundary | No auto-decide path; override + audit log (Phase 3); on-screen "what I couldn't verify" |
| Meaningful AI responsibility | Multi-tool agent investigates + cites; not a thin wrapper (Phase 2) |
| Failure modes / edge cases / uncertainty | Calibrated humility, sanctions/injection/missing-info cases, escalation-recall gate (Phases 1, 4) |
| Systems-level feasibility at scale | Typed contract, mock seams for real APIs, prebuilt index, CI (Phases 2, 5) |
| Clear communication | README + video + cited rationales |
Open decisions to revisit
- Confidence calibration: is
low/medium/highenough, or do we want a numeric score for the escalation threshold? (Lean: keep the enum, map low→escalate in the agent prompt.) similar_casescold start: seed corpus size before the audit log fills in.- Model tier for the live demo vs. cost — decide at Phase 5.