g5-eval-dashboard / docs /PHASE7_EXPERIMENT_GUIDE.md
juandavidrodriguezar's picture
Upload folder using huggingface_hub
eaae571 verified
|
Raw
History Blame Contribute Delete
2.65 kB

Phase 7 Experiment Guide

This guide explains the current comparison workflow for before-vs-after mitigation experiments.

Purpose

Phase 7 is where the team compares:

  • a baseline run
  • a defended run on the same slice
  • the metric differences between them

The goal is to produce evidence for what improved, what remained weak, and what still needs discussion in the final report.

Main Comparison Pairs

Larger active-corpus pairs

  • baseline: configs/hackaprompt_override_25.yaml

  • defended: configs/hackaprompt_override_25_mitigated.yaml

  • baseline: configs/tensortrust_hijacking_25.yaml

  • defended: configs/tensortrust_hijacking_25_mitigated.yaml

Mock validation pairs

  • configs/hackaprompt_override_25_mock.yaml
  • configs/hackaprompt_override_25_mock_mitigated.yaml
  • configs/tensortrust_hijacking_25_mock.yaml
  • configs/tensortrust_hijacking_25_mock_mitigated.yaml

Focused validation pairs

  • configs/hackaprompt_direct_10.yaml
  • configs/hackaprompt_direct_10_mitigated.yaml
  • configs/agentdojo_tool_10.yaml
  • configs/agentdojo_tool_10_mitigated.yaml
  • configs/hackaprompt_adaptive_1_mitigated.yaml

How To Run

  1. Run the baseline config.
  2. Run the matching defended config.
  3. Compare the two summary.json files.

For exact hosted-run commands and a fill-in template, use:

  • docs/GROQ_RUNBOOK.md
  • docs/RESULTS_TEMPLATE.md

How To Compare

Use:

python scripts/compare_runs.py --baseline results/<baseline_run_id>/summary.json --defended results/<defended_run_id>/summary.json

Optional saved outputs:

python scripts/compare_runs.py --baseline results/<baseline_run_id>/summary.json --defended results/<defended_run_id>/summary.json --output results/comparisons/run_comparison.json --markdown-output results/comparisons/run_comparison.md

Output Coverage

The comparison script reports:

  • overall metric deltas
  • by-category deltas
  • by-surface deltas
  • by-source deltas
  • mitigation activity counts

What To Report

For each pair, record:

  • attack success rate before and after mitigation
  • successful attacks before and after mitigation
  • prompt leakage changes
  • refusal changes
  • manual-review changes
  • mitigation activity such as blocked, transformed, sanitized, and response-filtered cases
  • category-level differences

Current Position

The Phase 7 workflow is in strong shape:

  • documented
  • runnable with both mock and Groq
  • able to generate report-ready comparison outputs
  • already validated on real HackAPrompt and TensorTrust comparison pairs

Interpretation still matters, because not every benchmark slice improves equally.