g5-eval-dashboard / docs /ARCHITECTURE.md
juandavidrodriguezar's picture
Upload folder using huggingface_hub
eaae571 verified
|
Raw
History Blame Contribute Delete
3.22 kB
# Architecture Notes
## Core Components
- Attack Generator
- Attack Execution Engine
- Mitigation Layer
- Evaluation and Metrics
- Results Logging and Analysis
## Design Principles
- Modular interfaces for attacks, mitigations, and models
- Reproducible experiment configuration
- Clear separation between execution and evaluation
## Current MVP
The current MVP implements a small end-to-end evaluation pipeline:
1. Load normalized attack cases from `data/normalized/attack_cases.jsonl`
2. Select a small reproducible sample from config
3. Execute the sample through a model adapter
4. Score each case with a minimal attack-success evaluator
5. Save run artifacts under `results/<run_id>/`
The active runtime path is:
`normalized attack cases -> model adapter -> evaluator -> run artifacts`
## Canonical Data Contract
The canonical normalized source of truth is:
- `data/normalized/attack_cases.jsonl`
The enriched derivative is:
- `data/normalized/attack_cases.enriched.jsonl`
The enriched file adds taxonomy and evaluation-support metadata, but it is not
the primary source of truth. The normalized file remains the base corpus used
by the framework contract.
## Implemented Adapters
- `mock`
Used for local pipeline verification without network calls or token costs.
- `groq`
Used as the first real hosted inference backend for the MVP.
The model adapter boundary is intentionally provider-agnostic so later work can
add `ollama`, `transformers`, or `vllm` without changing orchestration.
## Run Artifacts
Each experiment run writes:
- `manifest.json`: config, model, dataset, and git metadata
- `case_results.jsonl`: one result per evaluated case
- `summary.json`: aggregate metrics and output paths
The manifest now records the config source path, dataset SHA-256, selected case
IDs, and filter/selection settings so a run can be replayed from the same
contract with deterministic case selection.
The full schema contract for normalized cases, enriched cases, configs, and run
artifacts is documented in:
- `docs/SCHEMA_CONTRACT.md`
## Current Evaluation Scope
The MVP currently implements one minimal metric:
- `attack_success`
A simple substring-based check using the case `expected_target`
This is intentionally narrow so the team can validate the full execution path
before adding category-specific success logic, richer leakage checks, and
mitigation-aware scoring.
The runner now also records heuristic evaluation support fields:
- `refusal`
- `prompt_leakage`
- `manual_review`
- `matched_rule`
These are not final research metrics, but they make the MVP more useful for
debugging, output review, and richer analysis while taxonomy work continues.
## Near-Term Scaling Path
The next steps after the MVP are:
- improve category-specific evaluation logic
- expand focused attack subset configs for stronger attack pressure
- use focused active-corpus slices such as HackAPrompt override/leakage and TensorTrust hijacking/extraction
- compare prompt execution modes such as `plain`, `context_first`, and `attack_then_task`
- add mitigation plugin hooks
- add more hosted or local model backends
- expand experiment configs beyond the MVP smoke and 20-case Groq runs