ToiTenBao's picture
Upload hallucination folder
a2ffd07 verified
|
Raw
History Blame Contribute Delete
3.9 kB
# Evaluation Pipeline
Measures how well a knowledge-editing method suppresses toilet hallucination in LLaVA while preserving other capabilities.
## Quick Start
```bash
# LoRA adapter
python -m experiment.evaluation.validate \
--model_type lora \
--model_dir step3_lora_outputs/run_YYYYMMDD/lora_adapter
# Merged model
python -m experiment.evaluation.validate \
--model_type merged \
--model_dir step3_lora_outputs/run_YYYYMMDD/merged_model
# Delta-W checkpoint (transformers backend only)
python -m experiment.evaluation.validate \
--model_type delta_w \
--checkpoint step3_outputs/final_model.pt \
--inference_backend transformers
```
## Image Categories
| Category | bathroom | toilet | Role |
|----------|----------|--------|------|
| `bathroom_no_toilet` | 1 | 0 | **Edit target** — hallucination should be suppressed |
| `bathroom_with_toilet` | 1 | 1 | **Preserved** — should still mention toilet |
| `non_bathroom_with_toilet` | 0 | 1 | **Preserved** — toilet in other contexts must survive |
| `unrelated` | 0 | 0 | **Preserved** — general captioning must be unaffected |
## Metrics
### Per-Model
| Metric | Description |
|--------|-------------|
| `mention_rate_keyword` | Fraction of outputs containing an un-negated toilet synonym (fast regex). |
| `mention_rate_llm` | Fraction where VLM judge says toilet is present (more accurate, expensive). |
| `avg_clip_score` | Mean image–text cosine similarity (SigLIP). Only on "Describe this image." prompts. |
| `avg_caption_quality` | Mean caption quality score (1–5) from VLM judge. |
| `quality_gated_mention_rate` | Mention rate filtered to outputs with caption_quality >= 3. |
### KME Metrics
| Metric | Description | Category |
|--------|-------------|----------|
| **Efficacy** | `1 - mention_rate` on edit target. Did the edit work? | `bathroom_no_toilet` |
| **Generality** | Efficacy split by seen vs. unseen prompts. Does suppression generalize? | `bathroom_no_toilet` |
| **Locality** | Text similarity (exact match, ROUGE-L, BERTScore) between original and finetuned on preserved images. | all preserved categories |
| **Consistency** | Locality pooled across all preserved categories. | all preserved |
**Generality prompts:**
- Seen (train): `"Describe this image."`, `"Is there a toilet in this image?"`
- Unseen: `"What objects are in this picture?"`, `"Tell me what you see."`, `"What is in this bathroom?"`
## Files
| File | Purpose |
|------|---------|
| `validate.py` | Entry point, CLI, orchestration |
| `inference.py` | Generates outputs for both models (HuggingFace or vLLM) |
| `metrics.py` | `KeywordMentionDetector`, `CLIPScorer`, `LocalLLMJudge`, `TextSimilarityScorer` |
| `metric.py` | `evaluate_collected_outputs()`, `compute_kme_metrics()` |
| `summary.py` | Formats results to stdout |
## Output
Results saved to `--output_dir/<run_name>/`:
- `validation_results.json` — aggregate metrics + config
- `validation_details.json` — per-sample detail records
## Key CLI Arguments
| Argument | Default | Description |
|----------|---------|-------------|
| `--model_type` | *(required)* | `lora`, `merged`, or `delta_w` |
| `--base_model_name` | `llava-hf/llava-1.5-7b-hf` | Base model ID |
| `--model_dir` | — | LoRA adapter or merged model path |
| `--checkpoint` | — | `.pt` state dict (delta_w only) |
| `--inference_backend` | `vllm` | `transformers` or `vllm` |
| `--num_per_category` | `50` | Max images per category |
| `--mention_method` | `both` | `keyword`, `llm`, or `both` |
| `--output_dir` | `./step4_v2_outputs` | Root output directory |
## Optional Dependencies
| Package | Unlocks |
|---------|---------|
| `rouge-score` | ROUGE-L locality/consistency |
| `bert-score` | BERTScore locality/consistency |
| `peft` | LoRA model loading |
| `vllm` | Fast batched inference |
Missing packages degrade gracefully (metrics show `N/A`).