| # Evaluation Pipeline |
|
|
| Measures how well a knowledge-editing method suppresses toilet hallucination in LLaVA while preserving other capabilities. |
|
|
| ## Quick Start |
|
|
| ```bash |
| # LoRA adapter |
| python -m experiment.evaluation.validate \ |
| --model_type lora \ |
| --model_dir step3_lora_outputs/run_YYYYMMDD/lora_adapter |
| |
| # Merged model |
| python -m experiment.evaluation.validate \ |
| --model_type merged \ |
| --model_dir step3_lora_outputs/run_YYYYMMDD/merged_model |
| |
| # Delta-W checkpoint (transformers backend only) |
| python -m experiment.evaluation.validate \ |
| --model_type delta_w \ |
| --checkpoint step3_outputs/final_model.pt \ |
| --inference_backend transformers |
| ``` |
|
|
| ## Image Categories |
|
|
| | Category | bathroom | toilet | Role | |
| |----------|----------|--------|------| |
| | `bathroom_no_toilet` | 1 | 0 | **Edit target** — hallucination should be suppressed | |
| | `bathroom_with_toilet` | 1 | 1 | **Preserved** — should still mention toilet | |
| | `non_bathroom_with_toilet` | 0 | 1 | **Preserved** — toilet in other contexts must survive | |
| | `unrelated` | 0 | 0 | **Preserved** — general captioning must be unaffected | |
|
|
| ## Metrics |
|
|
| ### Per-Model |
|
|
| | Metric | Description | |
| |--------|-------------| |
| | `mention_rate_keyword` | Fraction of outputs containing an un-negated toilet synonym (fast regex). | |
| | `mention_rate_llm` | Fraction where VLM judge says toilet is present (more accurate, expensive). | |
| | `avg_clip_score` | Mean image–text cosine similarity (SigLIP). Only on "Describe this image." prompts. | |
| | `avg_caption_quality` | Mean caption quality score (1–5) from VLM judge. | |
| | `quality_gated_mention_rate` | Mention rate filtered to outputs with caption_quality >= 3. | |
| |
| ### KME Metrics |
| |
| | Metric | Description | Category | |
| |--------|-------------|----------| |
| | **Efficacy** | `1 - mention_rate` on edit target. Did the edit work? | `bathroom_no_toilet` | |
| | **Generality** | Efficacy split by seen vs. unseen prompts. Does suppression generalize? | `bathroom_no_toilet` | |
| | **Locality** | Text similarity (exact match, ROUGE-L, BERTScore) between original and finetuned on preserved images. | all preserved categories | |
| | **Consistency** | Locality pooled across all preserved categories. | all preserved | |
|
|
| **Generality prompts:** |
| - Seen (train): `"Describe this image."`, `"Is there a toilet in this image?"` |
| - Unseen: `"What objects are in this picture?"`, `"Tell me what you see."`, `"What is in this bathroom?"` |
|
|
| ## Files |
|
|
| | File | Purpose | |
| |------|---------| |
| | `validate.py` | Entry point, CLI, orchestration | |
| | `inference.py` | Generates outputs for both models (HuggingFace or vLLM) | |
| | `metrics.py` | `KeywordMentionDetector`, `CLIPScorer`, `LocalLLMJudge`, `TextSimilarityScorer` | |
| | `metric.py` | `evaluate_collected_outputs()`, `compute_kme_metrics()` | |
| | `summary.py` | Formats results to stdout | |
|
|
| ## Output |
|
|
| Results saved to `--output_dir/<run_name>/`: |
| - `validation_results.json` — aggregate metrics + config |
| - `validation_details.json` — per-sample detail records |
|
|
| ## Key CLI Arguments |
|
|
| | Argument | Default | Description | |
| |----------|---------|-------------| |
| | `--model_type` | *(required)* | `lora`, `merged`, or `delta_w` | |
| | `--base_model_name` | `llava-hf/llava-1.5-7b-hf` | Base model ID | |
| | `--model_dir` | — | LoRA adapter or merged model path | |
| | `--checkpoint` | — | `.pt` state dict (delta_w only) | |
| | `--inference_backend` | `vllm` | `transformers` or `vllm` | |
| | `--num_per_category` | `50` | Max images per category | |
| | `--mention_method` | `both` | `keyword`, `llm`, or `both` | |
| | `--output_dir` | `./step4_v2_outputs` | Root output directory | |
|
|
| ## Optional Dependencies |
|
|
| | Package | Unlocks | |
| |---------|---------| |
| | `rouge-score` | ROUGE-L locality/consistency | |
| | `bert-score` | BERTScore locality/consistency | |
| | `peft` | LoRA model loading | |
| | `vllm` | Fast batched inference | |
|
|
| Missing packages degrade gracefully (metrics show `N/A`). |
|
|